Method and apparatus for intentional programming for heterogeneous systems
The algorithm is promoted to DSL representation through code conversion and variant generator. Combined with the runtime scheduler to optimize the performance of heterogeneous systems, the problems of long compilation time and poor performance in the existing technology are solved, and efficient resource utilization of heterogeneous systems is achieved.
Patent Information
- Application Number
- CN202510356087.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-27
- Filing Date
- 2020-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
When compiling algorithms are used in heterogeneous systems, the prior art fails to effectively consider the load and real-time performance of each processing element, resulting in long compilation time and poor performance, making it difficult to fully utilize the computing resources of heterogeneous systems.
Using sample code converters and variant generators, the performance of heterogeneous systems is optimized by promoting the algorithm of imperative programming language representation to DSL representation and generating multiple variant binaries, combined with the runtime scheduler to dynamically select the best execution path.
提高了在异构系统上的运行时性能,动态选择最优调度策略,减少了编译时间,充分利用了异构系统的计算资源。
Smart Images

Figure CN120295633A_ABST
Abstract
Description
This application is a divisional application of the invention patent application with the application date of March 26, 2020, the application number of 202010223629.8, and the title of "Methods and Apparatuses for Intentional Programming for Heterogeneous Systems". Technical Field
[0001] This disclosure generally relates to machine learning, and more particularly, to methods and apparatuses for intentional programming for heterogeneous systems. Background Art
[0002] Computer hardware manufacturers develop hardware components for use in various components of computer platforms. For example, computer hardware manufacturers develop motherboards, chip sets for motherboards, central processing units (CPUs), graphics processing units (GPUs), vision processing units (VPUs), field programmable gate arrays (FPGAs), hard disk drives (HDDs), solid state drives (SSDs), and other computer components. Many computer hardware manufacturers develop programs and / or other methods to compile algorithms and / or other code for running on specific processing platforms. Brief Description of the Drawings
[0003] Figure 1 A block diagram is depicted that illustrates an example heterogeneous system.
[0004] Figure 2 A block diagram is depicted that illustrates an example software tuning system, which includes a first software tuning system and a second software tuning system for training an example machine learning / artificial intelligence model.
[0005] Figure 3 A block diagram is depicted that illustrates an example variant generation system, which includes an example variant application that includes an example variant generator and an example code converter, and the example variant generator and the example code converter can be used to implement Figure 2 the first and / or second software tuning systems.
[0006] Figure 4 A block diagram is depicted that illustrates Figure 3 an example implementation of the example variant generator.
[0007] Figure 5 An example FAT binary file is depicted, which includes an example variant library, an example jump table library, and an example runtime scheduler for implementing the examples disclosed herein.
[0008] Figure 6 A block diagram is depicted that illustrates Figure 3 an example implementation of the example code converter.
[0009] Figure 7Depicts an example application code including example intentional code.
[0010] Figure 8 Is an example workflow for converting Figure 7 the example application code into a representation in a domain-specific language.
[0011] Figure 9 Is an example workflow for converting Figure 7 the example application code into example variant code.
[0012] Figure 10 Is an example workflow for compiling Figure 3 and / or Figure 5 the example FAT binary file.
[0013] Figure 11 Is a flowchart representing example machine-readable instructions that can be executed to implement Figure 3 、 Figure 4 、 Figure 6 and / or Figure 10 the example code converter and / or Figure 3 、 Figure 4 、 Figure 6 、 Figure 9 and / or Figure 10 the example variant generator for invoking an application to execute the (multiple) workloads on a heterogeneous system.
[0014] Figure 12 Is a flowchart representing example machine-readable instructions that can be executed to implement Figure 3 、 Figure 4 、 Figure 6 、 Figure 9 and / or Figure 10 the variant generator for compiling the (multiple) variants.
[0015] Figure 13 Is a flowchart representing example machine-readable instructions that can be executed to implement Figure 1 、 Figure 3 、 Figure 4 and / or Figure 6 the heterogeneous system for invoking an application to execute the (multiple) workloads.
[0016] Figure 14 Is a flowchart representing example machine-readable instructions that can be executed to implement during the training phase Figure 3 、 Figure 4 、 Figure 6 、 Figure 9 and / or Figure 10 the example variant generator.
[0017] Figure 15 is a flowchart showing example machine-readable instructions that can be executed to implement during an inference phase Figure 3 、 Figure 4 、 Figure 6 、 Figure 9 and / or Figure 10 of an example variant generator.
[0018] Figure 16 is a flowchart showing example machine-readable instructions that can be executed to implement Figure 3 and / or Figure 10 of an example executable file and / or Figure 3 and / or Figure 5 of an example FAT binary file.
[0019] Figure 17 is configured to execute Figure 11 、 Figure 12 、 Figure 14 and / or Figure 15 of example machine-readable instructions to implement Figure 3 、 Figure 4 、 Figure 6 and / or Figure 10 of a variant application of an example processing platform.
[0020] Figure 18 is configured to execute Figure 13 and / or Figure 16 of example machine-readable instructions to implement Figure 3 and / or Figure 10 of an example executable file of an example processing platform.
[0021] These figures are not drawn to scale. In general, the same reference numerals will be used throughout the drawings and the accompanying written description to refer to the same or like components. The connecting lines and connectors shown in the various figures presented are intended to represent example functional relationships and / or physical or logical couplings between the various elements.
[0022] When identifying multiple elements or components, the descriptors "first", "second", "third", etc. are used herein. Unless otherwise specified or understood based on their context of use, such descriptors are not intended to impart any meaning of priority or temporal order, but are merely labels for referring to multiple elements or components separately to facilitate understanding of the disclosed examples. In some examples, the descriptor "first" may be used to refer to an element in the detailed description, while a different descriptor such as "second" or "third" may be used in the claims to refer to the same element. In such cases, it should be understood that such descriptors are only for ease of reference to multiple elements or components. Detailed implementation manners
[0023] Many computer hardware manufacturers and / or other suppliers develop programs and / or other methods to compile algorithms and / or other code for running on a specific processing platform. For example, some computer hardware manufacturers develop programs and / or other methods to compile algorithms and / or other code for running on a CPU, FPGA, GPU, or VPU. Such programs and / or other methods function using domain-specific languages (DSLs). DSLs (such as, Halide, OpenCL, etc.) utilize the principle of separation of concerns to separate how an algorithm (such as, a program, a code block, etc.) is written from how the algorithm is executed. For example, many DSLs prompt developers of an algorithm to implement high-level strategies to map the processing pipeline of the algorithm to a parallel machine (such as, scheduling).
[0024] For example, an algorithm can be defined to blur an image (such as, how to write the algorithm), and a developer may desire that the algorithm run efficiently on a CPU, FPGA, GPU, and VPU. To run the algorithm efficiently on various types of processing elements (such as, CPU, FPGA, GPU, VPU, heterogeneous systems, etc.), a schedule will be generated. To generate the schedule, the algorithm is transformed in different ways depending on the specific processing element. Many methods have been developed to automate the compile-time scheduling of algorithms. For example, compile-time automation scheduling can include auto-tuning, heuristic search, and hybrid scheduling.
[0025] Auto-tuning includes compiling an algorithm in a random manner, executing the algorithm, measuring the performance of the processing element, and repeating the process until a performance threshold (such as, power consumption, execution speed, etc.) is reached. However, to reach the desired performance threshold, a large amount of compile time is required, and as the algorithm complexity increases, the compile time becomes longer.
[0026] Heuristic search includes: (1) applying rules that define the types of algorithm transformations that will improve performance to meet the performance threshold; and (2) applying rules that define the types of algorithm transformations that will not improve performance to meet the performance threshold. Then, based on these rules, a search space can be defined and searched based on a cost model. However, the cost model is usually specific to a particular processing element. Similarly, it is difficult to define a cost model for an arbitrary algorithm. For example, the cost model works under predetermined conditions, but for unknown conditions, the cost model usually fails.
[0027] Hybrid scheduling involves using artificial intelligence (AI) to identify cost models for general-purpose processing elements. The cost model can correspond to representing, predicting, and / or otherwise determining the computational cost for one or more processing elements to execute a portion of code to facilitate the processing of one or more workloads. For example, artificial intelligence, including machine learning (ML), deep learning (DL), and / or other artificial machine-driven logic, enables a machine (e.g., a computer, a logic circuit, etc.) to use a model to process input data in order to generate an output based on patterns and / or associations previously learned by the model through a training process. For example, data can be used to train a model to identify patterns and / or associations, and when processing input data, follow such patterns and / or associations such that other input(s) result in output(s) consistent with the identified patterns and / or associations.
[0028] There are many different types of machine learning models and / or machine learning architectures. Some types of machine learning models include, for example, support vector machines (SVMs), neural networks (NNs), recurrent neural networks (RNNs), convolutional neural networks (CNNs), long short-term memory (LSTMs), gated recurrent units (GRUs), etc.
[0029] Generally speaking, implementing an ML / AI system involves two phases: a learning / training phase and an inference phase. In the learning / training phase, a training algorithm is used to train a model to operate based on patterns and / or associations, for example, based on training data. Typically, a model includes internal parameters that guide how to transform input data into output data, such as transforming input data into output data through a series of nodes and connections within the model. Additionally, hyperparameters are used as part of the training process to control how learning is performed (e.g., learning rate, number of layers to use in a machine learning model, etc.). Hyperparameters are defined as training parameters that are determined before initiating the training process.
[0030] Different types of training can be performed based on the type of the ML / AI model and / or the expected output. For example, supervised training uses input and corresponding expected (e.g., labeled) outputs to select parameters for the ML / AI model (e.g., by iterating over multiple combinations of selected parameters) to reduce model error. As used herein, a label refers to the expected output of a machine learning model (e.g., classification, expected output value, etc.). Alternatively, unsupervised training (e.g., for subsets of deep learning, machine learning, etc.) involves inferring patterns from input to select parameters for the ML / AI model (e.g., without the benefit of an expected (e.g., labeled) output).
[0031] Training is performed using training data. Once training is complete, the model is deployed as an executable construct that processes input and provides output based on the network of nodes and connections defined in the model.
[0032] Once trained, the deployed model can be operated during the inference phase to process data. During the inference phase, the data to be analyzed (e.g., real-time data) is input into the model, and the model is executed to create an output. This inference phase can be thought of as the AI "thinking" in order to generate an output based on what it learned from training (e.g., by executing the model to apply the learned patterns and / or associations to the real-time data). In some examples, the input data undergoes preprocessing before being used as input to the machine learning model. Additionally, in some examples, after the output data is generated by the AI model, the output data can be post-processed to transform the output into a useful result (e.g., data display, instructions to be executed by a machine, etc.).
[0033] In some examples, the output of the deployed model can be captured and provided as feedback. By analyzing this feedback, the accuracy of the deployed model can be determined. If the feedback indicates that the accuracy of the deployed model is below a threshold or other criterion, the feedback and an updated training dataset, hyperparameters, etc. can be used to trigger the training of an updated model to generate an updated deployed model.
[0034] Regardless of the ML / AI model used, once the ML / AI model is trained, the ML / AI model generates a cost model for a general-purpose processing element. Then, the auto-regulator utilizes the cost model to generate a schedule for the algorithm. Once the schedule is generated, the schedule is utilized by an integrated development environment (IDE) associated with the DSL to generate an executable file.
[0035] The executable file includes multiple different executable segments, where each executable segment can be executed by a specific processing element, and the executable file is referred to as a FAT binary file. For example, if a developer is developing code that will be used on a heterogeneous processing platform including a CPU, FPGA, GPU, and VPU, the associated FAT binary file will include executable segments for the CPU, FPGA, GPU, and VPU respectively. In such an example, the runtime scheduler is able to rely on the physical characteristics of the heterogeneous system and a function that defines successful execution (e.g., a function that specifies successful execution of an algorithm on the heterogeneous system) to execute the algorithm on at least one of the CPU, FPGA, GPU, or VPU using the FAT binary file. For example, such a success function can correspond to executing the function so as to meet and / or otherwise satisfy a power consumption threshold. In other examples, the success function can correspond to executing the function within a threshold amount of time. However, the runtime scheduler can utilize any suitable success function when determining how to execute the algorithm on the heterogeneous system using the FAT binary file.
[0036] While auto-tuning, heuristic search, and AI-based hybrid methods may be acceptable approaches for scheduling during compilation, such scheduling methods do not account for the load and real-time performance of the individual processing elements of a heterogeneous system. For example, when developing a cost model, the developer or AI system makes assumptions about how a particular processing element (e.g., CPU, FPGA, GPU, VPU, etc.) is constructed. Additionally, the developer or AI system can make assumptions about the specific computational elements, memory subsystems, interconnect structures, and / or other components of a particular processing element. However, these components of a particular processing element are volatile, sensitive to load and environmental conditions, including subtle hardware design details, having problematic drivers / compilers, and / or including performance behaviors that run counter to expected performance.
[0037] For example, when a heterogeneous system offloads one or more computational tasks (e.g., workloads, computational workloads, etc.) to a GPU, there are specific consequences for not offloading enough computations to the GPU. More specifically, if an insufficient number of computational tasks are offloaded to the GPU, one or more hardware threads of the GPU may stall and cause one or more execution units of the GPU to shut down, thus limiting the processing power of the GPU. An example effect of this consequence can be that a workload of size X offloaded to the GPU can have the same or substantially similar processing time as a workload of size 0.5X offloaded to the GPU.
[0038] In addition, even the movement of data from one processing element to another causes complexities. For example, a runtime scheduler can utilize the texture sampler of a GPU to process images in a workload. To offload the workload to the GPU, the image is converted from a linear format supported by the CPU to a tile format supported by the GPU. Although processing the image on the GPU may be faster, such a conversion incurs a computational cost on the CPU, and the overall operation of converting the image format on the CPU and then processing it on the GPU may be longer than simply processing the image on the CPU alone.
[0039] Furthermore, many compilers utilize auto-vectorization, which relies on the knowledge of human developers about the transformations and other scheduling techniques used to trigger the auto-vectorization feature. Thus, developers who are not aware of these techniques will have less satisfactory executables.
[0040] Developing code to fully utilize all available computing resources on heterogeneous systems presents significant programming challenges, especially in achieving the best possible performance under a specific power budget. The programming complexity can evolve from multiple sources, such as (1) the hardware organization of a particular processing element may be completely different from that of different processing elements, (2) the way in which CPU programs are associated with workload-specific processing elements that require specific offloading and synchronization semantics, and (3) processing bottlenecks due to the movement of data across the entire heterogeneous system.
[0041] Typical programming languages (e.g., imperative programming languages such as C / C++) and abstractions (e.g., compute application programming interfaces (APIs)) require developers to base their programming intent on the specific details of the target hardware. However, typical imperative programming languages are often overly descriptive (e.g., they have many ways to write slow code but few ways to write fast code) and are based on serial computing constructs (e.g., loops) that are difficult to map to some types of accelerators.
[0042] A typical way to develop high-performance workloads can be through direct tuning using the most well-known methods or practices or the given target hardware. For example, developers can use detailed knowledge of the hardware organization of the target hardware to generate code that understands data / thread parallelism, uses access patterns optimized for cache behavior, and exploits hardware-specific instructions. However, such code may not be easily portable to different target hardware and thus requires additional development, different hardware expertise, and code decomposition.
[0043] The examples disclosed herein include methods and apparatuses for intentional programming for heterogeneous systems. Contrary to some methods for developing algorithms for heterogeneous systems, the examples disclosed herein do not solely rely on the developer's knowledge of the theoretical understanding of processing elements, algorithm transformation, and other scheduling techniques and other pitfalls of some methods for heterogeneous system algorithm development.
[0044] The examples disclosed herein include an example code converter for obtaining application code corresponding to one or more algorithms represented in an imperative programming language to be executed on a heterogeneous system. The example code converter is capable of identifying one or more annotated code blocks in the application code. The example code converter is capable of elevating the algorithmic intent from the one or more annotated code blocks to an elevated intermediate representation (e.g., elevated intermediate code). The example code converter is capable of lowering one or more code blocks in the elevated intermediate representation to a DSL representation.
[0045] In some disclosed examples, a code converter is capable of replacing one or more annotated code blocks with corresponding function calls to a runtime scheduler that can be used to facilitate the execution of (a) workload(s) on a heterogeneous system. In response to replacing one or more annotated code blocks, the example code converter can call a compiler to generate modified application code to be included in an executable file. In some disclosed examples, a heterogeneous system can access the modified application code by calling the executable file to process (a) workload(s) using one or more processing elements of the heterogeneous system. For example, during runtime, a heterogeneous system can execute the executable file to process multiple workloads to dynamically select one or more of one or more variant binaries of the executable file based on the performance characteristics of the heterogeneous system, thereby improving the runtime performance of software executed on the heterogeneous system.
[0046] Figure 1 FIG. depicts a block diagram illustrating an example heterogeneous system 100. In Figure 1 the illustrated example, the heterogeneous system 100 includes an example CPU 102, an example memory 104, an example FPGA 106, an example VPU 108, and an example GPU 110. Figure 1 The memory 104 of Figure 1 includes an example executable file 105. Alternatively, the memory 104 can include more than one executable file. In
[0047] In Figure 1 the illustrated example, the heterogeneous system 100 is a system-on-chip (SoC). Alternatively, the heterogeneous system 100 can be any other type of computing or hardware system.
[0048] In Figure 1In the illustrated example, the CPU 102 is a processing element that executes instructions (e.g., machine-readable instructions included in and / or otherwise corresponding to the executable file 105) to perform, implement, and / or facilitate the completion of operations associated with a computer or computing device. In Figure 1 , the CPU 102 is the primary processing element of the heterogeneous system 100 and includes at least one core. Alternatively, the CPU 102 can be a co-primary processing element (e.g., in an example where more than one CPU is used), while in other examples, the CPU 102 can be a secondary processing element.
[0049] In Figure 1 the illustrated example, the storage 104 is a memory that includes the executable file 105. Additionally or alternatively, the executable file 105 can be stored in the CPU 102, FPGA 106, VPU 108, and / or GPU 110. In Figure 1 , the storage 104 is shared storage among at least one of the CPU 102, FPGA 106, VPU 108, or GPU 110. In Figure 1 , the storage 104 is physical storage local to the heterogeneous system 100. Alternatively, the storage 104 can be external to the heterogeneous system 100 and / or otherwise remote from the heterogeneous system 100. Alternatively, the storage 104 can be virtual storage. In Figure 1 the example, the storage 104 is persistent storage (e.g., read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.). Alternatively, the storage 104 can be a persistent basic input / output system (BIOS) or flash storage. Alternatively, the storage 104 can be volatile memory.
[0050] In Figure 1In the illustrated example, one or more of FPGA 106, VPU 108, and / or GPU 110 are processing elements that can be used by a program executing on heterogeneous system 100 for computing tasks such as hardware acceleration. For example, FPGA 106 is a general-purpose programmable processing element that can be used for computable operations or processes. In other examples, VPU 108 is a processing element that includes processing resources designed and / or otherwise configured or constructed to improve the processing speed and overall performance for handling machine vision tasks for AI. In still other examples, GPU 110 is a processing element designed and / or otherwise configured or constructed to improve the processing speed and overall performance for computer graphics and / or image processing. Although FPGA 106, VPU 108, and GPU 110 include functions to support specific processing tasks, one or more of FPGA 106, VPU 108, and / or GPU 110 can correspond to processing elements that support general-purpose processing tasks that can be offloaded from CPU 102 as needed.
[0051] Although Figure 1 The heterogeneous system 100 includes CPU 102, storage 104, FPGA 106, VPU 108, and GPU 110, but in some examples, the heterogeneous system 100 can include any number and / or type of processing elements, including application-specific instruction-set processors (ASIPs), physics processing units (PPUs), digital signal processors (DSPs), image processors, coprocessors, floating-point units, network processors, multi-core processors, and front-end processors.
[0052] Figure 2 FIG. depicts a block diagram of an illustrated example system (e.g., a software tuning system) 200 that includes an example administrator device 202, a first example software tuning system 204, an example network 206, an example database 208, and a second example software tuning system 210.
[0053] In Figure 2 the illustrated example, the administrator device 202 is a desktop computer. Alternatively, the administrator device 202 can be any suitable computing system or platform, such as a mobile phone, a tablet computer, a workstation, a laptop computer, or a server. In operation, an administrator or user can train the first software tuning system 204 via the administrator device 202. For example, the administrator can generate training data via the administrator device 202. In some examples, the training data is sourced from a randomly generated algorithm that is subsequently utilized by the first software tuning system 204. For example, the administrator can use the administrator device 202 to generate a large number of algorithms (e.g., hundreds, thousands, hundreds of thousands, etc.) and transmit them to the first software tuning system 204 to train the first software tuning system 204. In Figure 2In this case, the administrator device 202 communicates with the first software tuning system 204 via a wired connection. Alternatively, the administrator device 202 can communicate with the first software tuning system 204 via any suitable wired and / or wireless connection.
[0054] In Figure 2 the illustrated example, one or both of the first software tuning system 204 and / or the second software tuning system 210 generate and improve the execution of applications on a heterogeneous system (e.g., Figure 1 the heterogeneous system 100). One or both of the first software tuning system 204 and / or the second software tuning system 210 utilize ML / AI techniques to generate applications based on the received algorithms and the performance of the processing elements.
[0055] In Figure 2 the illustrated example, the first software tuning system 204 communicates with the administrator device 202 via a wired connection. Alternatively, the first software tuning system 204 can communicate with the administrator device 202 via any suitable wired and / or wireless connection. Additionally, the first software tuning system 204 communicates with the database 208 and the second software tuning system 210 via the network 206. The first software tuning system 204 can communicate with the network 206 via any suitable wired and / or wireless connection.
[0056] In Figure 2 the illustrated example, the system 200 includes a first software tuning system 204 for training an ML / AI model (e.g., an untrained ML / AI model) 212 to generate a trained ML / AI model 214 that can be used to develop code and / or other algorithms for execution on Figure 1 the heterogeneous system 100. In some examples, the trained ML / AI model 214 can be used to facilitate and / or otherwise perform code improvement techniques such as provenance improvement, inductive synthesis, syntax-guided synthesis, etc. For example, the trained ML / AI model 214 can: obtain a query including one or more code blocks (e.g., annotated code blocks) corresponding to an algorithm; identify one or more candidate code blocks or programs having a first algorithmic intent that matches a second algorithmic intent of the one or more code blocks (e.g., substantially matches, equivalently matches, etc.); and return one or more candidate code blocks or programs as intermediate code for further processing and / or verification.
[0057] In some examples, in response to training, the first software tuning system 204 transmits and / or stores the trained ML / AI model 214. For example, the first software tuning system 204 can transmit the trained ML / AI model 214 to the database 208 via the network 206. Additionally or alternatively, the first software tuning system 204 can transmit the trained ML / AI model 214 to the second software tuning system 210.
[0058] In Figure 2 the illustrated example of, the system 200 includes a second software tuning system 210 for executing code and / or other algorithms on a heterogeneous system using the trained ML / AI model 214. The second software tuning system 210 can obtain the trained ML / AI model 214 from the first software tuning system 204 or the database 208. Alternatively, the second software tuning system 210 can generate the trained ML / AI model 214.
[0059] In some examples, the second software tuning system 210 collects and / or otherwise obtains data associated with at least one of a heterogeneous system or a system-wide success function of a heterogeneous system. In response to collecting the data, the second software tuning system 210 can transmit the data to the first software tuning system 204 and / or the database 208. The second software tuning system 210 can format the data in various ways as described in connection with Figure 3 described.
[0060] In Figure 2 the illustrated example of, the network 206 is the Internet. However, any suitable wired and / or wireless network(s) can be used to implement the network 206, including for example one or more data buses, one or more local area networks (LANs), one or more wireless LANs (WLANs), one or more cellular networks, one or more private networks, one or more public networks, etc. The network 206 enables the first software tuning system 204, the database 208, and / or the second software tuning system 210 to communicate with each other.
[0061] In Figure 2In the illustrated example, system 200 includes a database 208 for recording and / or otherwise storing data (e.g., heterogeneous system performance data, system-wide success functions, trained ML / AI models 214, etc.). The database 208 can be implemented by volatile memory (e.g., synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS dynamic random access memory (RDRAM), etc.) and / or non-volatile memory (e.g., flash memory). The database 208 can additionally or alternatively be implemented by one or more double data rate (DDR) memories (such as DDR, DDR2, DDR3, DDR4, mobile DDR (mDDR), etc.). The database 208 can additionally or alternatively be implemented by one or more mass storage devices, such as (multiple) HDDs, (multiple) CD drives, (multiple) digital versatile disc (DVD) drives, (multiple) SSDs, etc. Although the database 208 is shown as a single database in the illustrated example, the database 208 can be implemented by any number and / or (multiple) types of databases. Additionally, the data stored in the database 208 can be in any data format, such as, for example, binary data, comma-separated data, tab-separated data, structured query language (SQL) structures, etc. In Figure 2 this example, the database 208 is stored on an electronically accessible computing system. For example, the database 208 can be stored on a server, a desktop computer, an HDD, an SSD, or any other suitable computing system.
[0062] Figure 3 FIG. depicts a block diagram of an illustrated example system (e.g., a variant generation system) 300 that can be used to implement Figure 2 the first software tuning system 204 and / or the second software tuning system 210 of Figure 3 The variant generation system 300 of
[0063] In Figure 3 the illustrated example, the variant generation system 300 includes a third example software tuning system 301, an example variant generator 302, an example code converter 303, an example heterogeneous system 304, and an example storage 306. In Figure 3 this example, the example variant application 305 includes the variant generator 302 and the code converter 303. Alternatively, the variant application 305 can include the variant generator 302 or the code converter 303.
[0064] Further depicted in the variant generation system 300 is Figure 2 the network 206 of Figure 2database 208. Alternatively, the variant generation system 300 may not include the network 206 and / or the database 208. Alternatively, the storage 306 may be external to the heterogeneous system 304. Figure 3 The storage 306 includes an example executable file 308. The executable file 308 includes an example FAT binary (e.g., a FAT binary file) 309, which includes an example variant library 310, an example jump table library 312, and an example runtime scheduler 314. Alternatively, the storage 306 may include more than one executable file.
[0065] In Figure 3 the illustrated example, the third software tuning system 301 can correspond to Figure 2 the first software tuning system 204 or Figure 2 the second software tuning system 210. For example, the third software tuning system 301 can include Figure 2 the untrained ML / AI model 212 and / or the trained ML / AI model 214. In some examples, the variant generator 302 includes the untrained ML / AI model 212 and / or the trained ML / AI model 214. In some examples, the code converter 303 includes the untrained ML / AI model 212 and / or the trained ML / AI model 214. In some examples, each of the variant generator 302 and the code converter 303 includes at least one of the untrained ML / AI model 212 or the trained ML / AI model 214.
[0066] In Figure 3 the illustrated example, the heterogeneous system 304 can correspond to Figure 1 the heterogeneous system 100. In Figure 3 it, the storage 306 can correspond to Figure 1 the storage 104. In Figure 3 it, the executable file 308 can correspond to Figure 1 the executable file 105. In Figure 3 it, the heterogeneous system 304 includes an example CPU 316, an example FPGA 318, an example VPU 320, and an example GPU 322. Alternatively, the heterogeneous system 304 may include fewer or more processing elements than Figure 3 depicted therein. Alternatively, the heterogeneous system 304 may include multiple ones of the CPU 316, FPGA 318, VPU 320, and / or GPU 322. In Figure 3 it, the CPU 316 can correspond to Figure 1 the CPU 102. In Figure 3 it, the FPGA 318 can correspond to Figure 1 the FPGA 106. In Figure 3In it, the VPU 320 can correspond to the Figure 1 VPU 108. In Figure 3 it, the GPU 322 can correspond to the Figure 1 GPU 110.
[0067] In Figure 3 the illustrated example, the third software tuning system 301 includes a variant application 305 for facilitating the execution of at least one of the variant generator 302 or the code converter 303. For example, the variant application 305 can be a computing application, an executable file, etc. In such an example, the variant application 305 can be used by a user (e.g., an administrator, a developer, etc.) to generate an algorithm to be deployed to the heterogeneous system 304 and / or otherwise executed on the heterogeneous system 304. In some examples, the variant application 305 can correspond to and / or otherwise be implemented by a computer cluster (e.g., a cloud computing environment, a server room, etc.). Alternatively, the variant application 305 can be included in the heterogeneous system 304 and / or otherwise be implemented by the heterogeneous system 304.
[0068] In Figure 3 the illustrated example, the variant application 305 includes a code converter 303 for transforming and / or otherwise converting code (e.g., a code block, a code section, etc.) in a format or representation of an imperative programming language into a DSL representation or format. The code converter 303 can be a device or a software executable file executed by the device. Advantageously, the user can use their existing programming knowledge to develop code for performing heterogeneous offloading (e.g., instructing one or more processing elements of the heterogeneous system 304 to perform a workload) with a minimal understanding of the specific processing elements of the heterogeneous system 304. By converting the code into a DSL representation, the variant generator 302 can generate multiple variants of an algorithm to be executed by one or more processing elements, as described in more detail below. The code converter 303 enables the user to develop code (e.g., basic code or simple code) to describe a minimal algorithm, and the code converter 303 can handle the characteristics of the hardware organization of the processing elements, heterogeneous offloading, and data movement between the processing elements.
[0069] In some examples, the code converter 303 uses imperative programming language compiler extensions (e.g., #pragma or %pragma, #pragma lift begin or %pragma lift begin, #pragma lift end or %pragma lift end, etc.) to annotate code blocks applicable to heterogeneous offloading. Alternatively, any other type of compiler-specific language extension may be used. The annotated code blocks correspond to and / or otherwise indicate the intention of the computation (e.g., algorithmic intention). For example, a user can describe the intention through a premature implementation of the desired algorithm, and in some examples, corresponding metadata regarding runtime scheduling and variant generation is provided. In such examples, the user can develop a premature implementation describing the algorithmic intention by using typical imperative programming language idioms (e.g., loops, variables, etc.).
[0070] In some examples, the code converter 303 lifts one or more annotated code blocks from an imperative programming language representation to a formal representation. For example, the code converter 303 can use proven lifting techniques such as inductive synthesis (e.g., inductive program synthesis) to identify and / or otherwise establish the algorithmic intention and remove details such as loop order and workload-specific optimizations from one or more annotated code blocks.
[0071] In some examples, the code converter 303 uses inductive synthesis techniques to reduce the lifted algorithm in formal representation to a DSL representation with separation of concerns (e.g., Halide, OpenCL, etc.). Thus, the code converter 303 can transmit the DSL representation of the algorithmic intention to the variant generator 302 to compile and / or otherwise generate one or more variants of the algorithmic intention.
[0072] In Figure 3 the illustrated example, the variant application 305 includes a variant generator 302 for compiling one or more variants of an algorithm to be executed by the heterogeneous system 304. For example, the variant generator 302 can obtain the algorithm in DSL representation from the code converter 303 and is capable of compiling one or more variants of the algorithm to offload them to one or more processing elements of the heterogeneous system 304. Alternatively, the variant generator 302 can be separate from the third software tuning system 301 and / or the variant application 305.
[0073] In Figure 3In the illustrated example, the variant generator 302 is depicted separately from the heterogeneous system 304. For example, the variant generator 302 can be located at a remote facility (e.g., remote with respect to the heterogeneous system 304). In such an example, the variant generator 302 can correspond to and / or be implemented by a computer cluster (e.g., a cloud computing environment, a server room, etc.). Alternatively, the variant generator 302 can be included in and / or be implemented by the heterogeneous system 304.
[0074] In Figure 3 the illustrated example, the variant generator 302 and / or, more generally, the variant application 305 is coupled to the network 206. In Figure 3 it, the variant generator 302 and / or, more generally, the variant application 305 is coupled to the storage 306, the variant library 310, the jump table library 312, and the runtime scheduler 314. The variant generator 302 can receive algorithms and / or machine learning models from one or more external devices (such as Figure 2 the administrator device 202, the first software tuning system 204, or the second software tuning system 210) of
[0075] In some examples, during the training phase, the variant generator 302 can receive and / or otherwise obtain an algorithm (e.g., a random algorithm or a randomly selected algorithm) from one or more external devices to train Figure 2 the untrained ML / AI model 212 of Figure 2 it. In other examples, during the inference phase, the variant generator 302 can receive and / or otherwise obtain user-generated algorithms and / or trained ML / AI models (e.g., Figure 2 the trained ML / AI model 214 of
[0076] In Figure 3In the illustrated example, variant generator 302 is a device that compiles algorithms received from a code converter 303, one or more external devices, etc. into an executable application that includes multiple variants of the algorithms or a software executable file executed by such a device. Additionally or alternatively, variant generator 302 can generate a trained ML / AI model (e.g., trained ML / AI model 214) associated with generating an application to be run on a heterogeneous system 304. For example, if the algorithm received from one or more external devices is written in a programming language such as C or C++ (e.g., an imperative programming language), variant generator 302 can compile the algorithm into an executable application for storage in Figure 3 storage 306. In some examples, the executable application compiled by variant generator 302 is a FAT binary file. Alternatively, the executable application compiled by variant generator 302 can be any other suitable binary or executable file.
[0077] In Figure 3 the illustrated example, variant generator 302 utilizes ML / AI techniques. In some examples, variant generator 302 utilizes a convolutional neural network (CNN) model, a deep neural network (DNN) model, etc. Generally, the machine learning models / architectures applicable to the examples disclosed herein will be supervised. However, other examples can include machine learning models / architectures that utilize unsupervised learning. In some examples, gradient descent is used to train the ML / AI models disclosed herein. In some examples, hyperparameters used to train the ML / AI models disclosed herein control the exponential decay rate of the moving average of the gradient descent. Such hyperparameters are selected, for example, by iterating through a grid of hyperparameters until the hyperparameters meet and / or otherwise satisfy an acceptable or predefined performance value. Additionally or alternatively, any other ML / AI training algorithm can be used.
[0078] In Figure 3In the illustrated example, during the training phase, the variant generator 302 executes, operates, and / or otherwise generates a trained ML / AI model 214 that is capable of generating an executable application that includes multiple variants of one or more algorithms that are capable of being executed on various processing elements. When in the training phase, the variant generator 302 selects a processing element (e.g., CPU 316, FPGA 318, VPU 320, or GPU 322) for which the variant generator 302 is to develop one or more variants and corresponding executable applications. In response to selecting the processing element of interest, e.g., FPGA 318, when in the training phase, the variant generator 302 selects an aspect of the processing element to optimize. For example, the variant generator 302 may select the execution speed of an algorithm on the FPGA 318 for optimization.
[0079] In Figure 3 the illustrated example, in response to selecting the aspect of the processing element to optimize, the variant generator 302 uses a machine learning model (e.g., CNN, DNN, trained ML / AI model 214, etc.) to generate a cost model for the processing element. The variant generator 302 uses an auto-tuning technique to develop a schedule to map the algorithm to the selected processing element to improve the selected aspect. For example, the variant generator 302 can use an auto-tuning technique to develop a schedule to map the algorithm to the FPGA 318 such that the mapping of the algorithm to the FPGA 318 improves the execution speed of the algorithm on the FPGA 318.
[0080] In Figure 3 the illustrated example, in response to developing a schedule for the selected processing element, the variant generator 302 compiles the algorithm (e.g., the algorithm in DSL representation from the code converter 303) into a variant (e.g., variant binary, variant binary file, etc.) according to the schedule. The compilation of the algorithm is different from the compilation of the executable application because the variant generator 302 compiles the algorithm into methods, classes, and / or objects that can be invoked or called out by the executable application (e.g., executable file 308). In response to compiling the variant, when in the training phase, the variant generator 302 transfers the variant to the executable file 308 in the storage 306. For example, the executable file 308 may include a FAT binary file 309 stored in the storage 306, and the variant generator 302 is capable of storing the variant in the variant library 310 of the executable file 308. In some examples, when in the training phase, the variant generator 302 transfers variant symbols to the executable file 308 in the storage 306. Variant symbols are data elements corresponding to the location of the variant in the variant library 310.
[0081] In Figure 3In the illustrated example, the variant is then executed on the heterogeneous system 304. In response to executing the variant on the heterogeneous system 304, the variant generator 302 can obtain performance characteristics associated with the selected processing element (e.g., FPGA 318). When in training mode, the performance characteristics can correspond to the characteristics of the selected processing element (e.g., FPGA 318), including, for example, the power consumption of the selected processing element, the time to run on the selected processing element, and / or other performance characteristics associated with the selected processing element.
[0082] In Figure 3 the illustrated example, the variant generator 302 analyzes the collected data and determines whether the variant used meets a performance threshold. In some examples, training is performed until the performance threshold is met. For example, the performance threshold can correspond to an acceptable amount of L2 (least squares regression) error achieved for the selected aspect. In response to meeting the performance threshold, the variant generator 302 can determine whether there are subsequent aspects to optimize. In response to determining that there is at least one subsequent aspect to be optimized, the variant generator 302 can generate additional variants corresponding to the subsequent aspect (e.g., the power consumption of FPGA 318) for the selected processing element. In response to determining that there is no other aspect to be optimized, the variant generator 302 can determine whether there is at least one processing element of interest to be processed in order to generate one or more corresponding variants (e.g., a first variant generated for CPU 316, a second variant generated for VPU 320, and / or a third variant generated for GPU 322 that is the opposite of the fourth variant generated (e.g., previously generated) for FPGA 318).
[0083] In Figure 3 the illustrated example, in response to the variant generator 302 generating variants for all processing elements of the heterogeneous system 304, the variant generator 302 determines whether there is at least one additional algorithm for which variants are to be generated. In response to determining that there is another algorithm to be processed, the variant generator 302 can generate variants of the additional algorithm for each processing element of the heterogeneous system 304 for any selection and / or any aspect of each processing element. In response to determining that there are no additional algorithms of interest to be processed, the variant generator 302 outputs the trained ML / AI model 214. For example, the variant generator 302 can output one or more files, including weights associated with the cost model of each processing element of the heterogeneous system 304. The model can be stored in the database 208, the storage 306, and / or in different variant generators depicted with Figure 3 The trained ML / AI model 214 can be executed by the variant generator 302 during subsequent execution, or by a different variant generator depicted with Figure 3 depicted.
[0084] In Figure 3In the illustrated example, in response to outputting and / or otherwise generating the trained ML / AI model 214, the variant generator 302 monitors any additional input data. For example, the input data may correspond to data associated with the execution of an application generated by the trained ML / AI model 214 on a target platform (e.g., the heterogeneous system 304). When executing the desired workload, the specific data obtained by the variant generator 302 can indicate the performance of the target platform and can reflect the actual system under the actual load as opposed to the test system under the simulated load. In response to receiving and / or otherwise obtaining the input data, the variant generator 302 can identify the success function of the heterogeneous system 304. Based on the success function, the variant generator 302 can determine a performance delta corresponding to the difference between the desired performance (e.g., a performance threshold) defined by the success function and the actual performance obtained during the execution of the algorithm during the inference phase.
[0085] In Figure 3 the illustrated example, in response to the variant generator 302 determining at least one of (1) the success function, (2) the relevant aspect(s) of the overall system (e.g., the heterogeneous system 304) being targeted, or (3) the performance delta associated with the success function, the variant generator 302 updates and / or otherwise adjusts the (multiple) cost models associated with the respective processing elements of the heterogeneous system 304 to account for the real-time characteristics and load (e.g., load distribution) of the heterogeneous system 304. This is described below in connection with Figure 4 updating and other adjustments to the (multiple) cost models associated with the respective processing elements of the heterogeneous system.
[0086] In Figure 3 the illustrated example, the variant library 310 is a data structure associated with the executable file 308 that stores different variants of the algorithms that the executable file 308 can execute. For example, the variant library 310 can correspond to the following data section of the FAT binary file 309 that includes different variants associated with a particular algorithm, such as variants associated with the respective processing elements of the heterogeneous system 304. For one or more processing elements, the variant library 310 may additionally include variants for different aspects of the performance of the respective one or more processing elements. In some examples, the variant library 310 is linked to at least one of the jump table library 312 or the runtime scheduler 314. In some examples, the variant library 310 is a static library during the execution of the executable file 308, but can be updated with new or changed variants between multiple executions of the executable file 308.
[0087] In Figure 3In the illustrated example, the jump table library 312 is a data structure associated with the executable file 308 that stores one or more jump tables that include variant symbols that point to the locations of corresponding variants in the variant library 310. For example, the jump table library 312 can correspond to a data section of the executable file 308 that includes a jump table that includes associations of variant symbols (e.g., pointers) with corresponding variants located in the variant library 310. In some examples, the jump table library 312 does not change during the execution of the executable file 308. In such examples, the jump table library 312 can be accessed to call, direct, and / or otherwise evoke the corresponding variants to be loaded onto one or more processing elements of the heterogeneous system 304.
[0088] In Figure 3 the illustrated example, the runtime scheduler 314 determines how to execute a workload (e.g., one or more algorithms) during runtime of the heterogeneous system 304. In some examples, the runtime scheduler 314 generates and / or transmits an execution graph to the processing elements for execution and / or otherwise implements it. In some examples, the runtime scheduler 314 determines whether a workload should be offloaded from one processing element to a different processing element to achieve performance goals associated with the overall heterogeneous system 304.
[0089] In some examples, the code converter 303 replaces an annotated code block with a function call to the runtime scheduler 314. In response to lifting an annotated code block associated with a user-generated algorithm from user-generated application code, the code converter 303 can insert a function call to the runtime scheduler 314 at the location of the lifted annotated code block in the user-generated application code. For example, during the execution of the executable file 308, the heterogeneous system 304 can use the inserted function call to call the runtime scheduler 314 in response to accessing a corresponding code site in the user-generated application code of the executable file 308.
[0090] In some examples, to facilitate the execution of the function call to the runtime scheduler 314, the code converter 303 includes a memory allocation routine for the user-generated application code to ensure compatibility across processing elements. For example, the memory allocation routine(s) can copy data used in heterogeneous offloading to a shared memory buffer visible to all processing elements of the heterogeneous system 304. In other examples, the code converter 303 can simplify the data aspect of heterogeneous offloading by restricting the lifting to specific imperative programming language functions with well-defined data capture semantics (e.g., C++ lambda functions).
[0091] In Figure 3In the illustrated example, during the execution of executable file 308, runtime scheduler 314 is capable of monitoring heterogeneous system 304 to analyze the performance of heterogeneous system 304 based on performance characteristics obtained from executable file 308. In some examples, runtime scheduler 314 is capable of determining to offload a workload from one processing element to another based on performance. For example, during the runtime of heterogeneous system 304, executable file 308 can be executed by CPU 316, and based on the performance of CPU 316 and / or more generally, heterogeneous system 304, runtime scheduler 314 can offload the workload scheduled to be executed by CPU 316 to FPGA 318. In some examples, CPU 316 executes executable file 308 from storage 306, while in other examples, CPU 316 can execute executable file 308 locally on CPU 316.
[0092] In Figure 3 the illustrated example, in response to CPU 316 executing executable file 308, runtime scheduler 314 determines a success function. For example, during a training phase, the success function can be associated with a processing element of interest (e.g., GPU 322), where an untrained ML / AI model 212 is being trained for the processing element of interest. Figure 2 Contrary to operating in a training phase where runtime scheduler 314 determines a success function for a processing element of interest, when operating in an inference phase, runtime scheduler 314 can determine a system-wide success function. For example, a first system-wide success function can be associated with executing an algorithm by consuming less than or equal to a threshold power amount using executable file 308. In other examples, a second system-wide success function can be associated with executing an algorithm as fast as possible using executable file 308 without considering power consumption.
[0093] In some examples, the system-wide success function(s) is based on the overall state of heterogeneous system 304. For example, if heterogeneous system 304 is included in a laptop computer in a low power mode (e.g., the battery of the laptop computer is below a threshold battery level, the laptop computer is not connected to a battery charging source, etc.), then the system-wide success function(s) can be associated with power savings. In other examples, if heterogeneous system 304 is included in a laptop computer operating in a normal power mode (e.g., the battery is fully charged or substantially charged) or under normal operating conditions of the laptop computer, then the system-wide success function can be associated with the execution speed of the algorithm, since power savings may not be an issue.
[0094] In Figure 3In the illustrated example, the success function can be specific to the processing elements of the heterogeneous system 304. For example, the success function can be associated with leveraging a GPU 322 that exceeds a threshold amount, preventing contention between CPU 316 threads, or leveraging high-speed memory of a VPU 320 that exceeds a threshold amount. In some examples, the success function can be a combination of simpler success functions, such as the overall performance of the heterogeneous system 304 per unit of power.
[0095] In Figure 3 the illustrated example, in response to identifying the success function, the runtime scheduler 314 executes the executable file 308 based on the variant(s) generated by the ML / AI model. For example, during the training phase, the untrained ML / AI model 212 that generated the variant is not trained, and the runtime scheduler 314 is related to the specific performance of the processing element in which the untrained ML / AI model 212 is being trained. However, during the inference phase, the trained ML / AI model 214 that generated the variant is trained, and the runtime scheduler 314 is related to the specific performance of the heterogeneous system 304 with respect to an overall or substantial portion of the heterogeneous system 304. For example, during the inference phase, the runtime scheduler 314 can collect specific performance characteristics associated with the heterogeneous system 304 and can store and / or transmit the performance characteristics in the database 208, storage 306, etc.
[0096] In Figure 3 the illustrated example, during the inference phase, the runtime scheduler 314 collects performance characteristics, including metadata and metric information associated with each variant included in the executable file 308. For example, such metadata and metric information can include an identifier of the workload (e.g., the name of the algorithm), compatibility constraints associated with the drivers and other hardware of the heterogeneous system 304, the version of the cost model used to generate the variant, the algorithm execution size, and other data that ensures compatibility between the execution of the workload and one or more processing elements. In such an example, the runtime scheduler 314 can determine an offloading decision based on the metadata and metric information. The performance characteristics collected by the runtime scheduler 314 during the inference phase can also include the average execution time of the variant on each processing element, the average occupancy rate of each processing element during runtime, the (multiple) stall rate, the power consumption of each processing element, the compute cycle count utilized by the processing element, the memory latency when offloading the workload, the risk of offloading the workload from one processing element to another, the system-wide battery life, the amount of memory utilized, metrics associated with the communication buses between various processing elements, metrics associated with the storage 306 of the heterogeneous system 304, etc. and / or combinations thereof.
[0097] In Figure 3In the illustrated example, during the inference phase, the runtime scheduler 314 collects data associated with state transition data corresponding to the load and environmental conditions of the heterogeneous system 304 (e.g., why the runtime scheduler 314 accesses the jump table library 312, where / why the runtime scheduler 314 offloads the workload, etc.). In some examples, the state transition data includes runtime scheduling rules associated with the thermal and power characteristics of the heterogeneous system 304, and runtime scheduling rules associated with any other conditions that may interfere with (e.g., affect) the performance of the heterogeneous system 304.
[0098] In Figure 3 the illustrated example, in response to monitoring and / or collecting performance characteristics, the runtime scheduler 314 adjusts the configuration of the heterogeneous system 304 based on the success function of the heterogeneous system 304. For example, periodically, during the entire operation of the runtime scheduler 314, during the inference phase, the runtime scheduler 314 can store and / or transmit performance characteristics for further use by the variant generator 302. In such an example, the runtime scheduler 314 is capable of identifying whether the heterogeneous system 304 includes persistent storage (e.g., ROM, PROM, EPROM, etc.), a persistent BIOS, or flash storage.
[0099] In Figure 3 the illustrated example, if the heterogeneous system 304 includes persistent storage, the runtime scheduler 314 writes to a data section (e.g., the FAT binary file 309) in the executable file 308 to store the performance characteristics. The performance characteristics can be stored in the executable file 308 to avoid the possibility of historical loss between different executions of the executable file 308. In some examples, the runtime scheduler 314 executes as an image of the executable file 308 on the CPU 316. In such an example, the runtime scheduler 314 is capable of storing the performance characteristics in the executable file 308 stored in the storage 306. If the heterogeneous system 304 does not include persistent storage, but instead includes flash storage or a persistent BIOS, a similar method of storing the performance characteristics in the executable file 308 can be implemented.
[0100] In Figure 3 the illustrated example, if there is no form or instance of available persistent storage, persistent BIOS, or flash storage (e.g., if the storage 306 is volatile memory), the runtime scheduler 314 can alternatively transmit the collected performance characteristics to an external device (e.g., the first software tuning system 204, the second software tuning system 206, the database 208, the variant generator 302, etc. and / or combinations thereof) using a port of the communication interface. For example, the runtime scheduler 314 can use a universal serial bus (USB), Ethernet, serial, or any other suitable communication interface to transmit the collected performance characteristics to the external device.
[0101] In Figure 3 the illustrated example of, in response to the heterogeneous system 304 executing the executable file 308, the runtime scheduler 314 transmits performance characteristics and a performance delta associated with a system-wide success function to an external device. The performance delta can indicate, for example, the difference between the expected performance and the achieved performance.
[0102] In Figure 3 the illustrated example of, on a subsequent execution of the executable file 308, the runtime scheduler 314 is able to access the stored performance characteristics and adjust and / or otherwise improve the ML / AI model (e.g., the trained ML / AI model 214) in order to improve the processing and / or convenience of offloading the workload to a processing element. For example, the runtime scheduler 314 can access the stored performance characteristics to adjust the trained ML / AI model 214. In such an example, the stored performance characteristics can include data corresponding to bus traffic under load, preemption actions taken by the operating system of the heterogeneous system 304, decoding latency associated with audio and / or video processing, and / or any other data that can be used as a basis for making offloading decisions. For example, if the runtime scheduler 314 encounters an algorithm that includes decoding a video, the runtime scheduler 314 can initially schedule the video decoding task on the GPU 322. In such instances, the runtime scheduler 314 can have a variant available for a different processing element (e.g., the VPU 320) that can individually process the video decoding task faster than the variant executed on the GPU 322. However, the runtime scheduler 314 can decide not to offload the video decoding task to a different processing element because the memory movement latency associated with moving the workload from the GPU 322 to a different processing element can take the same or increased amount of time compared to keeping the workload on the GPU 322.
[0103] Although an example manner of implementing the executable file 308 is shown in Figure 3 , one or more of the elements, processes, and / or devices shown in Figure 3 can be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other way. Additionally, Figure 3 the example variant library 310, the example jump table library 312, the example runtime scheduler 314, and / or more generally, the example executable file 308 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, Figure 3Any of the example variant library 310, example jump table library 312, example runtime scheduler 314, and / or more generally, example executable 308 can be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, GPUs, DSPs, application specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field programmable logic devices (FPLDs). When reading any of the apparatus or system claims of this patent that cover pure software and / or firmware implementations, Figure 3 At least one of the example variant library 310, example jump table library 312, example runtime scheduler 314, and / or more generally, example executable 308 is thus explicitly defined to include a non-transitory computer-readable storage device or storage disk, such as a memory, DVD, CD, Blu-ray disc, etc., including software and / or firmware. Additionally, example executable 308 can include one or more elements, processes, and / or devices in addition to or in place of Figure 3 those elements, processes, and / or devices shown in, and / or can include more than one of any or all of the elements, processes, and devices shown. As used herein, the phrase "communicate" includes its various variants, including direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but additionally includes selective communication at periodic intervals, predetermined intervals, aperiodic intervals, and / or one-time events.
[0104] Figure 4 depicts a block diagram illustrating an example implementation of Figure 3 variant generator 302. Variant generator 302 compiles one or more variant binaries corresponding to the algorithms in DSL representation from code converter 303. In Figure 4 the illustrated example of, variant generator 302 includes example variant manager 402, example cost model learner 404, example weight store 406, example compilation auto-scheduler 408, example variant compiler 410, example jump table 412, example application compiler 414, example feedback interface 416, and example profiler 418.
[0105] In some examples, one or more of the variant manager 402, cost model learner 404, weight store 406, compilation auto-scheduler 408, variant compiler 410, jump table 412, application compiler 414, feedback interface 416, and / or performance profiler 418 communicate with one or more of the other elements of the variant generator 302. For example, the variant manager 402, cost model learner 404, weight store 406, compilation auto-scheduler 408, variant compiler 410, jump table 412, application compiler 414, feedback interface 416, and performance profiler 418 can communicate with each other via an example communication bus 420. For example, the communication bus 420 can correspond to interface circuitry or any other type of hardware or logic circuitry for facilitating communication between components. In other examples, the communication bus 420 can be implemented in software.
[0106] In some examples, the variant manager 402, cost model learner 404, weight store 406, compilation auto-scheduler 408, variant compiler 410, jump table 412, application compiler 414, feedback interface 416, and performance profiler 418 can communicate via any suitable wired and / or wireless communication method. In some examples, each of the variant manager 402, cost model learner 404, weight store 406, compilation auto-scheduler 408, variant compiler 410, jump table 412, application compiler 414, feedback interface 416, and performance profiler 418 can communicate with any component external to the variant generator 302 (e.g., Figure 2 one or more external devices) via any suitable wired and / or wireless communication method.
[0107] In Figure 4 the illustrated example, the variant generator 302 includes a variant manager 402 for analyzing communications received and / or otherwise obtained from devices external to the variant generator 302 (e.g., heterogeneous system 304, Figure 2 database 208 of Figure 2 the administrator device 202, etc. and / or combinations thereof) and managing the operation of one or more components of the variant generator 302. In some examples, the variant manager 402 receives and / or otherwise obtains algorithms from the code converter 303. For example, the variant manager 402 can obtain an algorithm represented in DSL from the code converter 303, which is converted from a representation in an imperative programming language by the code converter 303.
[0108] In some examples, the variant manager 402 receives and / or otherwise obtains algorithms from external devices. For example, during a training phase, the variant manager 402 can obtain any algorithm in a series or set of algorithms for training the variant manager 402. Additionally or alternatively, during an inference phase, the variant manager 402 can obtain an algorithm associated with a workload to be executed on the Figure 3 heterogeneous system 304.
[0109] In Figure 4 the illustrated example, in response to obtaining an algorithm from an external device, the variant manager 402 is able to select a processing element for which to generate a cost model and / or variants. For example, the processing element can be Figure 3 one of the CPUs 316, FPGAs 318, VPUs 320, or GPUs 322 of Figure 3 the runtime scheduler 314).
[0110] In some examples, in response to generating a variant and determining that the variant meets a performance threshold associated with a success function, the variant manager 402 determines whether there are any additional aspects of the selected processing element to target, whether there are additional processing elements for which to generate variants, and / or whether there are any additional algorithms for training the cost model learner 404. In response to determining that there is one or more additional aspects, one or more additional processing elements, and / or one or more additional algorithms of interest to process, the variant manager 402 can repeat the above actions. In response to determining that there are no additional aspects, additional processing elements, and / or additional algorithms, the variant manager 402 can output weights associated with the respective trained ML / AI models (e.g., Figure 2 the trained ML / AI model 214) corresponding to the respective processing elements of the heterogeneous system 304.
[0111] In some examples, the variant generator 302 includes the variant manager 402 for obtaining the configuration of a hardware target of interest. For example, the variant manager 402 can obtain a configuration associated with Figure 3a configuration (e.g., a target configuration) associated with the heterogeneous system 304. In some examples, the configuration can correspond to a processing element API, a host architecture, code generation options, scheduling heuristics, etc. for a processing element of interest and / or combinations thereof. In some examples, the variant manager 402 can obtain the configuration from the database 208, the heterogeneous system 304, etc. and / or combinations thereof. The configuration can include information indicating and / or otherwise corresponding to the heterogeneous system 304, which includes Figure 3 the CPU 316, FPGA 318, VPU 320, and GPU 322.
[0112] In Figure 4 the illustrated example, the variant generator 302 includes a cost model learner 404 for implementing and / or otherwise facilitating the execution of one or more ML / AI techniques to generate a trained ML / AI model associated with generating an application to run on a heterogeneous system. In some examples, the cost model learner 404 implements a supervised DNN to learn and improve a cost model associated with a processing element. However, in other examples, the cost model learner 404 can utilize supervised and / or unsupervised learning to implement any suitable ML / AI model. In some examples, the cost model learner 404 implements a CNN, DNN, etc. for each processing element of the heterogeneous system 304.
[0113] In Figure 4 the illustrated example, the variant generator 302 includes a weight store 406 as a memory, where the weights can be associated with one or more cost models for respective processing elements of the plurality of processing elements of the heterogeneous system 304. In some examples, the weights are stored in a file structure, where each cost model has a corresponding weight file. Alternatively, a weight file can be used for more than one cost model. In some examples, during the compilation of an auto-scheduling event (e.g., an event performed by the compile auto-scheduler 408) and in response to the variant manager 402 outputting and / or otherwise generating the trained ML / AI model 214, the weight file is read. In some examples, in response to the cost model learner 404 generating a cost model, the weights are written to the weight file.
[0114] In Figure 4In the illustrated example, the variant generator 302 includes a compilation auto-scheduler 408 for generating a schedule associated with an algorithm for a selected processing element based on a cost model (e.g., a weight file) generated by the cost model learner 404. In some examples, the schedule can correspond to a method or order of operations associated with computing and / or otherwise processing the algorithm, where the schedule can include at least one of a selection or decision regarding memory locality, redundant computation, or parallelism. In some examples, the compilation auto-scheduler 408 generates the schedule by using auto-tuning. Alternatively, any suitable auto-scheduling method can be used to generate a schedule associated with an algorithm for a selected processing element.
[0115] In some examples, the compilation auto-scheduler 408 selects and / or otherwise identifies the processing element (e.g., the hardware target) to be processed. For example, the compilation auto-scheduler 408 can select Figure 3 the CPU 316 to process. The compilation auto-scheduler 408 can configure one or more auto-schedulers included in the compilation auto-scheduler 408 based on the corresponding configuration for the processing element. For example, the compilation auto-scheduler 408 can be configured based on the target configuration associated with the CPU 316. In such examples, the compilation auto-scheduler 408 can be configured based on the hardware architecture, scheduling heuristics, etc., associated with the CPU 316.
[0116] In some examples, the compilation auto-scheduler 408 can be configured based on metadata from the code converter 303. For example, the code converter 303 can embed and / or otherwise include the metadata within the annotated code block, or transmit the metadata to the compilation auto-scheduler 408. In some examples, the code converter 303 generates metadata that includes scheduling information or data corresponding to the power profile of the heterogeneous system 304, where the power profile indicates whether low power consumption is preferred over maximum performance and to what extent. In some examples, the code converter 303 generates metadata that includes instructions specifying which of the (one or more) processing elements and / or (one or more) compute APIs are to be used for the algorithm of interest. In some examples, the code converter 303 generates metadata that includes instructions specifying a portion of a particular cost model, multiple variants of one or more processing elements to be used, lossy or reduced bit-width variants, etc., and / or combinations thereof.
[0117] In Figure 4In the illustrated example, the variant generator 302 includes a variant compiler 410 for compiling the schedules generated by the compilation auto-scheduler 408. In some examples, the variant compiler 410 compiles the algorithms into methods, classes, or objects that can be invoked or called by an executable application. In response to compiling a variant, the variant compiler 410 can transmit the variant to the application to be compiled. In some examples, the variant compiler 410 transmits the variant to a jump table 412.
[0118] In some examples, the variant compiler 410 implements means for compiling variant binaries 502, 504, 506, 508, 510 based on the schedules, where each of the variant binaries 502, 504, 506, 508, 510 is associated with an algorithm of interest using a DSL, and where the variant binaries 502, 504, 506, 508, 510 include a first variant binary 502 corresponding to a first processing element (e.g., Figure 3 CPU 316) and a second variant binary 504 corresponding to a second processing element (e.g., Figure 3 GPU 322). For example, the means for compiling can be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, GPUs, DSPs, application specific integrated circuits (ASICs), PLDs, and / or FPLDs.
[0119] In Figure 4 the illustrated example, the variant generator 302 includes a jump table 412 for associating different variants generated by the variant compiler 410 with the locations where the corresponding variants are (e.g., are to be) located in an executable application (e.g., executable file 308, FAT binary file 309, etc.). In some examples, the jump table 412 associates different variants with the corresponding locations of the different variants in the executable application via variant symbols (e.g., pointers) that point to the locations of the corresponding variants in the executable application.
[0120] Turning to Figure 5 FIG., the illustrated example depicts a Figure 3 FAT binary file 309, including a Figure 3 variant library 310, a jump table library 312, and a runtime scheduler 314. In Figure 5In the illustrated example, the FAT binary file 309 includes a variant library 310 for storing example variant binaries 502, 504, 506, 508, 510, including a first example variant binary 502, a second example variant binary 504, a third example variant binary 506, a fourth example variant binary 508, and a fifth example variant binary 510. Alternatively, the variant library 310 may include one variant binary or a different number of variant binaries than depicted in Figure 5 In Figure 5 , the variant binaries 502, 504, 506, 508, 510 are based on Halide. Alternatively, one or more of the variant binaries 502, 504, 506, 508, 510 may be based on a different DSL, such as OpenCL.
[0121] In Figure 5 's illustrated example, the first variant binary 502 is a CPU variant binary and corresponds to the compilation of the first scheduling pair algorithm according to the first target configuration of the CPU 316 based on Figure 3 . In Figure 5 , the second variant binary 504 is a GPU variant binary and corresponds to the compilation of the second scheduling pair algorithm according to the second target configuration of the GPU 322 based on Figure 3 . In Figure 5 , the third variant binary 506 is an iStudio Publisher (ISPX) variant binary and corresponds to the compilation of the third scheduling pair algorithm according to the third target configuration corresponding to one of the CPU 316, FPGA 318, VPU 320, or GPU 322 based on Figure 3 . In Figure 5 , the fourth variant binary 508 is a VPU variant binary and corresponds to the compilation of the fourth scheduling pair algorithm according to the fourth target configuration of the VPU 320 based on Figure 3 . In Figure 5 , the fifth variant binary 510 is an FPGA variant binary and corresponds to the compilation of the fifth scheduling pair algorithm according to the fifth target configuration of the FPGA 318 based on Figure 3 . Thus, Figure 3 and / or Figure 5 's FAT binary file 309 contains multiple versions (e.g., multiple binary versions, multiple scheduling versions, etc.) of the same algorithm implemented on a series of processing elements available on the target heterogeneous system (e.g., the heterogeneous system 304 of Figure 3 ).
[0122] In Figure 5 's illustrated example, the jump table library 312 includes one or more jump tables, including Figure 4jump table 412. In Figure 5 the jump table 412 includes example metadata 512 corresponding to information related to the correct or desired usage, performance characteristics, and workload characteristics generated in response to executing a workload on a corresponding processing element. In Figure 5 the metadata 512 includes a data structure "IP_CPU_PERF" that can correspond to data associated with at least one of the following: (1) the correct or desired usage of the CPU 316, (2) the performance characteristics of the CPU 316, which can represent performance counter or execution graph values, or (3) workload characteristics associated with data generated by the CPU 316 or a performance monitoring unit (PMU) that monitors the CPU 316, where the data is generated in response to the CPU 316 executing a workload.
[0123] In Figure 5 the illustrated example, the jump table 412 includes example symbols (e.g., entry point symbols, variant symbols, etc.) 514. In Figure 5 each of the symbols 514 corresponds to a respective one of the variant binaries 502, 504, 506, 508, 510. For example, the symbol "_halide_algox_cpu" can correspond to the first variant binary 502. In Figure 5 the symbols 514 are based on Halide because Figure 5 the variant binaries 502, 504, 506, 508, 510 are based on Halide. Alternatively, when a corresponding one of the variant binaries 502, 504, 506, 508, 510 is based on a different DSL, one or more of the symbols 514 can be based on a different DSL, such as OpenCL.
[0124] In Figure 5 the illustrated example, the FAT binary file 309 includes a runtime scheduler 314 for accessing and / or otherwise invoking the jump table 412 of the jump table library 312. In Figure 5 the variant library 310 is linked to the jump table 412 via the symbol 514. In Figure 5 the runtime scheduler 314 is linked to the jump table library 312 via an example jump function 516. In operation, the runtime scheduler 314 executes an example application function 518 to execute a workload associated with an algorithm. In operation, the runtime scheduler 314 executes the jump function 516 to determine a workload allocation decision based on the metadata 512 (e.g., metadata 512 collected in real time or near real time). In operation, the runtime scheduler 314 makes a determination by identifying Figure 3a processing element among the processing elements for executing a workload to determine a workload distribution decision. For example, the runtime scheduler 314 can determine to use the GPU 322 based on the metadata 512. In such an example, the runtime scheduler 314 is capable of selecting "_halide_algox_gpu" from the symbol 514 to invoke the second variant binary 504 to execute the workload.
[0125] Return to Figure 4 the illustrated example of, the variant generator 302 includes an application compiler 414 for compiling an algorithm, each variant, variant symbols, and / or a runtime scheduler (e.g., Figure 3 the runtime scheduler 314) into one or more executable applications (e.g., executable file 308) for storage. For example, the application compiler 414 can communicate with the heterogeneous system 304 and store the one or more executable applications in Figure 3 the storage 306 of the Figure 3 heterogeneous system 304. In some examples, the application compiler 414 compiles the algorithm, each variant, and the runtime scheduler into a compiled version of the original algorithm (e.g., code, human-readable instructions, etc.) received by the variant generator 302 (e.g., from the code converter 303, Figure 2 one or more external devices, etc.). For example, if the algorithm is written in an imperative programming language such as C or C++, the application compiler 414 can compile the algorithm, each variant, variant symbols, and the runtime scheduler into an executable C or C++ application that includes the variants written in their respective languages for execution on the respective processing elements.
[0126] In some examples, the executable application compiled by the application compiler 414 is a FAT binary file. For example, the application compiler 414 can assemble Figure 3 the executable file 308 by linking the runtime scheduler 314, the jump table library 312, and the variant library 310 to compile the application for processing one or more algorithms. For example, the application compiler 414 can generate Figure 3 and / or Figure 5 the FAT binary file 309 based on one or more links. Alternatively, the executable application compiled by the application compiler 414 can be any suitable executable file.
[0127] In some examples, the application compiler 414 generates a jump table to be included in the jump table library. For example, the application compiler 414 can add Figure 4 the jump table 412 to Figure 3 and / or Figure 4in the jump table library 312. Accordingly, the application compiler 414 can generate the jump table library 312 in response to generating one or more jump tables, such as jump table 412.
[0128] In some examples, the application compiler 414 implements means for compiling the executable file 308 to include the runtime scheduler 314 to select, based on a schedule (e.g., a schedule generated by the compilation auto-scheduler 408), Figure 5 one or more of the variant binaries 502, 504, 506, 508, 510 to execute the workload of interest.
[0129] In Figure 4 the illustrated example of, the variant generator 302 includes a feedback interface 416 for interfacing between an executable application (e.g., the executable file 308) running on a heterogeneous system (e.g., Figure 3 the heterogeneous system 304) and / or a storage facility (e.g., the database 208). For example, the feedback interface 416 can be a network interface, a USB port interface, an Ethernet port interface, or a serial port interface. During the training phase, the feedback interface 416 can collect performance characteristics associated with the selected processing element. During the training phase, the collected performance characteristics can correspond to a quantification of the power consumption of the selected processing element, the time of running parameters on the selected processing element, and / or other performance characteristics associated with the selected processing element.
[0130] In Figure 4 the illustrated example of, during the inference phase, the feedback interface 416 can be configured to collect performance characteristics and performance deltas associated with a system-wide success function. In some examples, the feedback interface 416 obtains (e.g., directly obtains) performance characteristics from an application executing on the heterogeneous system and / or from a storage device external to the heterogeneous system.
[0131] In Figure 4 the illustrated example of, the variant generator 302 includes a performance analyzer 418 for identifying received data (e.g., performance characteristics). During the training phase, the performance analyzer 418 can determine whether the selected variant meets and / or otherwise satisfies a performance threshold. Additionally, during the training phase, the performance analyzer 418 can analyze the performance of the processing element to meet and / or otherwise satisfy the success function. During the initial training phase, the performance analyzer 418 can analyze the performance of a single processing element in isolation and may not consider the overall context of the processing elements in the heterogeneous system. The analysis of the single processing element can be fed back into the cost model learner 404 to help analyze and develop, for a specific processing element, a cost model that is more accurate than the previous cost model for that specific processing element for CNNs, DNNs, etc.
[0132] In response to outputting and / or otherwise generating a trained ML / AI model 214 for deployment (e.g., for use by an administrator, developer, etc.), the performance analyzer 418, after receiving an indication (e.g., from the feedback interface 416) that input data (e.g., runtime characteristics on a heterogeneous system under load) has been received, is able to identify aspects of the heterogeneous system that are targeted based on the success function and performance characteristics of the system. In some examples, the performance analyzer 418 determines a performance delta by determining the difference between the expected performance (e.g., a performance threshold) defined by the success function and the actual performance achieved during the execution of the algorithm during the inference phase.
[0133] In some examples, during a subsequent training phase (e.g., a second training phase after completion of a first training phase), additional empirical data obtained by the feedback interface 416 and utilized by the performance analyzer 418 can be reinserted into the cost model learner 404 to adjust one or more cost models of individual processing elements based on context data associated with the system as a whole (e.g., performance characteristics such as runtime load and environmental characteristics).
[0134] In some examples, the cost model learner 404 performs and / or otherwise implements various executable actions associated with different cost models for individual processing elements based on the context data. For example, based on the collected empirical data, the cost model learner 404 can adjust the cost models of individual processing elements to cause the compilation auto-scheduler 408 to generate a schedule using the adjusted cost models so as to execute a specified workload in a more desirable manner (e.g., by using less power, taking less time to execute, etc.).
[0135] In some examples, in response to determining that the performance characteristics indicate that a particular variant is not frequently selected, the performance analyzer 418 is able to determine that the variant targeted at a particular aspect associated with that variant is not a satisfactory candidate for offloading workloads during runtime. Based on this determination, the performance analyzer 418 can instruct the variant manager 402 not to generate variants for the associated aspect and / or associated processing element. Advantageously, by not generating additional variants, the variant manager 402 can reduce the space on the application (e.g., a FAT binary file) generated by the application compiler 414, reducing the memory consumed by the application when stored in memory.
[0136] In Figure 4In the illustrated example, when using the collected empirical data, the cost model learner 404 can additionally generate multiple cost models associated with a specific processing element by using additional CNNs, DNNs, etc. Each cost model can focus on a specific aspect of a specific processing element, and at runtime, a runtime scheduler (e.g., runtime scheduler 314) can select from multiple variants to use on Figure 3 the heterogeneous system 304. For example, if the overall system success function is associated with power savings, the runtime scheduler 314 can utilize variants targeted at reducing power consumption across all processing elements. However, when understanding the overall system performance under runtime execution (e.g., by collecting empirical data), the cost model learner 404 can generate multiple variants targeted at at least reducing power consumption and increasing speed. At runtime, the runtime scheduler 314 implementing the examples disclosed herein can determine and even execute that a variant targeted at increasing speed is still within the bounds of the success function associated with power savings. Advantageously, the runtime scheduler 314 can improve the performance of the entire heterogeneous system while still maintaining the functionality of meeting the required success function.
[0137] Although Figure 4 illustrates an example manner of implementing Figure 3 the variant generator 302, one or more of the elements, processes, and / or devices shown Figure 4 can be combined, split, rearranged, omitted, eliminated, and / or implemented in any way. Additionally, the example variant manager 402, example cost model learner 404, example weight store 406, example compile auto-scheduler 408, example variant compiler 410, example jump table 412, example application compiler 414, example feedback interface 416, example profiler 418, and / or more generally, Figure 3 the example variant generator 302 and / or more generally, Figure 3 the example variant application 305 can be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, the example variant manager 402, example cost model learner 404, example weight store 406, example compile auto-scheduler 408, example variant compiler 410, example jump table 412, example application compiler 414, example feedback interface 416, example profiler 418, and / or more generally, Figure 3The example variant generator 302 and / or, more generally, any party in the example variant application 305 can be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, GPUs, DSPs, ASICs, PLDs, and / or FPLDs. When reading any claim of the apparatus or system of this patent to cover a pure software and / or firmware implementation, the example variant manager 402, the example cost model learner 404, the example weight store 406, the example compilation auto-scheduler 408, the example variant compiler 410, the example jump table 412, the example application compiler 414, the example feedback interface 416, the example profiler 418, and / or, more generally, Figure 3 the example variant generator 302 and / or, more generally, Figure 3 at least one of the example variant applications 305 is thus explicitly defined to include a non-transitory computer-readable storage device or storage disk, such as a memory, a DVD, a CD, a Blu-ray disc, etc., including software and / or firmware. Additionally, the example variant generator 302 can include one or more elements, processes, and / or devices attached to or in place of Figure 3 the elements, processes, and / or devices shown in
[0138] Figure 6 depicts a block diagram illustrating Figure 3 an example implementation of the code converter 303 shown in
[0139] In Figure 6 the illustrated example of Figure 2 the code converter 303 obtains one or more algorithms from Figure 6In the illustrated example, the code converter 303 includes an example code interface 602, an example code lifter 604, an example domain-specific language (DSL) generator 606, an example metadata generator 608, an example code replacer 610, and an example database 612. In Figure 6 the database 612 includes example annotated code 614 and example unannotated code 616. In Figure 6 the code converter 303 further depicted is an example communication bus 618.
[0140] In Figure 6 the illustrated example, the code converter 303 includes a code interface 602 for obtaining code (e.g., application code, a part of the code, a code block, human-readable instructions, machine-readable instructions corresponding to the human-readable instructions, etc.) corresponding to (multiple) algorithms represented in an imperative programming language to be executed on the Figure 3 heterogeneous system 304. For example, the code can correspond to the example application code 700 described below in connection with Figure 7 For example, the code can be associated with one or more workloads to be processed by the Figure 3 heterogeneous system 304. In some examples, the code interface 602 obtains one or more algorithms from one or more external devices, database 208, database 612, etc. of Figure 2 For example, the code interface 602 can obtain one or more algorithms from a user (e.g., an administrator, a developer, a computing device associated with the administrator or developer, etc.).
[0141] In Figure 6 the illustrated example, the code converter 303 includes a code lifter 604 for identifying example annotated code 614 (e.g., one or more annotated code blocks, one or more annotated code portions, etc.) included in the code obtained by the code interface 602. In some examples, the code lifter 604 identifies (multiple) annotated code blocks based on the presence of an identifier (e.g., an annotated code identifier, an offload identifier, etc.) corresponding to an imperative programming language compiler extension (e.g., %pragma, #pragma, etc.). For example, a user or a computing device associated with the user can insert an identifier at one or more code sites, locations, positions, etc. within the application code to specify, mark, and / or otherwise indicate the intention to offload the execution of the specified code on one or more processing elements of the heterogeneous system 304.
[0142] In some examples, the code lifter 604 identifies uncommented code 616 (e.g., one or more uncommented code blocks, one or more uncommented code sections, etc.) included in the code obtained by the code interface 602. In such examples, the code lifter 604 is able to identify the uncommented code 616 based on the presence of a lack of identifiers corresponding to imperative programming language compiler extensions.
[0143] In some examples, in response to identifying the commented code 614, the code lifter 604 elevates the algorithmic intent of the commented code 614 in a first representation corresponding to an imperative programming language to a second representation corresponding to a formal or elevated intermediate representation. For example, the code lifter 604 can elevate the algorithmic intent by performing and / or otherwise implementing one or more verified elevation techniques. The code lifter 604 can perform a verified elevation by taking as input a block of commented code 614 written in an imperative general-purpose language (e.g., an imperative programming language) (e.g., code potentially optimized for heterogeneous offloading) and inferring a summary expressed in a high-level or formal specification language (e.g., an elevated intermediate representation, a predicate-logic-based programming language, a functional (e.g., pure functional) programming language, etc.) that is provably semantically equivalent to the original program or code. For example, the code lifter 604 can determine that the commented code 614 is a complex nested loop sequence that can be replaced by a simple five point stencil.
[0144] In some examples, the code lifter 604 infers the summary by searching for possible candidate programs in a target language (e.g., DSL) into which the commented code 614 can be converted. For example, the code lifter 604 can search Figure 2 database 208 of Figure 6 database 612 of Figure 2 database 208 of Figure 6 database 612 of
[0145] In some examples, the code lifter 604 finds (e.g., automatically finds) a lifted summary expressed in a predicate language by using inductive synthesis. For example, the code lifter 604 can find the lifted summary by querying an ML / AI model (e.g., Figure 2 the trained ML / AI model 214), Figure 2 the database 208, Figure 6 the database 612, etc. and / or combinations thereof. For example, the database 612 can include the trained ML / AI model 214. In such an example, the code lifter 604 can query the trained ML / AI model 214 with the annotated code 614 to identify the lifted summary. In other examples, the code lifter 604 can transmit the annotated code 614 to Figure 2 the database 208 to obtain the lifted summary in response to the query. In response to obtaining the lifted summary, the code lifter 604 can transmit and / or otherwise provide the lifted summary to the DSL generator 606 to be converted into a high-performance DSL and then redirected by the variant generator 302 for execution on different architectures or target hardware as needed.
[0146] Figure 6 The code lifter 604 in the illustrated example of
[0147] In some examples, the code lifter 604 determines that the intermediate code is a correct summary of the annotated code 614 by constructing verification conditions. For example, the intermediate code can correspond to a postcondition for the annotated code 614. The postcondition can correspond to a predicate that is true at the end of a code block under all possible executions, as long as the input to the code block satisfies certain given preconditions. In some examples, the code lifter 604 determines that the intermediate code is valid with respect to the given preconditions and postconditions by constructing verification conditions. The verification condition can correspond to a formula (e.g., a mathematical formula), and if the formula is true, it means that the intermediate code is valid with respect to the preconditions and postconditions.
[0148] In some examples, the code lifter 604 uses syntax-guided synthesis to search a large range of possible invariants and postconditions. For example, the code lifter 604 can search the database 208, query the trained ML / AI model 214, etc., in order to identify invariants (e.g., loop invariants) for each loop of the annotated code 614, and the corresponding postconditions, which can cause the theorem prover to determine valid verification conditions. For example, the code lifter 604 can identify one or more candidate programs or candidate code blocks stored in the repository, and these candidate programs or candidate code blocks include and / or otherwise correspond to invariant and postcondition pairs.
[0149] In some examples, the code lifter 604 confirms the intermediate code by validating the verification conditions. For example, the code lifter 604 can search a large range of possible invariants (e.g., loop invariants) when validating one or more logical statements. An invariant can refer to a formal statement about the relationships between variables in a code loop, which is true before running the code loop and is again true at the end of the code loop each time the code loop is traversed. As a result of the search, the code lifter 604 can find the invariants and postconditions for each loop (e.g., each loop of the intermediate code), and the invariants and postconditions can jointly cause the theorem prover (e.g., an executable file or application code that can prove theorems) to prove valid verification conditions. In response to confirming the intermediate code, the code lifter 604 can transmit the intermediate code to the DSL generator 606 for processing.
[0150] In some examples, the code lifter 604 implements means for identifying annotated code corresponding to an algorithm to be executed on the heterogeneous system 304 based on identifiers associated with the annotated code 614, where the annotated code 614 is in a first representation. In some examples, the code lifter 604 implements means for converting the annotated code in the first representation to intermediate code by identifying intermediate code in a second representation as having a first algorithmic intent corresponding to a second algorithmic intent of the annotated code. For example, the means for identifying and / or the means for converting may be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, GPUs, DSPs, application specific integrated circuits (ASICs), PLDs, and / or FPLDs.
[0151] In Figure 6 the illustrated example of, the code converter 303 includes a DSL generator 606 for transforming, converting, and / or otherwise converting the intermediate code into DSL code (e.g., code in a DSL representation). In response to the code lifter 604 identifying a valid post-condition (e.g., a lifted summary, an inferred summary, etc.), the DSL generator 606 can convert the intermediate code into DSL code (e.g., Halide code, OpenCL code, etc.). For example, the DSL generator 606 can generate a C++ source file that, when compiled and executed, produces an object file (e.g., a compiled version of the unannotated code 616) that can be linked with the original application. In some examples, the DSL generator 606 generates an additional code portion (e.g., glue code or stitching code) that can interface between the code blocks of the original application and the DSL code.
[0152] In some examples, the DSL generator 606 implements means for converting the intermediate code into DSL code in a DSL representation when the first algorithmic intent of the annotated code 614 matches the second algorithmic intent of the intermediate code. For example, the means for converting may be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, GPUs, DSPs, application specific integrated circuits (ASICs), PLDs, and / or FPLDs.
[0153] In Figure 6 the illustrated example of, the code converter 303 includes a metadata generator 608 for generating metadata that can be used to configure Figure 4Metadata (e.g., scheduling metadata) of the compilation automatic scheduler 408. In some examples, the metadata generator 608 embeds and / or otherwise includes metadata within a code block of the annotated code 614. For example, the metadata generator 608 can embed metadata in the annotated code 614 before the code lifter 604 obtains the annotated code 614. In such examples, the code lifter 604 can lift the algorithmic intent from the annotated code 614 while leaving the corresponding metadata unprocessed. The code lifter 604 can transmit the intermediate code and the corresponding metadata to the variant generator 302 to configure the compilation automatic scheduler 408 to compile the variant. Alternatively, the metadata generator 608 can transmit the metadata to the compilation automatic scheduler 408.
[0154] In some examples, the metadata generator 608 generates metadata that includes scheduling information or data corresponding to the power profile of the heterogeneous system 304, where the power profile indicates whether low power consumption is preferred over maximum performance and to what extent it is preferred over maximum performance. For example, the power profile can correspond to the heterogeneous system 304 operating in a first power state, a second power state, etc., where the first power state has a higher power consumption capacity than the second power state. In some examples, the metadata generator 608 generates metadata that includes instructions specifying which of the (one or more) processing elements and / or (one or more) computing APIs are to be used for an algorithm of interest. In some examples, the metadata generator 608 generates metadata that includes instructions specifying to use a particular cost model, multiple variants of one or more processing elements to be used, lossy or reduced bitwidth variants, etc. and / or a combination thereof.
[0155] In some examples, the metadata generator 608 implements means for generating scheduling metadata corresponding to the power profile of the heterogeneous system 304. For example, the means for generating can be implemented by one or more analog or digital circuits, logic circuits, (one or more) programmable processors, (one or more) programmable controllers, (one or more) GPUs, (one or more) DSPs, (one or more) application specific integrated circuits ((one or more)(ASICs)), (one or more) PLDs, and / or (one or more) FPLDs.
[0156] In Figure 6 the illustrated example, the code converter 303 includes a code replacer 610 for modifying the original application code obtained by the code interface 602 to bind the (one or more) sites of the annotated code 614 so as to call the algorithm of interest via Figure 3 the runtime scheduler 314. In some examples, the code replacer 610 replaces a code block of the annotated code 614 with a function call to the runtime scheduler 314.
[0157] In some examples, the code replacer 610 modifies the original application code to capture data semantics (e.g., data dependencies) for heterogeneous offloading. For example, the code replacer 610 may generate (one or more) memory allocation routines to include in the original application code to ensure compatibility across processing elements. For example, the code replacer 610 may generate (one or more) memory allocation routines to copy data used in heterogeneous offloading to a shared memory buffer visible to all processing elements of the heterogeneous system 304.
[0158] In some examples, the code replacer 610 invokes a compiler (e.g., an imperative programming language compiler, a C / C++ compiler, etc.) to compile Figure 3 the executable file 308. For example, the code replacer 610 can invoke the compiler to compile the executable file 308 by compiling a modified version of the application code into the executable file 308. For example, the modified version can correspond to the application code where the annotated code 614 is replaced with one or more function calls to the runtime scheduler 314. In such an example, the annotated code 614 can be replaced with one or more function calls based on data semantics associated with and / or otherwise used for heterogeneous offloading.
[0159] In some examples, the code replacer 610 implements means for replacing the annotated code 614 in the application code 700 with function calls to the runtime scheduler 314 in order to call one of the variant binaries 502, 504, 506, 508, 510 to load it onto one of the processing elements of the heterogeneous system 304 for execution of the workload of interest. For example, the means for replacement can be implemented by one or more analog or digital circuits, logic circuits, (one or more) programmable processors, (one or more) programmable controllers, (one or more) GPUs, (one or more) DSPs, (one or more) application specific integrated circuits ((one or more)(ASICs)), (one or more) PLDs, and / or (one or more) FPLDs.
[0160] In some examples, the code replacer 610 implements means for invoking a compiler to generate the executable file 308 including the variant binaries 502, 504, 506, 508, 510 based on DSL code, where each of the variant binaries 502, 504, 506, 508, 510 can call a corresponding one of the processing elements of the heterogeneous system 304 to execute the algorithm of interest. For example, the means for invocation can be implemented by one or more analog or digital circuits, logic circuits, (one or more) programmable processors, (one or more) programmable controllers, (one or more) GPUs, (one or more) DSPs, (one or more) application specific integrated circuits ((one or more)(ASICs)), (one or more) PLDs, and / or (one or more) FPLDs.
[0161] In Figure 6 In the illustrated example of Figure 6 , the code converter 303 includes a database 612 for recording and / or otherwise storing data (e.g., annotated code 614, unannotated code 616, loop invariants, preconditions, postconditions, etc.). The database 612 can be implemented by volatile memory (e.g., SDRAM, DRAM, RDRAM, etc.) and / or non-volatile memory (e.g., flash memory). The database 612 can additionally or alternatively be implemented by one or more DDR memories (such as DDR, DDR2, DDR3, DDR4, mDDR, etc.). The database 612 can additionally or alternatively be implemented by one or more mass storage devices (such as (multiple) HDDs, (multiple) CD drives, (multiple) DVD drives, (multiple) SSD drives, etc.). Although the database 612 is shown as a single database in the illustrated example, the database 612 can be implemented by any number and / or (multiple) types of databases. Additionally, the data stored in the database 612 can be in any data format, such as, for example, binary data, comma-separated data, tab-separated data, SQL structures, etc. In Figure 6 Figure 6 , the database 612 is stored on an electronically accessible computing system. For example, the database 612 can be stored on a server, a desktop computer, an HDD, an SSD, or any other suitable computing system.
[0162] In Figure 6 In the illustrated example of Figure 6 , the code converter 303 includes a communication bus 618 for facilitating communication operations associated with the code converter 303. For example, the communication bus 618 can correspond to an interface circuit or any other type of hardware or logic circuit for facilitating communication between components. In other examples, the communication bus 618 can be implemented in software. In some examples, one or more of the code interface 602, code lifter 604, DSL generator 606, metadata generator 608, and / or code replacer 610 communicate via any suitable wired and / or wireless communication method. In some examples, one or more of the code interface 602, code lifter 604, DSL generator 606, metadata generator 608, and / or code replacer 610 are capable of communicating with any processing element or hardware component external to the runtime scheduler 314 via any suitable wired and / or wireless communication method.
[0163] Although Figure 6 illustrates an example manner of implementing Figure 3 the code converter 303 of Figure 3 , Figure 6One or more of the illustrated components, processes, and / or devices may be combined, split, rearranged, omitted, eliminated, and / or implemented in any manner. Additionally, example code interface 602, example code booster 604, example DSL generator 606, example metadata generator 608, example code replacer 610, example database 612, example communication bus 618, and / or more generally, Figure 3 example code converter 303 of Figure 3 and / or more generally, Figure 3 example code converter 303 of Figure 3 and / or more generally, Figure 3 example variant application 305 of may be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, example code interface 602, example code booster 604, example DSL generator 606, example metadata generator 608, example code replacer 610, example database 612, example communication bus 618, and / or more generally, Figure 3 example code converter 303 of Figure 6 and / or more generally,
[0164] Figure 7 Example application code 700 is depicted, which includes Figure 6 annotated code 614 and unannotated code 616 of Figure 2 database 208 of Figure 2 and / or one or more external devices of Figure 7Among them, the code converter 303 can identify the annotated code 614 by identifying and / or otherwise determining the existence of an imperative programming language compiler extension. In Figure 7 's example, the code "#pragma intent" can correspond to an imperative programming language compiler extension. For example, the code converter 303 can identify the annotated code 614 as intentional code. Therefore, the code converter 303 can lift the intentional code from the imperative programming language representation to a formal representation corresponding to the intermediate code. In Figure 7 's example, the code converter 303 can identify the unannotated code 616 by determining that the unannotated code 616 is not annotated and / or otherwise marked by an imperative programming language compiler extension.
[0165] Figure 8 is an example workflow 800 for converting Figure 7 's application code 700 into a DSL representation. In Figure 8 's workflow 800, during the first operation 802, in response to identifying "#pragma intent" in the application code 700, the code converter 303 extracts the annotated code 614 from the application code 700. During the first operation 802, the code converter 303 identifies the annotated code 614 as an intentional block 804 (e.g., an intentional code block, a code block corresponding to an algorithmic intent, etc.).
[0166] In Figure 's illustrated example, during the second operation 806, the code converter 303 uses inductive synthesis to convert the annotated code 614 from the imperative programming language representation to an example intermediate code (e.g., a lifted intermediate code) 808 in a lifted intermediate representation. For example, the intermediate code 808 can be represented and / or otherwise expressed in a high-level or formal specification language (e.g., a programming language based on predicate logic, a functional (e.g., pure functional) programming language, etc.) that is provably semantically equivalent to the annotated code 614. During the third operation 810, the code converter 303 can transfer the intermediate code 808 to the variant generator 302. During the third operation 810, the variant generator 302 converts the intermediate code 808 from the lifted intermediate representation to an example DSL code 812 in a DSL representation. Alternatively, the order of execution of the operations can be changed, and / or some of the described operations can be changed, eliminated, or combined.
[0167] is an example workflow 900 for converting 's application code 700 into an example variant code 902. For example, 's workflow 900 can correspond to an extended workflow of 's workflow 800. In In the illustrated example, in response to the DSL code 812 is generated, and the code converter 303 can transmit and / or otherwise provide the DSL code 812 to the variant generator 302. In it, the variant generator 302 generates an example cost model 904, which includes a CPU cost model 906, a GPU cost model 908, and an FPGA cost model 910. For example, the cost model learner 404 can generate the cost model 904 based on the DSL code 812. Alternatively, the cost model 904 can include fewer or more cost models than depicted in
[0168] In the workflow 900, the variant generator 302 can generate an example schedule 912, which includes an example CPU schedule 914, an example GPU schedule 916, and an example FPGA schedule 918. In the workflow 900, the variant generator 302 can generate variant code 902 based on the schedule 912. For example, the variant code 902 can correspond to one of the variant binaries 502, 504, 506, 508, 510 that can be executed by a corresponding one of the processing elements of the heterogeneous system 304. Alternatively, the execution order of the operations can be changed, and / or some of the described operations can be changed, eliminated, or combined.
[0169] is an example workflow 1000 for compiling and / or the FAT binary file 309. In the illustrated example, during a first operation, a user (e.g., an administrator, a developer, etc.) develops, generates, and / or obtains the application code 700. During a second operation, the code converter 303 identifies the annotated code 614 and extracts the identified annotated code 614. During a third operation, the code converter 303 elevates the algorithmic intent from the annotated code 614. For example, the code converter 303 can convert the annotated code 614 from a first representation to a second representation, where the second representation corresponds to an elevated intermediate representation. In such an example, the code converter 303 converts the annotated code 614 to the intermediate code 808.
[0170] In the illustrated example, during a fourth operation, the code converter 303 lowers the intermediate code 808 to correspond to a DSL representation The DSL code 812. In In the illustrated example, the variant generator 302 implements and / or otherwise facilitates the execution of the aspect separation compiler 1002. In In, the variant generator 302 includes the aspect separation compiler 1002 for outputting the variant library 310 of the jump table library 312 of etc. for generating the executable file 308 of. The aspect separation compiler 1002 utilizes the principle of aspect separation to separate how an example algorithm 1004 represented in DSL is written from how the algorithm 1004 is executed. For example, the aspect separation compiler 1002 can implement
[0171] In In the illustrated example, during the fifth operation, the aspect separation compiler 1002 selects the variant of interest (e.g., variant binary) corresponding to the processing element of interest. For each variant (e.g., executed in parallel, in sequence, etc.), the aspect separation compiler 1002 selects a scheduling heuristic corresponding to the variant of interest during the sixth operation and selects a target (e.g., hardware target, processing element, etc.) during the seventh operation. In response to selecting the scheduling heuristic at the sixth operation, the aspect separation compiler 1002 configures the compilation auto-scheduler 408 of. For example, the compilation auto-scheduler 408 can be configured based on the example scheduling metadata 1006 corresponding to the annotated code 614.
[0172] In In the illustrated example, during the ninth operation, the aspect separation compiler 1002 compiles the variant binary in advance (AOT) for each variant of interest. During the ninth operation, the aspect separation compiler 1002 stores the compiled variant binary in the variant library 310 of. In In, during the tenth operation, the aspect separation compiler 1002 calls the variant generator 302 to add jump table entries corresponding to the compiled variant binary. For example, jump table entries of the jump table 412 can be generated, and a corresponding variant symbol in the variant symbol 514 can be stored at the jump table entry.
[0173] In In the illustrated example, during the eleventh operation, after each variant has been compiled and / or otherwise processed, the variant generator 302 generates a jump table (e.g., jump table 412), and stores the jump table in the jump table library 312. During the twelfth operation, the variant generator 302 generates , and / or the FAT binary file 309 by linking at least and / or the variant library 310, the jump table library 312, and the runtime scheduler 314. During the thirteenth operation, the code converter 303 replaces the annotated code 614 with one or more runtime scheduler calls. For example, the code converter 303 can replace a code block of the annotated code 614 with a function call to the runtime scheduler 314 to call the jump table library 312. During the fourteenth operation, the code converter 303 captures and / or otherwise determines the data dependencies between the application code 700 and the annotated code 614. For example, the code converter 303 can link the replaced portion of the annotated code 614 with the application code 700.
[0174] In the illustrated example, during the fifteenth operation, the variant generator 302 compiles the FAT binary file 309. For example, the variant generator 302 can compile the FAT binary file 309 by linking at least one of the variant library 310, the jump table library 312, or the runtime scheduler 314. During the sixteenth operation, the code converter 303 invokes an example imperative programming language compiler 1008 (e.g., a C / C++ compiler) to compile the executable file 308. For example, the code converter 303 can call the compiler 1008 to compile the executable file 308 by compiling a modified version of the application code 700 into the executable file 308. For example, the modified version can correspond to the application code 700, where the annotated code 614 is replaced with one or more function calls to the runtime scheduler 314. In such an example, the annotated code 614 can be replaced with one or more function calls based on the data dependencies captured during the fourteenth operation.
[0175] In In the illustrated example, compiler 1008 can compile executable file 308 by at least linking variant library 310, jump table library 312, runtime scheduler 314, and a modified version of application code 700. Alternatively, compiler 1008 can be included in code converter 303, variant generator 302, and / or more generally, variant application 305. Alternatively, the order of execution of operations can be changed, and / or some of the described operations can be changed, eliminated, or combined.
[0176] In , , and / or shown are diagrams representing example hardware logic, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing , , , and / or of variant generator 302 and / or , , and / or of code converter 303. The machine-readable instructions can be one or more executable programs or (multiple) portions of an executable program executed by a computer processor, such as processor 1712 shown in example processor platform 1700 discussed below. Although the program can be embodied in software stored on a non-transitory computer-readable storage medium associated with processor 1712, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disc, or memory, all or portions of the program can alternatively be executed by a device other than processor 1712 and / or embodied in firmware or dedicated hardware. Additionally, although the example program is described with reference to the flowcharts shown in , , and / or , many other methods for implementing example variant generator 302 and / or code converter 303 can alternatively be used. For example, the order of execution of the blocks can be changed, and / or some of the described blocks can be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks can be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGA, ASIC, comparator, operational amplifier (op-amp), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware.
[0177] Additionally, in and / or shown are diagrams representing for implementing and / or A flowchart of example hardware logic, machine-readable instructions, a hardware-implemented state machine, and / or any combination thereof for an example executable file 308. The machine-readable instructions can be one or more executable programs or portions of an executable program that are executed by a computer processor, such as processor 1812 shown in example processor platform 1800 discussed below. Although the program can be embodied in software stored on a non-transitory computer-readable storage medium associated with processor 1812, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disc, or memory, all or portions of the program can alternatively be executed by a device other than processor 1812 and / or embodied in firmware or dedicated hardware. Further, although the example program is described with reference to the flowchart illustrated in and / or many other methods of implementing example executable file 308 can alternatively be used. For example, the order of execution of the blocks can be changed, and / or some of the blocks described can be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks can be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware.
[0178] The machine-readable instructions described herein can be stored in one or more of a compressed format, an encrypted format, a fragmented format, a compiled format, an executable format, a packaged format, etc. The machine-readable instructions described herein can be stored as data (e.g., portions of instructions, code, code representations, etc.) that can be used to create, manufacture, and / or generate machine-executable instructions. For example, the machine-readable instructions can be segmented and stored on one or more storage devices and / or computing devices (e.g., servers). The machine-readable instructions may need to be installed, modified, adapted, updated, combined, supplemented, configured, decrypted, decompressed, unpacked, distributed, redistributed, compiled, etc. in order for them to be directly readable, interpretable, and / or executable by a computing device and / or other machine. For example, the machine-readable instructions can be stored in multiple portions that are individually compressed, encrypted, and stored on separate computing devices, where the portions form a set of executable instructions that implement a program as described herein when decrypted, decompressed, and combined.
[0179] In another example, the machine-readable instructions may be stored in a state in which they are readable by a computer, but libraries (e.g., dynamic link libraries (DLLs)), software development kits (SDKs), application programming interfaces (APIs), etc. need to be added in order to execute the instructions on a particular computing device or other device. In another example, it may be necessary to configure the machine-readable instructions (e.g., stored settings, data inputs, recorded network addresses, etc.) before the machine-readable instructions and / or corresponding program(s) can be executed, in whole or in part. Accordingly, the disclosed machine-readable instructions and / or corresponding program(s) are intended to encompass such machine-readable instructions and / or program(s), regardless of the particular format or state of the machine-readable instructions and / or program(s) when stored or otherwise at rest or in transit.
[0180] The machine-readable instructions described herein may be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, the machine-readable instructions may be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, HyperText Markup Language (HTML), Structured Query Language (SQL), Swift, etc.
[0181] As mentioned above, the example processes may be implemented using executable instructions (e.g., computer-readable instructions and / or machine-readable instructions) stored on a non-transitory computer and / or machine-readable medium such as a hard disk drive, flash memory, read-only memory, CD, DVD, cache, random access memory, and / or any other storage device or storage disk that stores information therein for any length of time (e.g., for an extended period of time, permanently, during a brief instance, during a temporary buffer and / or information cache). As used herein, the term non-transitory computer-readable storage medium is expressly defined to include any type of computer-readable storage device and / or storage disk and to exclude propagated signals and to exclude transmission media.
[0182] "Comprising" and "including" (and all forms and tenses thereof) are used herein as open-ended terms. Thus, whenever a claim uses any form of "comprising" or "including" (e.g., includes, comprises, has, etc.) as a preamble or within any kind of claim recitation, it is to be understood that additional elements, items, etc. may exist without falling outside the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used as a transitional term, e.g., in conjunction with a claim, it is as open-ended as the terms "comprising" and "including". When the term "and / or" is used in the form such as A, B, and / or C, it refers to any combination or subset of A, B, and C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, and (7) A and B and C. As used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A and B" is intended to represent an implementation including any one of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A or B" is intended to represent an implementation including any one of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. As used herein in the context of describing the processing or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A and B" is intended to represent an implementation including any one of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein in the context of describing the processing or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" is intended to represent an implementation including any one of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B.
[0183] As used herein, singular references (e.g., "a", "an", "first", "second", etc.) do not exclude a plurality. The term "a" or "an" entity as used herein refers to one or more of that entity. The terms "a" (or "an"), "one or more", and "at least one" may be used interchangeably herein. Additionally, although listed separately, multiple devices, elements, or method acts may be implemented by, for example, a single unit or processor. Further, although individual features may be included in different examples or claims, these features may be combined, and the inclusion in different examples or claims does not imply that the combination of features is not feasible and / or advantageous.
[0184] is a flowchart depicting example machine-readable instructions 1100 that may be executed to implement , , and / or of the code converter 303 and / or , , , and / or of the variant generator 302 for invoking an application to execute a (multiple) workload on heterogeneous system 100 and / or , and / or heterogeneous system 304. The machine-readable instructions 1100 start at block 1102, where the code converter 303 and / or more generally, the variant application 305 ( ) obtains application code corresponding to the (multiple) algorithms represented in an imperative programming language to be executed on the heterogeneous system. For example, the code interface 602 ( ) may obtain from the database 208 of , one or more external devices, etc. and / or a combination thereof of application code 700. In such an example, the application code 700 can correspond to one or more algorithms written and / or otherwise developed in an imperative programming language (such as C or C++). Thus, one or more algorithms can be represented in a first representation corresponding to the imperative programming language.
[0185] At block 1104, the code converter 303 and / or more generally, the variant application 305 identifies the (multiple) annotated code blocks in the application code. For example, the code lifter 604 ( ) may identify and / or of the annotated code 614 as intentional code in response to identifying the imperative programming language compiler extension "#pragma". In such an example, the code lifter 604 can identify the annotated code 614 based on the beginning of the annotated code 614 being annotated with the "#pragma" identifier and / or otherwise identified.
[0186] At block 1106, the code converter 303 and / or more generally, the variant application 305 selects the annotated code blocks of interest to process. For example, the code lifter 604 may identify the annotated code blocks corresponding to the annotated code 614 depicted in to process. In such an example, the code lifter 604 can identify the algorithm associated with the annotated code block to be processed.
[0187] At block 1108, the code converter 303 and / or, more generally, the variant application 305 elevates the algorithmic intent from the annotated code block to the elevated intermediate representation. For example, the code lifter 604 can use the proven lift to elevate the algorithmic intent from the annotated code 614 to generate intermediate code 808.
[0188] At block 1110, the code converter 303 and / or, more generally, the variant application 305 reduces the annotated code block to a DSL representation. For example, the DSL generator 606 can convert the intermediate code 808 to DSL code 812, where the DSL code 812 is in the DSL representation.
[0189] At block 1112, the code converter 303 and / or, more generally, the variant application 305 replaces the annotated code block with a function call to the runtime scheduler. For example, the code replacer 610( ) can replace the annotated code 614 in the application code 700 with one or more function calls to the runtime scheduler 314. In such an example, in response to the heterogeneous system 304 executing the application code 700 compiled in the executable file 308 at the site of the replaced code (e.g., the code site) (e.g., the code site of the annotated code 614), the executable file 308 can execute one or more function calls to invoke the runtime scheduler 314 in order to execute the algorithm corresponding to the annotated code 614 using at least one of the processing elements of the heterogeneous system 304.
[0190] At block 1114, the code converter 303 and / or, more generally, the variant application 305 determines whether to select another annotated code block of interest to process. For example, the code lifter 604 can identify another code block in the application code 700 that has been marked, identified, and / or otherwise annotated with a "#pragma".
[0191] If, at block 1114, the code converter 303 and / or, more generally, the variant application 305 determines that there is another annotated code block of interest to process, control returns to block 1106 to select another annotated code block of interest to process. If, at block 1114, the code converter 303 and / or, more generally, the variant application 305 determines that there is no other annotated code block of interest to process, then, at block 1116, the variant generator 302 and / or, more generally, the variant application 305 compiles the variant(s). For example, the variant compiler 410( ) can compile based on the DSL code 812 one or more of the variant binaries 502, 504, 506, 508, 510. The following, in conjunction with describes an example process that can be used to implement block 1116.
[0192] At block 1118, the code converter 303 and / or, more generally, the variant application 305 modifies the application code to capture data semantics for offloading to a heterogeneous system. For example, the code replacer 610 can generate one or more memory allocation routines to be included in the application code 700 to ensure compatibility across processing elements. For example, the code replacer 610 can generate the one or more memory allocation routines to copy data used in the heterogeneous offload to a shared memory buffer visible to all processing elements of the heterogeneous system 304.
[0193] At block 1120, the code converter 303 and / or, more generally, the variant application 305 compiles the application. For example, the code replacer 610 can call the compiler 1008 of
[0194] At block 1122, the code converter 303 and / or, more generally, the variant application 305 determines whether additional application code corresponding to another algorithm is to be selected for processing. If, at block 1122, the code converter 303 and / or, more generally, the variant application 305 determines that there is additional application code and / or additional algorithms of interest to process, control returns to block 1102 to obtain the additional application code and / or the additional algorithms. If, at block 1122, the code converter 303 and / or, more generally, the variant application 305 determines that there is no additional code and / or algorithms of interest to process, then, at block 1124, the heterogeneous system 304 invokes the application to execute the workload(s). The following, in conjunction with describes an example process that can be used to implement block 1124. In response to executing the application at block 1124 to execute the workload(s), the machine-readable instructions 1100 of
[0195] is a flowchart representing the machine-readable instructions 1116 that can be executed to implement the variant generator 302 of The process of can be used to implement block 1116 of Machine-readable instructions 1116 start at block 1202, where the variant generator 302 and / or more generally, the variant application 305 obtains the configuration(s) of the hardware target(s) of interest. For example, the variant manager 402 ( ) may obtain the configuration (e.g., the target configuration) associated with the heterogeneous system 304. In such an example, the variant manager 402 can obtain the target configuration from the database 208, one or more external devices, , and / or the heterogeneous system 304, etc. and / or combinations thereof. The target configuration may include information indicating the heterogeneous system 304, which includes the CPU 316, FPGA 318, VPU 320, and / or GPU 322.
[0196] At block 1204, the variant generator 302 and / or more generally, the variant application 305 selects the hardware target of interest. For example, the compilation auto-scheduler 408 ( ) may select the CPU 316 for processing. At block 1206, the variant generator 302 and / or more generally, the variant application 305 configures the auto-scheduler for the hardware target based on the scheduling metadata and the corresponding configuration. For example, the compilation auto-scheduler 408 may be configured based on the target configuration associated with the CPU 316 and / or the scheduling metadata 1006. In such examples, the compilation auto-scheduler 408 can be configured based on the hardware architecture, scheduling heuristics, etc. associated with the CPU 316. In response to this configuration, the compilation auto-scheduler 408 can generate a schedule, one or more execution graphs, etc. that can be used by the CPU 316 to execute the workload.
[0197] At block 1208, the variant generator 302 and / or more generally, the variant application 305 compiles the variant. For example, the variant compiler 410 ( ) may compile the first variant binary 502 based on the target configuration, schedule, one or more execution graphs, scheduling metadata 1006, etc. associated with the CPU 316 and / or combinations thereof. At block 1210, the variant generator 302 and / or more generally, the variant application 305 adds the variant to the variant library. For example, the variant compiler 410 may add the first variant binary 502 to the variant library 310.
[0198] At block 1212, the variant generator 302 and / or more generally, the variant application 305 adds variant symbols to the jump table. For example, the variant compiler 410 can add a variant symbol corresponding to the first variant binary 502 to the and / or jump table library 312 of the jump table 412. In such an example, the variant compiler 410 can add the variant symbol “_halide_algox_cpu” to correspond to the first variant binary 502 of “ALGOX_CPU.O” as depicted in the illustrated example such as .
[0199] At block 1214, the variant generator 302 and / or more generally, the variant application 305 determines whether to select another hardware target of interest. For example, the compilation auto-scheduler 408 can select the FPGA 318 to process. If at block 1214, the variant generator 302 and / or more generally, the variant application 305 determines to select another hardware target of interest, the control returns to block 1204 to select another hardware target of interest. If at block 1214, the variant generator 302 and / or more generally, the variant application 305 determines not to select another hardware target of interest, then at block 1216, the variant generator 302 and / or more generally, the variant application 305 adds the jump table to the jump table library. For example, the application compiler 414 ( ) can add the jump table 412 to and / or the jump table library 312. In response to adding the jump table to the jump table library at block 1216, the machine-readable instructions 1116 return to block 1118 of the machine-readable instructions 1100 to modify the application code to capture the data semantics for offloading to the heterogeneous system.
[0200] is a flowchart representing example machine-readable instructions 1124 that can be executed to implement , and / or the heterogeneous system 304 for invoking an application to execute a (multiple) workload. The process of can be used to implement block 1124. The machine-readable instructions 1124 begin at block 1302 where the heterogeneous system 304 obtains the workload to be processed. For example, the heterogeneous system 304 can obtain from database 208, one or more external devices, obtaining algorithms to be processed from storage 306, storage 306, etc.
[0201] At block 1304, heterogeneous system 304 obtains performance characteristics of heterogeneous system 304. For example, runtime scheduler 314 ( ) can obtain metadata 512 to obtain performance characteristics associated with CPU 316, FPGA 318, VPU 320, and / or GPU 322 of
[0202] At block 1306, heterogeneous system 304 determines (a) processing element(s) for deploying a workload based on the performance characteristics. For example, runtime scheduler 314 can determine to use GPU 322 to execute the workload based on the performance characteristics corresponding to GPU 322 indicating that GPU 322 has available bandwidth to execute the workload. In other examples, even if GPU 322 cannot be used to execute the workload at a first moment, runtime scheduler 314 can still determine to use GPU 322 at the first moment. For example, runtime scheduler 314 can determine that GPU 322 can execute the workload at a second moment after the first moment. In such an example, runtime scheduler 314 can determine that waiting until the second moment to use GPU 322 to execute the workload is faster or provides different benefits (e.g., different amounts of power consumption, more efficient use of processing elements, etc.) compared to using one of the available processing elements among multiple processing elements to execute the workload at the first moment.
[0203] At block 1308, heterogeneous system 304 calls one or more variant binaries by invoking the corresponding symbol(s) in a jump table. For example, runtime scheduler 314 can execute jump function 516 of to call second variant binary 504 in variant library 310. In such an example, runtime scheduler 314 can call second variant binary 504 by invoking and / or otherwise calling variant symbol “_halide_algox_gpu” in variant symbols 514 of jump table 412.
[0204] At block 1310, heterogeneous system 304 executes the workload using the one or more called variant binaries. For example, runtime scheduler 314 can execute The application function 518 is used to execute the workload using the second variant binary 504. For example, the runtime scheduler 314 can load the second variant binary 504 onto the GPU 322 by accessing the corresponding variant symbol from the jump table library 312. In other examples, the runtime scheduler 314 can execute the application function 518 to execute the first part of the workload using the first variant binary among the variant binaries 502, 504, 506, 508, 510 and execute the second part of the workload using the second variant binary among the variant binaries 502, 504, 506, 508, 510. For example, the runtime scheduler 314 can load the second variant binary 504 onto the GPU 322 by accessing the corresponding variant symbol from the jump table library 312 and load the fourth variant binary 508 onto the VPU 320.
[0205] At block 1312, the heterogeneous system 304 determines whether there is another workload to process. For example, the runtime scheduler 314 can determine that there is another workload to process, or in other examples, can determine that there is no additional workload to process. If at block 1312, the heterogeneous system 304 determines that there is another workload of interest to process, the control returns to block 1302 to obtain another workload to process. Otherwise, the machine-readable instructions 1124 return to the machine-readable instructions 1100 and end.
[0206] is a flowchart representing the example machine-readable instructions 1400 that can be executed to implement , , , and / or the variant generator 302 during the training phase. The machine-readable instructions 1400 start at block 1402, where the variant manager 402 ( ) obtains an algorithm. For example, the external device can correspond to the administrator device 202, and the algorithm can correspond to any algorithm among a set of arbitrary algorithms.
[0207] At block 1404, the variant generator 302 selects the processing element for which the algorithm is to be developed. For example, the variant generator 302 can be developing a variant for use on a heterogeneous system including four processing elements. In such an example, the variant manager 402 can select one of the multiple processing elements of the heterogeneous system 304 for which to generate a variant.
[0208] At block 1406, the variant generator 302 selects an aspect of the processing element as the target for the success function of the selected processing element. For example, the variant manager 402 may select the execution speed of the obtained algorithm on the FPGA 318 as the target.
[0209] At block 1408, the variant generator 302 generates a cost model for the selected processing element and the selected aspect as the target. For example, in an initial run, the cost model learner 404 ( ) may use the general weights for the DNN to generate one or more of the cost models 904. At block 1410, the variant generator 302 generates a schedule for implementing the obtained algorithm using the success function associated with the selected aspect on the selected processing element. For example, the compilation auto-scheduler 408 ( ) may use the success function associated with the selected aspect on the selected processing element to generate the Figure 9 schedules 912 for implementing the obtained algorithm.
[0210] At block 1412, the variant generator 302 compiles the variant. For example, the variant compiler 410 ( Figure 4 ) may compile the variant according to the schedule generated by the compilation auto-scheduler 408. In such an example, the compiled variant can be loaded as an executable file (e.g., a binary file) into the application compiled by the application compiler 414 ( Figure 4 ).
[0211] At block 1414, after subsequently executing the variant on a training system (e.g., a training heterogeneous system), the variant generator 302 collects the performance characteristics associated with the implementation of the variant on the selected processing element. For example, the feedback interface 416 ( Figure 4 ) may obtain the performance characteristics associated with the implementation of the variant on the selected processing element.
[0212] At block 1416, the variant generator 302 determines whether the execution of the variant meets the performance threshold. For example, the performance analyzer 418 ( Figure 4) It can be determined whether the execution of the variant meets and / or otherwise meets the performance threshold. If the execution of the variant does not meet the performance threshold (e.g., the desired performance level) (block 1416: No), the control returns to block 1408, where the collected performance characteristics are fed back into the cost model learner 404. If the execution of the variant meets the performance threshold (block 1416: Yes), then at block 1418, the variant generator 302 (e.g., the variant manager 402) determines whether there are any other aspects targeted for the success function for the selected processing element. If there are subsequent aspects targeted for the success function (block: 1418: Yes), the control returns to block 1406. If there are no subsequent aspects targeted for the success function (block: 1418: No), then at block 1420, the variant generator 302 (e.g., the variant manager 402) determines whether there are any other processing elements for which one or more variants are to be developed.
[0213] If there are subsequent processing elements (block: 1420: Yes), the control returns to block 1404. If there are no subsequent processing elements (block: 1420: No), then at block 1422, the variant generator 302 (e.g., the variant manager 402) determines whether there is an additional algorithm. If there is an additional algorithm (block: 1422: Yes), the control returns to block 1402. If there is no additional algorithm (block: 1422: No), then at block 1424, the variant generator 302 (e.g., the variant manager 402) outputs the corresponding trained DNN models (e.g., Figure 2 and / or Figure 3 the trained ML / AI model 214) corresponding to the respective processing elements of the heterogeneous system for use. For a algorithms targeting m different aspects to be executed on n processing elements, the variant generator 302 is capable of generating a*n*m DNNs to generate and analyze various cost models. For example, the variant manager 402 can output the trained DNN models to a database, another variant generator, and / or the heterogeneous system in the field or the system.
[0214] At block 1426, the variant generator 302 monitors the input data. For example, the feedback interface 416 can monitor the database, the heterogeneous system in the field or the system, or other data sources that can provide empirically collected performance characteristics.
[0215] At block 1428, the variant generator 302 (e.g., feedback interface 416) determines whether input data has been received and / or otherwise obtained. If the feedback interface 416 determines that the input data has not been received (block 1428: No), control returns to block 1426. If the feedback interface 416 determines that the input data has been received (block 1428: Yes), then at block 1430, the variant generator 302 (e.g., performance analyzer 418) identifies aspects of the heterogeneous system that are targeted based on the success function and performance characteristics of the system.
[0216] At block 1432, the variant generator 302 (e.g., performance analyzer 418) determines the difference between the desired performance (e.g., performance threshold) defined by the success function and the actual performance obtained during execution of the algorithm during the inference phase. In response to determining the difference at block 1432, Figure 14 the machine-readable instructions 1400 return to block 1408, where the empirical data is reinserted into the variant generator 302 (e.g., cost model learner 404) to adjust the cost model of individual processing elements based on context data (e.g., performance characteristics such as runtime load and environmental characteristics) associated with the system as a whole.
[0217] Figure 15 is a flowchart representing example machine-readable instructions 1500 that may be executed to implement Figure 3 、 Figure 4 、 Figure 6 、 Figure 9 and / or Figure 10 the variant generator 302 during the inference phase. Figure 15 The machine-readable instructions 1500 of
[0218] begin at block 1502, where the variant generator 302 (e.g., variant manager 402) obtains an algorithm from an external device. For example, the external device may correspond to a laptop computer of a program developer, user, administrator, etc. Figure 4 At block 1504, the variant generator 302 (e.g.,
[0219] the variant manager 402 of Figure 3targets the power consumption of execution on a GPU such as GPU 322.
[0220] At block 1508, the variant generator 302 (e.g., Figure 4 the cost model learner 404) uses the trained DNN model to generate at least one cost model of the algorithm for execution on at least one processing element of the heterogeneous system. At block 1510, the variant generator 302 (e.g., Figure 4 the compilation auto-scheduler 408) generates a schedule for implementing the obtained algorithm using a success function associated with a selected aspect on the selected processing element. At block 1512, the variant generator 302 (e.g., Figure 4 the variant compiler 410) compiles the variant according to the schedule generated by the compilation auto-scheduler 408.
[0221] At block 1514, the variant generator 302 (e.g., the variant compiler 410) adds the variant to the variant library of the application to be compiled. At block 1516, the variant generator 302 (e.g., the variant compiler 410) adds a variant symbol (e.g., a pointer) to Figure 4 and / Figure 5 the jump table 412 by transmitting the variant to the jump table 412, and the jump table 412 generates a corresponding symbol associated with the position of the variant in the variant library of the application to be compiled.
[0222] At block 1518, the variant generator 302 (e.g., the variant manager 402) determines whether there are any other aspects targeted for the success function for the selected processing element. If there are subsequent aspects targeted for the success function (block: 1518: yes), the control returns to block 1506. If there are no subsequent aspects targeted for the success function (block: 1518: no), then at block 1520, the variant generator 302 (e.g., the variant manager 402) determines whether there are any other processing elements for which one or more variants are to be developed. If there are subsequent processing elements (block: 1520: yes), the control returns to block 1504. If there are no subsequent processing elements (block: 1520: no), then at block 1522, the variant generator 302 (e.g., the jump table 412) adds the current state of the jump table 412 to the jump table library of the application to be compiled. In block 1524, the variant generator 302 (e.g., the application compiler 414) compiles the different variants for each processing element in the variant library, the variant symbols in the jump table library, and the runtime scheduler into an executable application.
[0223] At block 1526, a variant generator (e.g., variant manager 402) determines whether there are additional algorithms. If there are additional algorithms (block: 1526: yes), control returns to block 1502. If there are no additional algorithms (block: 1526: no), then Figure 15 the machine-readable instructions 1500 end.
[0224] Figure 16 is a flowchart of example machine-readable instructions 1600 that may be executed to implement Figure 3 and / or Figure 10 the executable file 308. The machine-readable instructions 1600 begin at block 1602, where a runtime scheduler 314 ( Figure 3 ) determines a system-wide success function for a heterogeneous system.
[0225] At block 1604, the runtime scheduler 314 executes an algorithm on the heterogeneous system according to a variant generated by a trained ML / AI model. At block 1606, the runtime scheduler 314 monitors performance characteristics of the heterogeneous system under load and environmental conditions.
[0226] At block 1608, the runtime scheduler 314 adjusts the configuration of the heterogeneous system to meet the system-wide success function. For example, based on the performance characteristics, the runtime scheduler 314 may offload a workload executed on the CPU 316 to the GPU 322. To this end, the runtime scheduler 314 is able to access a variant of a specific algorithm for the workload corresponding to the GPU 322, which is stored in the variant library 310. The runtime scheduler 314 is able to load the variant onto the GPU 322 by accessing a corresponding variant symbol from the jump table library 312.
[0227] At block 1610, the runtime scheduler 314 determines whether the heterogeneous system includes persistent storage. If the runtime scheduler 314 determines that the heterogeneous system does include persistent storage (block 1610: yes), then at block 1612, the runtime scheduler 314 periodically stores the monitored data on an executable file on the persistent storage (e.g., Figure 3 and / or Figure 5In the FAT binary file 309). After box 1612, control proceeds to box 1624. If the runtime scheduler 314 determines that the heterogeneous system does not include persistent storage (box 1610: No), then at box 1614, the runtime scheduler 314 determines whether the heterogeneous system includes flash storage. If the runtime scheduler 314 determines that the heterogeneous system does include flash storage (box 1614: Yes), then at box 1616, the runtime scheduler 314 periodically stores the monitored data in an executable file (e.g., FAT binary file 309) on the flash storage. After box 1616, control proceeds to box 1624. If the runtime scheduler 314 determines that the heterogeneous system does not include flash storage (box 1614: No), then at box 1618, the runtime scheduler 314 determines whether the heterogeneous system includes persistent storage. If the runtime scheduler 314 determines that the heterogeneous system does include persistent BIOS (box 1618: Yes), then at box 1620, the runtime scheduler 314 periodically stores the monitored data in an executable file (e.g., FAT binary file 309) on the persistent BIOS. After box 1620, control proceeds to box 1624. If the runtime scheduler 314 determines that the heterogeneous system does not include persistent storage (box 1618: No), then at box 1622, the runtime scheduler 314 transfers the monitored data (e.g., empirical performance characteristics) to external storage (e.g., Figure 2 database 208).
[0228] At box 1624, the runtime scheduler 314 determines whether the algorithm has finished executing. If the runtime scheduler 314 determines that the algorithm has not finished executing (box 1624: No), then control returns to box 1606. If the runtime scheduler 314 determines that the algorithm has finished executing (box 1624: Yes), then at box 1626, the runtime scheduler 314 transfers the monitored data (e.g., empirical performance characteristics) to external devices (e.g., database 208, variant generator 302, etc.). At box 1628, the runtime scheduler 314 determines whether there are additional algorithms. If there are additional algorithms (box: 1628: Yes), then control returns to box 1602. If there are no additional algorithms (box: 1628: No), then Figure 16 the machine-readable instructions 1600 end.
[0229] Figure 17 is constructed to execute Figure 11 , Figure 12 , Figure 14 and / or Figure 15 instructions in order to implement Figure 3 , Figure 4 , Figure 6 and / or Figure 10 variant application 305 and / or more generally,Figure 3 Block diagram of an example processor platform 1700 of the third software adjustment system 301. The processor platform 1700 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smart phone, a tablet computer such as an iPad TM or any other type of computing device).
[0230] The illustrated example of the processor platform 1700 includes a processor 1712. The illustrated example of the processor 1712 is hardware. For example, the processor 1712 can be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs or controllers from any desired family or manufacturer. The hardware processor can be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 1712 implements Figure 3 , Figure 4 , Figure 6 , Figure 9 and / or Figure 10 example variant generator 302, and Figure 4 example variant manager 402, example cost model learner 404, example weight store 406, example compilation auto-scheduler 408, example variant compiler 410, example jump table 412, example application compiler 414, example feedback interface 416 and example profiler 418. In this example, the processor 1712 implements Figure 3 , Figure 4 , Figure 6 and / or Figure 10 example code converter 303, and Figure 6 example code interface 602, example code lifter 604, example DSL generator 606, example metadata generator 608 and / or example code replacer 610. In this example, the processor 1712 implements Figure 3 , Figure 4 , Figure 6 and / or Figure 10 variant application 305.
[0231] The illustrated example of the processor 1712 includes local memory 1713 (e.g., cache). The illustrated example of the processor 1712 communicates with main memory including volatile memory 1714 and non-volatile memory 1716 via a bus 1718. The volatile memory 1714 can be implemented by SDRAM, DRAM, and / or any other type of random access memory device. The non-volatile memory 1716 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 1714, main memory 1716 is controlled by a memory controller.
[0232] The illustrated example's processor platform 1700 also includes interface circuitry 1720. The interface circuitry 1720 can be implemented by any type of interface standard, such as, an Ethernet interface, a Universal Serial Bus (USB), interface, a Near Field Communication (NFC) interface, and / or a Peripheral Component Interconnect Express (PCI express) interface. In this example, the interface circuitry 1720 implements Figure 4 communication bus 420 and Figure 6 communication bus 618. Alternatively, the interface circuitry 1720 can implement Figure 6 code interface 602.
[0233] In the illustrated example, one or more input devices 1722 are connected to the interface circuitry 1720. The (multiple) input devices 1722 permit a user to input data and / or commands into the processor 1712. The (multiple) input devices can be implemented by, for example, audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, trackpads, trackballs, isopoint mice, and / or voice recognition systems.
[0234] One or more output devices 1724 are also connected to the interface circuitry 1720 of the illustrated example. The output devices 1724 can be implemented by, for example, display devices (e.g., light-emitting diodes (LEDs), organic light-emitting diodes (OLEDs), liquid crystal displays (LCDs), cathode ray tube displays (CRTs), in-plane switching (IPS) displays, touchscreens, etc.), haptic output devices, printers, and / or speakers. Thus, the interface circuitry 1720 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0235] The interface circuitry 1720 of the illustrated example also includes communication devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate the exchange of data with external machines (e.g., any kind of computing device) via a network 1726. The communication can be via, for example, an Ethernet connection, a Digital Subscriber Line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-sight wireless system, a cellular phone system, etc.
[0236] The illustrated example's processor platform 1700 also includes one or more mass storage devices 1728 for storing software and / or data. Examples of such mass storage devices 1728 include floppy disk drives, hard disk drives, compact disc drives, Blu-ray disc drives, redundant array of independent disks (RAID) systems, and DVD drives. In this example, one or more mass storage devices 1728 implement Figure 6 database 612, which includes Figure 6and / or Figure 7 the annotated code 614 and the unannotated code 616 of
[0237] Figure 11 , Figure 12 , Figure 14 and / or Figure 15 the machine-executable instructions 1732 of
[0238] Figure 18 is configured to execute Figure 13 and / or Figure 16 the instructions of Figure 3 and / or Figure 10 to implement the example processor platform 1800 of the executable file 308. The processor platform 1800 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smart phone, a tablet computer such as an iPad TM or any other type of computing device).
[0239] The illustrated example of the processor platform 1800 includes a processor 1812. The illustrated example of the processor 1812 is hardware. For example, the processor 1812 can be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor can be a semiconductor-based (e.g., silicon-based) device. Additionally, the processor platform 1800 can include additional processing elements such as Figure 3 the example CPU 316, the example FPGA 318, the example VPU 320, and / or the example GPU 322 of
[0240] The illustrated example of the processor 1812 includes local memory 1813 (e.g., a cache). In this example, the local memory 1813 and / or more generally, the processor 1812 includes and / or otherwise implements Figure 3 the example executable file 308, the example variant library 310, the example jump table library 312, and the example runtime scheduler 314. The illustrated example of the processor 1812 communicates with a main memory including volatile memory 1814 and non-volatile memory 1816 via a bus 1818. The volatile memory 1814 can be SDRAM, DRAM, and / or any other type of random access memory device. The non-volatile memory 1816 can be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 1814 and the main memory 1816 is controlled by a memory controller.
[0241] The illustrated example of the processor platform 1800 also includes interface circuitry 1820. The interface circuitry 1820 can be implemented by any type of interface standard, such as, an Ethernet interface, a Universal Serial Bus (USB), interface, a Near Field Communication (NFC) interface, and / or a Peripheral Component Interconnect Express (PCI express) interface.
[0242] In the illustrated example, one or more input devices 1822 are connected to the interface circuitry 1820. The input device(s) 1822 permit a user to input data and / or commands into the processor 1812. The input device(s) can be implemented by, for example, an audio sensor, a microphone, a camera (still or video), a keyboard, a button, a mouse, a touch screen, a track pad, a track ball, an isotropic mouse, and / or a voice recognition system.
[0243] One or more output devices 1824 are also connected to the interface circuitry 1820 of the illustrated example. The output device 1824 can be implemented, for example, by a display device (e.g., an LED, an OLED, an LCD, a CRT, an IPS display, a touch screen, etc.), a haptic output device, a printer, and / or a speaker. Thus, the interface circuitry 1820 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0244] The illustrated example of the interface circuitry 1820 also includes communication devices, such as a transmitter, a receiver, a transceiver, a modem, a residential gateway, a wireless access point, and / or a network interface, to facilitate the exchange of data with external machines (e.g., any kind of computing device) via a network 1826. The communication can be via, for example, an Ethernet connection, a Digital Subscriber Line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-sight wireless system, a cellular telephone system, etc.
[0245] The illustrated example of the processor platform 1800 also includes one or more mass storage devices 1828 for storing software and / or data. Examples of such mass storage devices 1828 include a floppy disk drive, a hard disk drive (HDD), a CD drive, a Blu-ray drive, a RAID system, and a DVD drive.
[0246] Figure 13 and / or Figure 16The machine-executable instructions 1832 can be stored in the mass storage device 1828, the volatile memory 1814, the non-volatile memory 1816, and / or on a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0247] In view of the foregoing, it will be understood that example methods, apparatuses, and articles have been disclosed that do not rely solely on the developer's knowledge of the theoretical understanding of processing elements, algorithm transformations, and other scheduling techniques, and other pitfalls of certain methods for compiling schedules. The examples disclosed herein use algorithms based on relatively simple semantics for representing performance- and / or power-intensive code blocks using conventional imperative programming language idioms. In addition to the mere algorithm definition, such semantics can include metadata regarding the developer's intent. The examples disclosed herein use proven promotion techniques to transform imperative programming language code into a promoted intermediate representation or formal description of the program intent. The examples disclosed herein transform the formal description of the program intent into a domain-specific language representation that can be used to generate variants of the original algorithm for deployment across one or more of a plurality of processing elements of a heterogeneous system.
[0248] The examples disclosed herein collect empirical performance characteristics and the differences between the desired performance (e.g., success function) and the actual performance obtained. Additionally, the examples disclosed herein allow for the continuous and automatic improvement of the performance of a heterogeneous system without developer intervention. The disclosed methods, apparatuses, and articles improve the efficiency of using a computing device at least by reducing the power consumption of an algorithm executed on the computing device by deploying workloads based on the use of processing elements, by increasing the execution speed of the algorithm on the computing device by deploying the workload to processing elements that may be unutilized or underutilized, and by increasing the utilization rate of the various processing elements of the computing system by allocating the workload to available processing elements. The disclosed methods, devices, and articles accordingly relate to one or more improvements in computer functionality.
[0249] Example methods, apparatuses, systems, and articles for intentional programming for heterogeneous systems are disclosed herein. Other examples and combinations thereof include the following:
[0250] Example 1 includes an apparatus for intentional programming for a heterogeneous system, the apparatus including: a code lifter for: identifying annotated code corresponding to an algorithm to be executed on the heterogeneous system based on an identifier associated with the annotated code, the annotated code being in a first representation; and converting the annotated code in the first representation to intermediate code by identifying intermediate code in a second representation as having a first algorithmic intent corresponding to a second algorithmic intent of the annotated code; a domain-specific language (DSL) generator for converting the intermediate code in the second representation to DSL code in a third representation corresponding to a DSL representation when the first algorithmic intent matches the second algorithmic intent; and a code replacer for invoking a compiler to generate an executable file including variant binaries based on the DSL code, each of the variant binaries for invoking a corresponding one of the processing elements of the heterogeneous system to execute the algorithm.
[0251] Example 2 includes the apparatus of Example 1, wherein the identifier is an imperative programming language compiler extension and the first representation is an imperative programming language representation.
[0252] Example 3 includes the apparatus of Example 1, wherein the executable file includes a runtime scheduler and the code replacer is for replacing the annotated code in the application code with a function call to the runtime scheduler to invoke a variant binary to be loaded onto one or more of the processing elements of the heterogeneous system to execute the workload.
[0253] Example 4 includes the apparatus of Example 1, further including: a metadata generator for generating scheduling metadata corresponding to a power profile of the heterogeneous system; a compiled auto-scheduler for generating a schedule based on the scheduling metadata for the processing elements of the heterogeneous system, the processing elements including at least a first processing element and a second processing element; a variant compiler for compiling variant binaries based on the schedule, each of the variant binaries being associated with the algorithm in the DSL, the variant binaries including a first variant binary corresponding to the first processing element and a second variant binary corresponding to the second processing element; and an application compiler for compiling the executable file to include a runtime scheduler to select one or more of the variant binaries based on the schedule to execute the workload.
[0254] Example 5 includes the apparatus of Example 1, and further includes: a feedback interface for obtaining performance characteristics of the heterogeneous system from the executable file, the performance characteristics being associated with the processing elements of the heterogeneous system that execute the workload during a first runtime, the executable file being executed according to a function that specifies successful execution of the executable file on the heterogeneous system, the processing elements including a first processing element and a second processing element; and a performance analyzer for: determining a performance delta based on the performance characteristics and the function; and before a second runtime, using a machine learning model to adjust a cost model of the first processing element based on the performance delta.
[0255] Example 6 includes the apparatus of Example 5, wherein the cost model is a first cost model, and the apparatus further includes a cost model learner that, before the second runtime, adjusts a second cost model of the second processing element based on the performance delta by using a neural network.
[0256] Example 7 includes the apparatus of Example 6, wherein the performance analyzer is configured to determine the performance delta by determining a difference between the performance obtained during the first runtime and the performance defined by the function that specifies successful execution of the executable file on the heterogeneous system.
[0257] Example 8 includes a non-transitory computer-readable storage medium that includes instructions which, when executed, cause a machine to at least: identify, based on an identifier, annotated code corresponding to an algorithm to be executed on a heterogeneous system, the identifier being associated with the annotated code, the annotated code being in a first representation; convert the annotated code in the first representation to intermediate code by identifying intermediate code in a second representation as having a first algorithmic intent corresponding to a second algorithmic intent of the annotated code; when the first algorithmic intent matches the second algorithmic intent, convert the intermediate code in the second representation to domain-specific language (DSL) code in a third representation, the third representation corresponding to a DSL representation; and invoke a compiler to generate an executable file including variant binaries based on the DSL code, each of the variant binaries being for invoking a corresponding one of the processing elements of the heterogeneous system to execute the algorithm.
[0258] Example 9 includes the non-transitory computer-readable storage medium of Example 8, wherein the identifier is an imperative programming language compiler extension, and the first representation is an imperative programming language representation.
[0259] Example 10 includes the non-transitory computer-readable storage medium of Example 8, wherein the executable file includes a runtime scheduler, and when the instructions are executed, cause the machine to replace the annotated code in the application code with a function call to the runtime scheduler to call a variant binary to be loaded onto the processing element of the heterogeneous system to execute the workload.
[0260] Example 11 includes the non-transitory computer-readable storage medium of Example 8, wherein when the instructions are executed, cause the machine to: generate scheduling metadata corresponding to the power profile of the heterogeneous system; generate a schedule based on the scheduling metadata for the processing elements of the heterogeneous system, the processing elements including at least a first processing element and a second processing element; compile variant binaries based on the schedule, each of the variant binaries being associated with the algorithm using the DSL, the variant binaries including a first variant binary corresponding to the first processing element and a second variant binary corresponding to the second processing element; and compile the executable file to include a runtime scheduler to select one or more of the variant binaries based on the schedule to execute the workload.
[0261] Example 12 includes the non-transitory computer-readable storage medium of Example 8, wherein when the instructions are executed, cause the machine to: obtain performance characteristics of the heterogeneous system from the executable file, the performance characteristics being associated with the processing element of the heterogeneous system that executes the workload at a first runtime, the executable file being executed according to a function that specifies successful execution of the executable file on the heterogeneous system, the processing element including a first processing element and a second processing element; determine a performance delta based on the performance characteristics and the function; and before a second runtime, use a machine learning model to adjust a cost model of the first processing element based on the performance delta.
[0262] Example 13 includes the non-transitory computer-readable storage medium of Example 12, wherein the cost model is a first cost model, and wherein when the instructions are executed, cause the machine to, before the second runtime, use a neural network to adjust a second cost model of the second processing element based on the performance delta.
[0263] Example 14 includes the non-transitory computer-readable storage medium of Example 13, wherein when the instructions are executed, cause the machine to determine the performance delta by determining a difference between the performance obtained at the first runtime and the performance defined by the function that specifies successful execution of the executable file on the heterogeneous system.
[0264] Example 15 includes a method for intentional programming for a heterogeneous system, the method including: identifying, based on an identifier, annotated code corresponding to an algorithm to be executed on the heterogeneous system, the identifier being associated with the annotated code, the annotated code being in a first representation; converting the annotated code in the first representation to intermediate code by identifying intermediate code in a second representation as having a first algorithmic intent corresponding to a second algorithmic intent of the annotated code; converting the intermediate code in the second representation to domain-specific language (DSL) code in a third representation when the first algorithmic intent matches the second algorithmic intent, the third representation corresponding to a DSL representation; and invoking a compiler to generate an executable file including variant binaries based on the DSL code, each of the variant binaries for invoking a corresponding one of the processing elements of the heterogeneous system to execute the algorithm.
[0265] Example 16 includes the method of Example 15, wherein the identifier is an imperative programming language compiler extension, and the first representation is an imperative programming language representation.
[0266] Example 17 includes the method of Example 15, wherein the executable file includes a runtime scheduler, and the method further includes: replacing the annotated code in the application code with a function call to the runtime scheduler to invoke a variant binary to be loaded onto one of the processing elements of the heterogeneous system to execute the workload.
[0267] Example 18 includes the method of Example 15, further including: generating scheduling metadata corresponding to a power profile of the heterogeneous system; generating a schedule based on the scheduling metadata for the processing elements of the heterogeneous system, the processing elements including at least a first processing element and a second processing element; compiling the variant binaries based on the schedule, each of the variant binaries being associated with the algorithm in the DSL, the variant binaries including a first variant binary corresponding to the first processing element and a second variant binary corresponding to the second processing element; and compiling the executable file to include a runtime scheduler to select one or more of the variant binaries based on the schedule to execute the workload.
[0268] Example 19 includes the method of Example 15 and further includes: obtaining performance characteristics of the heterogeneous system from the executable file, the performance characteristics being associated with the processing elements of the heterogeneous system that execute the workload at a first runtime, the executable file being executed according to a function that specifies successful execution of the executable file on the heterogeneous system, the processing elements including a first processing element and a second processing element; determining a performance delta based on the performance characteristics and the function; and before a second runtime, adjusting a cost model of the first processing element based on the performance delta using a machine learning model.
[0269] Example 20 includes the method of Example 19, wherein the cost model is a first cost model, and the method further includes adjusting a second cost model of the second processing element based on the performance delta using a neural network before the second runtime.
[0270] Example 21 includes the method of Example 20, wherein determining the performance delta includes determining a difference between the performance obtained at the first runtime and the performance defined by the function that specifies successful execution of the executable file on the heterogeneous system.
[0271] Example 22 includes an apparatus for intentional programming for a heterogeneous system, the apparatus including: means for identifying annotated code corresponding to an algorithm to be executed on the heterogeneous system based on an identifier associated with the annotated code, the annotated code being in a first representation; means for converting the annotated code in the first representation to intermediate code by identifying intermediate code in a second representation as having a first algorithmic intent corresponding to a second algorithmic intent of the annotated code; means for converting the intermediate code in the second representation to domain-specific language (DSL) code in a third representation corresponding to a DSL representation when the first algorithmic intent matches the second algorithmic intent; and means for invoking a compiler to generate an executable file including variant binaries based on the DSL code, each of the variant binaries for invoking a corresponding one of the processing elements of the heterogeneous system to execute the algorithm.
[0272] Example 23 includes the apparatus of Example 22, wherein the identifier is an imperative programming language compiler extension and the first representation is an imperative programming language representation.
[0273] Example 24 includes the apparatus of Example 22, wherein the executable file includes a runtime scheduler, and the apparatus further includes: means for replacing the annotated code in the application code with a function call to the runtime scheduler so as to call a variant binary to be loaded onto one of the processing elements in the heterogeneous system to execute the workload.
[0274] Example 25 includes the apparatus of Example 22, and further includes: a first means for generating scheduling metadata corresponding to a power profile of the heterogeneous system; a second means for generating a schedule based on the scheduling metadata for the processing elements of the heterogeneous system, the processing elements including at least a first processing element and a second processing element; a first means for compiling variant binaries based on the schedule, each of the variant binaries being associated with the algorithm using the DSL, the variant binaries including a first variant binary corresponding to the first processing element and a second variant binary corresponding to the second processing element; and a second means for compiling the executable file to include a runtime scheduler so as to select one or more of the variant binaries based on the schedule to execute the workload.
[0275] Example 26 includes the apparatus of Example 22, and further includes: means for obtaining performance characteristics of the heterogeneous system from the executable file, the performance characteristics being associated with the processing element that executes the workload at a first runtime, the executable file being executed according to a function that specifies successful execution of the executable file on the heterogeneous system, the processing element including a first processing element and a second processing element; means for determining a performance delta based on the performance characteristics and the function; and means for adjusting a cost model of the first processing element based on the performance delta using a machine learning model before a second runtime.
[0276] Example 27 includes the apparatus of Example 26, wherein the cost model is a first cost model, and the apparatus further includes means for adjusting a second cost model of the second processing element based on the performance delta using a neural network before the second runtime.
[0277] Example 28 includes the apparatus of Example 26, wherein the means for determining is configured to determine the performance delta by determining a difference between the performance obtained at the first runtime and the performance defined by the function that specifies successful execution of the executable file on the heterogeneous system.
[0278] Although certain example systems, methods, devices, and articles are disclosed herein, the scope covered by this patent is not limited thereto. Instead, this patent covers all systems, methods, devices, and articles that fall within the scope of the claims of this patent.
Claims
1. At least one non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that cause a processor circuit to at least perform the following operations: Identify a first code block having a first algorithmic intent based on a second code block having a second algorithmic intent, the second algorithmic intent corresponding to the first algorithmic intent; Translate the first code block into an executable domain-specific language code using a machine learning model; and Output the executable domain-specific language code.
2. The at least one non-transitory computer-readable storage medium according to claim 1, wherein, The machine learning model is a first machine learning model, and the instructions cause the processor circuit to use a second machine learning model to determine that the second algorithmic intent corresponds to the first algorithmic intent.
3. The at least one non-transitory computer-readable storage medium according to claim 1, wherein, The instructions cause the processor circuit to: Determine that the first algorithmic intent corresponds to a first operation using a first representation, the first operation being for generating an output based on an input; and Determine that the second algorithmic intent corresponds to a second operation using a second representation, the second operation being for generating the output based on the input.
4. The at least one non-transitory computer-readable storage medium according to claim 1, wherein, The instructions cause the processor circuit to: Determine that the second code block includes a loop; Determine that the first code block is a template code block; and Translate the first code block into the executable domain-specific language code based on replacement of the loop with the template code block.
5. The at least one non-transitory computer-readable storage medium according to claim 1, wherein, The processor circuit is a first processor circuit, and the instructions cause the first processor circuit to: Generate a first schedule for a second processor circuit to execute a first portion of a workload based on the executable domain-specific language code; Generate a second schedule for a third processor circuit to execute a second portion of the workload based on the executable domain-specific language code; Cause the second processor circuit to execute the first portion based on the first schedule; And Cause the third processor circuit to execute the second portion based on the second schedule.
6. The at least one non-transitory computer-readable storage medium of claim 5, wherein the instructions cause the first processor circuit to determine that the second processor circuit is to execute the first portion in parallel with the third processor circuit's execution of the second portion.
7. The at least one non-transitory computer-readable storage medium according to claim 5, wherein, The third processor circuit is a graphics processing unit, and the instructions cause the first processor circuit to: Determine that the second processor circuit is to execute the first portion and the second portion; And Based on a determination that the utilization of the second processor circuit meets a threshold, cause the second processor circuit to migrate the second portion to the graphics processing unit.
8. The at least one non-transitory computer-readable storage medium of claim 5, wherein at least one of the second processor circuit or the third processor circuit is one of the following: a special instruction set processor, a central processing unit, a digital signal processor, a field programmable gate array, a graphics processing unit, a physics processing unit, or a vision processing unit.
9. An apparatus, comprising: A memory; Instructions; And A processor circuitry that is configured to execute the instructions to perform the following operations: Determining that the first code block has the first algorithmic intent based on determining that a second algorithmic intent of a second code block corresponds to the first algorithmic intent, the determining being based on an output from a machine learning model; Converting the first code block into an executable domain-specific language code; And Generating an executable file based on the executable domain-specific language code.
10. The device according to claim 9, wherein, The output is a first output, the machine learning model is a first machine learning model, and the processor circuitry is operative to translate the first code block into the executable domain-specific language code based on a second output from a second machine learning model.
11. The device according to claim 9, wherein, The output is a first output, and the processor circuitry is operative to: Determine that the first algorithmic intent corresponds to a first function in a first representation, the first function being operative to generate a second output based on an input; and Determine that the second algorithmic intent corresponds to a second function in a second representation, the second function being operative to generate the second output based on the input.
12. The device according to claim 9, wherein The processor circuitry is operative to: Identify a loop in the second code block, the loop being in a first representation; Identify a template code block in the first code block, the template code block being in a second representation; And Translate the first code block into the executable domain-specific language code based on a translation of the first representation into the executable domain-specific language code.
13. The device according to claim 9, wherein, The processor circuitry is a first processor circuitry, and the first processor circuitry is operative to: Compile a first schedule for a second processor circuitry to execute a first portion of a workload based on the executable domain-specific language code; Compile a second schedule for a third processor circuitry to execute a second portion of the workload based on the executable domain-specific language code; Invoke the second processor circuitry to execute the first portion based on the first schedule; And Invoke the third processor circuitry to execute the second portion based on the second schedule.
14. The device according to claim 13, wherein, The first processor circuitry is operative to determine that the second processor circuitry is to execute the first portion in parallel with the third processor circuitry's execution of the second portion.
15. The device according to claim 13, wherein, The third processor circuitry is a graphics processing unit, and the first processor circuitry is operative to: Determine that the second processor circuitry is to execute the first portion and the second portion; and Based on a determination that utilization of the second processor circuitry meets a threshold, instruct the second processor circuitry to migrate the second portion to the graphics processing unit.
16. The apparatus of claim 13, wherein at least one of the second processor circuitry or the third processor circuitry is one of the following: a special instruction set processor, a central processing unit, a digital signal processor, a field programmable gate array, a graphics processing unit, a physics processing unit, or a vision processing unit.
17. A system, comprising: A mass storage device for storing a first code block; And A server for generating an executable file, the executable file being operative to: Identify a first code block associated with a first algorithmic intent based on a second code block associated with a second algorithmic intent, the second algorithmic intent corresponding to the first algorithmic intent; Determine that the second code block includes a loop; Determine that the first code block is a template code block; And Convert the first code block into an executable domain-specific language code based on adapting a portion of the loop to the template code block; And Output the executable domain-specific language code.
18. The system according to claim 17, wherein, The server is configured to generate the executable file to execute a machine learning model to determine that the second algorithmic intent corresponds to the first algorithmic intent based on an output of the machine learning model.
19. The system according to claim 17, wherein, The server is configured to generate the executable file to execute a machine learning model to translate the first code block into the executable domain-specific language code based on an output of the machine learning model.
20. The system according to claim 17, wherein, The server is configured to generate the executable file for: Determine that the first algorithmic intent corresponds to a first operation in a first representation, the first operation being for generating an output based on an input; and Determine that the second algorithmic intent corresponds to a second operation in a second representation, the second operation being for generating the output based on the input.
21. The system according to claim 17, wherein, The server is configured to generate the executable file for: Generate a first schedule for a first processor circuit to execute a first portion of a workload based on the executable domain-specific language code; Generate a second schedule for a second processor circuit to execute a second portion of the workload based on the executable domain-specific language code; Cause the first processor circuit to execute the first portion based on the first schedule; And Cause the second processor circuit to execute the second portion based on the second schedule.
22. The system according to claim 21, wherein, The server is configured to generate the executable file to determine that the first processor circuit is to execute the first portion in parallel with the second processor circuit's execution of the second portion.
23. The system according to claim 21, wherein, The second processor circuit is a graphics processing unit, and the server is configured to generate the executable file for: Determine that the first processor circuit is to execute the first portion and the second portion; And Based on a determination that a utilization rate of the first processor circuit meets a threshold, cause the second processor circuit to migrate the second portion to the graphics processing unit.
24. The system of claim 21, wherein at least one of the first processor circuit or the second processor circuit is one of the following: a special instruction set processor, a central processing unit, a digital signal processor, a field programmable gate array, a graphics processing unit, a physics processing unit, or a vision processing unit.