Apparatus, system and method for compiling code for processor
By introducing VMP OpenCL compiler and LLVM compilation scheme into the compiler, the intermediate code is optimized to support the instruction set of vector processors, which solves the problem that existing compilers are difficult to efficiently compile source code, and achieves efficient vector processing performance.
Patent Information
- Application Number
- CN202380072160.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-27
AI Technical Summary
Existing compilers have difficulty compiling source code efficiently into object code suitable for vector processors, resulting in inefficient processing of functional requirements.
By introducing a vector microcode processor (VMP) open computing language (OpenCL) compiler into the compiler, combined with a low-level virtual machine (LLVM) compilation scheme, the intermediate code is optimized to support the instruction set of vector processors.
It realizes efficient compilation of source code into object code suitable for vector processors, improving the performance and efficiency of processing functional requirements.
Smart Images

Figure CN120051761A_ABST
Abstract
Description
[0001] Cross-reference
[0002] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 415,307, filed on October 12, 2022, entitled "APPARATUS, SYSTEM, AND METHOD OF COMPILING CODE FOR A PROCESSOR", the entire disclosure of which is incorporated herein by reference. BACKGROUND OF THE INVENTION
[0003] A compiler may be configured to compile source code into object code configured for execution by a processor.
[0004] There is a need to provide technical solutions to support efficient processing functionality. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity of presentation. Additionally, reference numerals may be repeated in the figures to indicate corresponding or similar elements. The drawings are listed below.
[0006] Figure 1 is a schematic block diagram illustration of a system in accordance with some exemplary aspects.
[0007] Figure 2 is a schematic illustration of a compiler in accordance with some exemplary aspects.
[0008] Figure 3 is a schematic illustration of a vector processor in accordance with some exemplary aspects.
[0009] Figure 4 is a schematic flowchart illustration of a method of compiling code for a processor in accordance with some exemplary aspects.
[0010] Figure 5 is a schematic flowchart illustration of a method of compiling code for a processor in accordance with some exemplary aspects.
[0011] Figure 6 is a schematic illustration of a product in accordance with some exemplary aspects. DETAILED DESCRIPTION
[0012] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of some aspects. However, one of ordinary skill in the art will understand that some aspects may be practiced without these specific details. In other instances, well-known methods, procedures, components, units, and / or circuits have not been described in detail so as not to obscure the discussion.
[0013] The following presents a detailed description of some parts in terms of algorithms and symbolic representations of operations on data bits or binary digital signals within a computer memory. These algorithmic descriptions and representations can be techniques used by those skilled in the data processing art to convey the essence of their work to other technicians in the field.
[0014] An algorithm is here and generally considered to be a self-consistent sequence of actions or operations leading to a desired result. These include physical manipulations of physical quantities. Usually, but not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient, mainly for general reasons, to sometimes refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc. However, it should be understood that all such terms and similar terms should be associated with appropriate physical quantities and are merely convenient labels applied to these quantities.
[0015] Discussions herein that utilize terms such as, for example, "process," "compute," "calculate," "determine," "establish," "analyze," "verify," etc., may refer to operations and / or processes of a computer, computing platform, computing system, or other electronic computing device that manipulates data represented as physical (e.g., electrical) quantities within the registers and / or memory of the computer and / or converts such data into other data similarly represented as physical quantities within the registers and / or memory of the computer or other information storage media that can store instructions for performing the operations and / or processes.
[0016] As used herein, the terms "plurality" and "varieties" include, for example, "a plurality" or "two or more." For example, "a plurality of items" includes two or more items.
[0017] References to "an aspect," "one aspect," "exemplary aspect," "various aspects," etc., indicate that the aspect so described may include a particular feature, structure, or characteristic, but not every aspect necessarily includes the particular feature, structure, or characteristic. Moreover, repeated use of the phrase "in one aspect" does not necessarily refer to the same aspect, although it may.
[0018] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc., to describe a common object merely indicates different instances of the same object being referred to and is not intended to imply that the objects so described must be in a given sequence in time, space, rank, or in any other way.
[0019] For example, some aspects may capture the form of entirely hardware aspects, entirely software aspects, or aspects that include both hardware and software elements. Some aspects may be implemented in software, including but not limited to firmware, resident software, microcode, etc.
[0020] In addition, some aspects may capture the form of a computer program product that is accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For example, a computer-usable or computer-readable medium may be or may include any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0021] In some exemplary aspects, the medium may be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium.
[0022] In some exemplary aspects, a data processing system suitable for storing and / or executing program code may include, for example, at least one processor directly or indirectly coupled to memory elements through a system bus. The memory elements may include, for example, local memory, mass storage devices, and cache memory employed during the actual execution of the program code, and the cache memory may provide temporary storage of at least some program code to reduce the number of times code must be retrieved from the mass storage device during execution.
[0023] In some exemplary aspects, input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) may be directly or indirectly coupled to the system through an intermediate I / O controller. In some exemplary aspects, a network adapter may be coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices, for example, through an intermediate private or public network. In some exemplary aspects, modems, cable modems, and Ethernet cards are exemplary instances of network adapter types. Other suitable components may be used.
[0024] Some aspects may be used in conjunction with various devices and systems, such as, for example, computing devices, computers, mobile computers, non-mobile computers, server computers, etc.
[0025] As used herein, the term "circuitry" may refer to, be part of, or include an application specific integrated circuit (ASIC), an integrated circuit, an electronic circuit, a processor (shared, dedicated, or grouped), and / or a memory (shared, dedicated, or grouped) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable hardware components that provide the described functionality. In some aspects, some of the functionality associated with the circuitry may be implemented by one or more software or firmware modules. In some aspects, the circuitry may include logic that is at least partially operable in hardware.
[0026] The term "logic" may refer to, for example, computing logic embedded in circuitry of a computing device and / or stored in a memory of a computing device. For example, the logic may be accessed by a processor of the computing device to execute the computing logic for performing computing functions and / or operations. In one instance, the logic may be embedded in various types of memories and / or firmware, such as silicon blocks of various chips and / or processors. The logic may be included in and / or implemented as part of various circuitry, such as processor circuitry, control circuitry, and / or the like. In one instance, the logic may be embedded in volatile and / or non-volatile memories, including random access memory, read-only memory, programmable memory, magnetic memory, flash memory, persistent memory, and the like. The logic may be executed by one or more processors using a memory (e.g., registers, caches, buffers, and / or the like) coupled to the one or more processors, for example, to execute the logic as needed.
[0027] Now refer to Figure 1 , which schematically illustrates a block diagram of a system 100 according to some exemplary aspects.
[0028] As Figure 1 shown, in some exemplary aspects, the system 100 may include a computing device 102.
[0029] In some exemplary aspects, the device 102 may be implemented using suitable hardware components and / or software components, such as a processor, a controller, a memory unit, a storage unit, an input unit, an output unit, a communication unit, an operating system, an application, and the like.
[0030] In some exemplary aspects, the device 102 may include, for example, a computer, a mobile computing device, a non-mobile computing device, a laptop computer, a notebook computer, a tablet computer, a handheld computer, a personal computer (PC), and the like.
[0031] In some exemplary aspects, device 102 may include, for example, one or more of the following: a processor 191, an input unit 192, an output unit 193, a memory unit 194, and / or a storage unit 195. Device 102 may optionally include other suitable hardware components and / or software components. In some exemplary aspects, some or all of the components of one or more of the devices in device 102 may be enclosed in a common housing or package and may be interconnected or operably associated using one or more wired or wireless links. In other aspects, the components of one or more of the devices in device 102 may be distributed among multiple or separate devices.
[0032] In some exemplary aspects, processor 191 may include, for example, a central processing unit (CPU), a digital signal processor (DSP), one or more processor cores, a single-core processor, a dual-core processor, a multi-core processor, a microprocessor, a host processor, a controller, multiple processors or controllers, a chip, a microchip, one or more circuits, circuitry, a logic unit, an integrated circuit (IC), an application-specific IC (ASIC), or any other suitable general-purpose or specific processor or controller. Processor 191 may execute, for example, the instructions of an operating system (OS) of device 102 and / or the instructions of one or more suitable applications.
[0033] In some exemplary aspects, input unit 192 may include, for example, a keyboard, a keypad, a mouse, a touch screen, a touchpad, a trackball, a stylus, a microphone, or other suitable pointing or input devices. Output unit 193 may include, for example, a monitor, a screen, a touch screen, a flat panel display, a light-emitting diode (LED) display unit, a liquid crystal display (LCD) display unit, a plasma display unit, one or more audio speakers or headphones, or other suitable output devices.
[0034] In some exemplary aspects, memory unit 194 includes, for example, random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous DRAM (SD-RAM), flash memory, volatile memory, non-volatile memory, a cache, a buffer, a short-term memory unit, a long-term memory unit, or other suitable memory units. Storage unit 195 may include, for example, a hard disk drive, a solid-state drive (SSD), or other suitable removable or non-removable storage units. Memory unit 194 and / or storage unit 195 may store, for example, data processed by device 102.
[0035] In some exemplary aspects, device 102 may be configured to communicate with one or more other devices via at least one network 103 (e.g., a wireless and / or wired network).
[0036] In some exemplary aspects, network 103 may include a wired network, a local area network (LAN), a wireless network, a wireless LAN (WLAN) network, a radio network, a cellular network, a WiFi network, an IR network, a Bluetooth (BT) network, and the like.
[0037] In some exemplary aspects, device 102 may be configured to perform and / or execute one or more operations, modules, processes, procedures, and / or the like, as described herein, for example.
[0038] In some exemplary aspects, device 102 may include a compiler 160 that may be configured to generate object code 115 based on source code 112, for example, as described below.
[0039] In some exemplary aspects, compiler 160 may be configured to translate source code 112 into object code 115, for example, as described below.
[0040] In some exemplary aspects, compiler 160 may include or may be implemented as software, software modules, applications, programs, subroutines, instructions, instruction sets, computing code, words, values, symbols, and / or the like.
[0041] In some exemplary aspects, source code 112 may include computer code written in a source language.
[0042] In some exemplary aspects, the source language may include a programming language. For example, the source language may include a high-level programming language, such as, for example, the C language, the C++ language, and / or the like.
[0043] In some exemplary aspects, object code 115 may include computer code written in an object language.
[0044] In some exemplary aspects, the object language may include a low-level language, such as, for example, assembly language, object code, machine code, and the like.
[0045] In some exemplary aspects, object code 115 may include one or more object files, which may create and / or form an executable program, for example.
[0046] In some exemplary aspects, the executable program may be configured to be executed on a target computer. For example, the target computer may include specific computer hardware, a specific machine, and / or a specific operating system.
[0047] In some exemplary aspects, the executable program may be configured to be executed on processor 180, for example, as described below.
[0048] In some exemplary aspects, the processor 180 may include a vector processor 180, for example, as described below. In other aspects, the processor 180 may include any other type of processor.
[0049] Some exemplary aspects are described herein with respect to a compiler (e.g., compiler 160) that is configured to compile source code 112 into target code 115 that is configured to be executed by a vector processor 180, for example, as described below. In other aspects, the compiler (e.g., compiler 160) is configured to compile source code 112 into target code 115 that is configured to be executed by any other type of processor 180.
[0050] In some exemplary aspects, the processor 180 may be implemented as part of the device 102.
[0051] In other aspects, the processor 180 may be implemented as part of any other device that is, for example, separate from the device 102.
[0052] In some exemplary aspects, the vector processor 180 (also referred to as an "array processor") may include processors that can be configured to process an entire vector in one instruction, for example, as described below.
[0053] In other aspects, the executable program may be configured to be executed on any other additional or alternative type of processor.
[0054] In some exemplary aspects, the vector processor 180 may be designed to support high-performance image and / or vector processing. For example, the vector processor 180 may be configured to process 1 / 2 / 3 / 4D arrays of fixed-point data and / or floating-point arrays very quickly and / or efficiently, for example.
[0055] In some exemplary aspects, the vector processor 180 may be configured to process arbitrary data, such as a structure with a pointer to a structure. For example, the vector processor 180 may include a scalar processor to compute non-vector data, assuming the non-vector data is minimal, for example.
[0056] In some exemplary aspects, the compiler 160 may be implemented as a local application to be executed by the device 102. For example, the memory unit 194 and / or the storage unit 195 may store the instructions obtained in the compiler 160, and / or the processor 191 may be configured to execute the instructions obtained in the compiler 160 and / or perform one or more calculations and / or processes of the compiler 160, for example, as described below.
[0057] In other aspects, the compiler 160 may include a remote application to be executed by any suitable computing system (e.g., server 170).
[0058] In some exemplary aspects, server 170 may include at least a remote server, a network-based server, a cloud server, and / or any other server.
[0059] In some exemplary aspects, server 170 may include a suitable memory and / or storage unit 174 and a suitable processor 171, the memory and / or storage unit having instructions obtained in compiler 160 stored thereon, the processor being configured to execute the instructions, for example, as described below.
[0060] In some exemplary aspects, compiler 160 may include a combination of a remote application and a local application.
[0061] In one instance, compiler 160 may be downloaded and / or received by a user of device 102 from another computing system (e.g., server 170) such that compiler 160 may be executed locally by the user of device 102. For example, the instructions may be received and stored, for example, temporarily, in the memory of device 102 or any suitable short-term memory or buffer, for example, before being executed by processor 191 of device 102.
[0062] In another instance, compiler 160 may include a client module to be executed locally by device 102 and a server module to be executed by server 170. For example, the client module may include and / or may be implemented as a local application, a web application, a website, a web client, for example, a HyperText Markup Language (HTML) web application, etc.
[0063] For example, one or more first operations of compiler 160 may be executed locally by device 102, and / or one or more second operations of compiler 160 may be executed remotely by server 170.
[0064] In other aspects, compiler 160 may include any other suitable computing arrangement and / or scheme, or may be implemented by any other suitable computing arrangement and / or scheme.
[0065] In some exemplary aspects, system 100 may include an interface 110 (e.g., a user interface) to interface between a user of device 102 and one or more elements of system 100 (e.g., compiler 160).
[0066] In some exemplary aspects, interface 110 may be implemented using any suitable hardware components and / or software components, such as a processor, a controller, a memory unit, a storage unit, an input unit, an output unit, a communication unit, an operating system, and / or an application.
[0067] In some aspects, interface 110 may be implemented as part of any suitable module, system, device, or component of system 100.
[0068] In other aspects, interface 110 may be implemented as a separate element of system 100.
[0069] In some exemplary aspects, interface 110 may be implemented as part of device 102. For example, interface 110 may be associated with and / or included as part of device 102.
[0070] In one instance, interface 110 may be implemented as part of, for example, middleware and / or any suitable application of device 102. For example, interface 110 may be implemented as part of compiler 160 and / or as part of the OS of device 102.
[0071] In some exemplary aspects, interface 110 may be implemented as part of server 170. For example, interface 110 may be associated with and / or included as part of server 170.
[0072] In one instance, interface 110 may include or be part of: web-based applications, websites, web pages, plugins, ActiveX controls, rich content components (e.g., Flash or Shockwave components), etc.
[0073] In some exemplary aspects, interface 110 may be associated with and / or may include: for example, gateway (GW) 113 and / or application programming interface (API) 114, for example, to transfer information and / or communication between elements of system 100 and / or to one or more other parties (e.g., internal or external parties), users, applications, and / or systems.
[0074] In some aspects, interface 110 may include any suitable graphical user interface (GUI) 116 and / or any other suitable interface.
[0075] In some exemplary aspects, interface 110 may be configured to receive source code 112 from a user of, for example, device 102 via, for example, GUI 116 and / or API 114.
[0076] In some exemplary aspects, interface 110 may be configured to transfer source code 112 to, for example, compiler 160, for example, to generate object code 115, as described below.
[0077] Reference Figure 2 , which schematically illustrates compiler 200 according to some exemplary aspects. For example, compiler 160 ( Figure 1 ) may implement one or more elements of compiler 200, and / or may perform one or more operations and / or functionality of compiler 200.
[0078] In some exemplary aspects, asFigure 2 As shown, the compiler 200 can be configured to generate object code 233, for example, by compiling source code 212 in a source language.
[0079] In some exemplary aspects, as Figure 2 shown, the compiler 200 can include a front end 210 that is configured to receive and analyze source code 212 in a source language.
[0080] In some exemplary aspects, the front end 210 can be configured to generate intermediate code 213, for example, based on the source code 212.
[0081] In some exemplary aspects, the intermediate code 213 can include a lower-level representation of the source code 212.
[0082] In some exemplary aspects, the front end 210 can be configured to perform, for example, lexical analysis, syntactic analysis, semantic analysis, and / or any other additional or alternative type of analysis on the source code 212.
[0083] In some exemplary aspects, the front end 210 can be configured to utilize the analysis results of the source code 212 to identify errors and / or problems. For example, the front end 210 can be configured to generate error information, for example, including error and / or warning messages, for example, the error information can identify the location in the source code 212, for example, the location where the error or problem is detected.
[0084] In some exemplary aspects, as Figure 2 shown, the compiler 200 can include a middle end 220 that is configured to receive and process the intermediate code 213 and generate adjusted (e.g., optimized) intermediate code 223.
[0085] In some exemplary aspects, the middle end 220 can be configured to perform one or more adjustments (e.g., optimizations) on the intermediate code 213, for example, to generate the adjusted intermediate code 223.
[0086] In some exemplary aspects, the middle end 220 can be configured to perform one or more optimizations on the intermediate code 213, for example, independent of the type of target computer on which the object code 233 is to be executed.
[0087] In some exemplary aspects, the middle end 220 can be implemented to support the use of the optimized intermediate code 223, for example, for different machine types.
[0088] In some exemplary aspects, the middle end 220 can be configured to optimize the intermediate representation of the intermediate code 223, for example, to improve the performance and / or quality of the resulting object code 233.
[0089] In some exemplary aspects, one or more optimizations to the intermediate code 213 may include, for example, inlining expansion, dead code elimination, constant propagation, loop transformation, parallelization, and / or the like.
[0090] In some exemplary aspects, as Figure 2 shown, the compiler 200 may include a backend 230 that is configured to receive and process the adjusted intermediate code 213 and to generate target code 233 based on the adjusted intermediate code 213.
[0091] In some exemplary aspects, the backend 230 may be configured to perform one or more operations and / or processes that may be specific to the target computer for executing the target code 233. For example, the backend 230 may be configured to process the optimized intermediate code 213 by applying analysis, transformation, and / or optimization operations to the adjusted intermediate code 213, which may be configured, for example, based on the target computer for executing the target code 233.
[0092] In some exemplary aspects, one or more analysis, transformation, and / or optimization operations applied to the adjusted intermediate code 213 may include, for example, resource and storage decisions such as, for example, register allocation, instruction scheduling, and / or the like.
[0093] In some exemplary aspects, the target code 233 may include target - specific assembly code that may be specific to the target computer for executing the target code 233 and / or the target operating system of the target computer.
[0094] In some exemplary aspects, the target code 233 may include target - specific assembly code for a processor (e.g., vector processor 180( Figure 1 ))
[0095] In some exemplary aspects, the compiler 200 may include a vector microcode processor (VMP) Open Computing Language (OpenCL) compiler, as described below, for example. In other aspects, the compiler 200 may include any other type of vector processor compiler or may be implemented as part of any other type of vector processor compiler.
[0096] In some exemplary aspects, the VMP OpenCL compiler may include a low - level virtual machine (LLVM) - based compiler that may be configured according to an LLVM - based compilation scheme, for example, to degrade OpenCL C code to VMP accelerator assembly code, for example, suitable for execution by the vector processor 180( Figure 1 )
[0097] In some exemplary aspects, compiler 200 may include one or more techniques that may be required to compile code into a format suitable for the VMP architecture, e.g., in addition to the open-source LLVM compiler passes.
[0098] In some exemplary aspects, FE 210 may be configured to parse OpenCL C code and translate it, e.g., via an Abstract Syntax Tree (AST), into, e.g., LLVM Intermediate Representation (IR).
[0099] In some exemplary aspects, compiler 200 may include a dedicated API, e.g., to detect the correct patterns for compiler pattern matching, e.g., patterns suitable for VMP. For example, VMP may be configured as a Complex Instruction Set Computer (CISC) machine that implements a very complex Instruction Set Architecture (ISA) that may be difficult to target from standard C code. Accordingly, compiler pattern matching may not easily detect the correct patterns, and for such cases, the compiler may require a dedicated API.
[0100] In some exemplary aspects, FE 210 may implement one or more vendor extensions built-ins that may be targeted at the VMP-specific ISA, e.g., in addition to the standard OpenCL built-ins that may be optimized for the VMP machine.
[0101] In some exemplary aspects, FE 210 may be configured to implement OpenCL constructs and / or work-item functions.
[0102] In some exemplary aspects, ME 220 may be configured to process LLVM IR code, which may be generic and target-independent, e.g., although it may include one or more hooks for a specific target architecture.
[0103] In some exemplary aspects, ME 220 may perform one or more custom passes, e.g., to support the VMP architecture, e.g., as described below.
[0104] In some exemplary aspects, ME 220 may be configured to perform one or more operations of Control Flow Graph (CFG) linearization analysis, e.g., as described below.
[0105] In some exemplary aspects, CFG linearization analysis may be configured to linearize code, e.g., in cases where VMP vector code does not support standard control flow, e.g., by converting if statements to select patterns.
[0106] In one instance, ME 220 may receive a given code, e.g., as follows:
[0107]
[0108] According to this example, ME 220 can be configured to apply CFG linearization analysis to a given code, for example, as follows:
[0109] tmpA = A + 5;
[0110] tmpB = B * 2;
[0111] mask = x > 0;
[0112] A = Select mask, tmpA, A
[0113] B = Select not mask, tmpB, B
[0114] Example (1)
[0115] In some exemplary aspects, ME 220 can be configured to perform one or more operations of auto-vectorization analysis, for example, as described below.
[0116] In some exemplary aspects, auto-vectorization analysis can be configured to vectorize (e.g., auto-vectorize) a given code, for example, to utilize the vector capabilities of the VMP.
[0117] In some exemplary aspects, ME 220 can be configured to perform auto-vectorization analysis, for example, to vectorize the code into scalar form. For example, some or all operations of auto-vectorization analysis may not be performed if the code is already provided in vectorized form.
[0118] In some exemplary aspects, for example, in some use cases and / or scenarios, the compiler may not always be able to auto-vectorize the code, for example, due to data dependencies between loop iterations.
[0119] In one example, ME 220 can receive a given code, for example, as follows:
[0120]
[0121] According to this example, ME 220 can be configured to perform CFG auto-vectorization analysis by applying a first transformation, for example, as follows:
[0122]
[0123] Example (2a)
[0124] For example, ME 220 can be configured to perform CFG auto-vectorization analysis by applying a second transformation, for example, after the first transformation, for example, as follows:
[0125]
[0126] Example (2b)
[0127] In some exemplary aspects, ME 220 may be configured to perform one or more operations of Scratch Pad Memory Loop Access Analysis (SPMLAA), for example, as described below.
[0128] In some exemplary aspects, SPMLAA may define a Processing Block (PB), for example, which should later be outlined and compiled for the VMP.
[0129] In some exemplary aspects, the processing block may include an acceleration loop, which may be executed by the vector unit of the VMP.
[0130] In some exemplary aspects, a PB (e.g., each PB) may include memory references. For example, some or all of the memory accesses may refer to local memory banks.
[0131] In some exemplary aspects, the VMP may enable access to the memory banks through an AGU (e.g., the AGU 320 described below with reference to Figure 3 and a Scatter-Gather unit (SG).
[0132] In some exemplary aspects, the AGU may be pre-configured, for example, before the loop execution. For example, the loop trip count may be calculated, for example, before running the processing block.
[0133] In some exemplary aspects, image references may be created at this stage, for example, some or all of the image references, and then the stride and offset may be calculated, for example, the per-dimension stride and offset for each reference.
[0134] In some exemplary aspects, ME 220 may be configured to perform one or more operations of AGU Planner Analysis, for example, as described below.
[0135] In some exemplary aspects, the AGU Planner Analysis may include iterator specification, which may cover image references from the entire processing block, for example, all image references.
[0136] In some exemplary aspects, the iterator may cover a single reference or a group of references.
[0137] In some exemplary aspects, one or more memory references may be merged through shuffle instructions and / or reuse the same access, and / or save the values read from previous iterations.
[0138] In some exemplary aspects, other memory references without a linear access pattern, for example, can be processed using scatter-gather (SG) units, which may have a performance penalty, e.g., because they may need to maintain indices and / or masks.
[0139] In some exemplary aspects, scheduling can be configured as an arrangement of iterators in a processing block. For example, a processing block can, in theory, have multiple schedules, for example.
[0140] In some exemplary aspects, the AGU scheduler analysis can be configured to build all possible schedules for all PBs and, for example, select a combination from all valid combinations, e.g., the best combination.
[0141] In some exemplary aspects, the total number of iterators in a valid combination may be limited, e.g., not exceeding the number of available AGUs on the VMP.
[0142] In some exemplary aspects, one or more parameters can be defined for an iterator (e.g., for each iterator), including, for example, stride, width, and / or base, e.g., as part of the AGU scheduler analysis. For example, the minimum-maximum range for an iterator can be defined in terms of dimensions, e.g., for each dimension, e.g., as part of the AGU scheduler analysis.
[0143] In some exemplary aspects, the AGU scheduler analysis can be configured to track and evaluate memory references to an image, e.g., each memory reference, e.g., to understand its access pattern.
[0144] In one instance, according to Instance 2a / 2b, image "a" as the base address can be accessed with a 32-byte stride for 64 iterations.
[0145] In some exemplary aspects, LLVM can include Scalar Evaluation Analysis (SCEV), which can compute access patterns, e.g., to understand each image reference.
[0146] In some exemplary aspects, ME 220 can utilize the masking capabilities of the AGU, e.g., to avoid maintaining induction variables, which may have a performance penalty.
[0147] In some exemplary aspects, ME 220 can be configured to perform one or more operations of rewrite analysis, e.g., as described below.
[0148] In some exemplary aspects, rewrite analysis can be configured to transform the code of a processing block, e.g., when setting iterators and / or modifying memory access instructions.
[0149] In some exemplary aspects, the setup of iterators (e.g., all iterators) can be implemented in IR in a target-specific intrinsic function. For example, the setup of iterators can reside in the pre-header of the outermost loop.
[0150] In some exemplary aspects, rewrite analysis can include loop perfection analysis, as described below, for example.
[0151] In some exemplary aspects, code can be compiled with the goal that substantially all computations should be performed within the innermost loop.
[0152] For example, loop perfection analysis can promote instructions, e.g., to move operations that occur after the last iteration of a loop into the loop.
[0153] For example, loop perfection analysis can sink instructions, e.g., to move operations that occur before the first iteration of a loop into the loop.
[0154] For example, loop perfection analysis can promote instructions and / or sink instructions such that substantially all instructions are moved from outer loops to the innermost loop.
[0155] For example, loop perfection analysis can be configured to provide a technical solution to support VMP iterators, e.g., to work only on perfectly nested loops.
[0156] For example, loop perfection analysis may result in a situation where there are no instructions between "for" statements that form a loop, e.g., to support a VMP iterator that cannot simulate such a situation.
[0157] In some exemplary aspects, loop perfection analysis can be configured to fold nested loops into a single folded loop.
[0158] In one instance, ME 220 can receive a given code, e.g., as follows:
[0159]
[0160]
[0161] According to this instance, ME 220 can be configured to perform loop perfection analysis to fold the nested loops in the code into a single folded loop, e.g., as follows:
[0162]
[0163] Instance (3)
[0164] In some exemplary aspects, ME 220 can be configured to perform one or more operations of vector loop demarcation analysis, as described below, for example.
[0165] In some exemplary aspects, vector loop demarcation analysis may be configured to partition code between a scalar subsystem and a vector subsystem, e.g., between a vector processing block 310 ( Figure 3 as described below) and a scalar processor 330 ( Figure 3 ). Figure 3 )
[0166] In some exemplary aspects, a VMP accelerator may include scalar and / or vector subsystems, e.g., as described below. For example, each of the subsystems may have different computing units / processors. Accordingly, scalar code may be compiled on a scalar compiler (e.g., an SSC compiler), and / or accelerated vector code may run on a VMP vector processor.
[0167] In some exemplary aspects, vector loop demarcation analysis may be configured to create separate functions for loop bodies of accelerated vector code. For example, these functions may be marked for the VMP and / or may proceed to the VMP backend, e.g., while the rest of the code may be compiled by the SSC compiler.
[0168] In some exemplary aspects, one or more portions of a vector loop (e.g., configuration of vector units and / or initialization of vector registers) may be performed by a scalar unit. However, these portions may be performed at a later stage, e.g., by backpatching scalar code, e.g., because scalar code may still be in LLVM IR before being processed by the SSC compiler.
[0169] In some exemplary aspects, BE 230 may be configured to translate LLVM IR into machine instructions. For example, BE 230 may not be target-agnostic and may be familiar with target-specific architectures and optimizations, e.g., as compared to ME 220 which may be agnostic to target-specific architectures.
[0170] In some exemplary aspects, BE 230 may be configured to perform one or more analyses that may be specific to the target machine (e.g., a VMP machine) to which the code is being degraded, e.g., although BE 230 may use general-purpose LLVM.
[0171] In some exemplary aspects, BE 230 may be configured to perform one or more operations of instruction degradation analysis, e.g., as described below.
[0172] In some exemplary aspects, instruction degradation analysis may be configured to translate LLVM IR into target-specific instruction machine IR (MIR), e.g., by translating LLVM IR into a directed acyclic graph (DAG).
[0173] In some exemplary aspects, the DAG may undergo a process of instruction legalization, for example, based on data types and / or VMP instructions, which may be supported by VMP HW.
[0174] In some exemplary aspects, instruction demotion analysis may be configured to perform a pattern matching process, for example, after the instruction legalization process, e.g., to demote nodes (e.g., each node) in the DAG into, for example, VMP-specific machine instructions.
[0175] In some exemplary aspects, instruction demotion analysis may be configured to generate MIR, for example, after the pattern matching process.
[0176] In some exemplary aspects, instruction demotion analysis may be configured to demote instructions according to a machine application binary interface (ABI) and / or calling convention.
[0177] In some exemplary aspects, BE 230 may be configured to perform one or more operations of unit balance analysis, for example, as described below.
[0178] In some exemplary aspects, unit balance analysis may be configured to balance instructions among VMP computing units, for example, among the data processing units 316 ( Figure 3 as referred to below Figure 3 ).
[0179] In some exemplary aspects, unit balance analysis may be familiar with some or all available arithmetic transformations, and / or may perform transformations according to an optimal algorithm.
[0180] In some exemplary aspects, BE 230 may be configured to perform one or more operations of modulo scheduler (pipeliner) analysis, for example, as described below.
[0181] In some exemplary aspects, the pipeliner may be configured to schedule instructions according to one or more constraints (e.g., data dependencies, resource bottlenecks, and / or any other constraints), for example, using the swing modulo scheduling (SMS) heuristic and / or any other additional and / or alternative heuristic.
[0182] In some exemplary aspects, the pipeliner may be configured to schedule, for example, a set of very long instruction word (VLIW) instructions (e.g., of the initiation interval (II)) that the program will iterate over during the steady state.
[0183] In some exemplary aspects, a performance metric may be measured, which may be based on the number of cycles executable by a typical loop, for example, as follows:
[0184] (Input data size in bytes) * II / (Bytes consumed / produced per iteration)
[0185] In some exemplary aspects, the pipeliner may attempt to minimize the II as much as possible, e.g., to improve performance.
[0186] In some exemplary aspects, the pipeliner may be configured to compute the minimum II and schedule accordingly. For example, if the pipeliner scheduling fails, the pipeliner may attempt to increase the II and retry the scheduling, e.g., until a predefined II threshold is violated.
[0187] In some exemplary aspects, the BE 230 may be configured to perform one or more operations of register allocation analysis, e.g., as described below.
[0188] In some exemplary aspects, the register allocation analysis may be configured to attempt to assign registers in an efficient (e.g., optimal) manner.
[0189] In some exemplary aspects, the register allocation analysis may assign values to bypass vector registers, general-purpose vector registers, and / or scalar registers.
[0190] In some exemplary aspects, the values may include private variables, constants, and / or values rotated across iterations.
[0191] In some exemplary aspects, the register allocation analysis may implement an optimal heuristic suitable for one or more VMP register file (regfile) constraints. For example, in some use cases, the register allocation analysis may not use standard LLVM register allocation.
[0192] In some exemplary aspects, in some cases, the register allocation analysis may fail, which may mean that the loop cannot be compiled. Accordingly, the register allocation analysis may implement a retry mechanism that may return to the modulo scheduler and may attempt to reschedule the loop, e.g., with an increased startup interval. For example, in many cases, increasing the startup interval may reduce register shortages and / or may support the compilation of vector loops.
[0193] In some exemplary aspects, the BE 230 may be configured to perform one or more operations of SSC configuration analysis, e.g., as described below.
[0194] In some exemplary aspects, the SSC configuration analysis may be configured to set the configuration for executing the kernel, e.g., the AGU configuration.
[0195] In some exemplary aspects, the SSC configuration analysis may be performed at a later stage, e.g., due to the configuration computed after legalization, register allocation analysis, and / or modulo scheduling analysis.
[0196] In some exemplary aspects, the SSC configuration analysis may include a zero-overhead loop (ZOL) mechanism in a vector loop. For example, the ZOL mechanism may configure the loop trip count based on the access pattern of memory references in the loop, e.g., to avoid executing instructions that check the loop exit condition on each iteration.
[0197] In some exemplary aspects, the VMP compilation flow may include one or more (e.g., a few) steps that may be invoked during the compilation flow in a test library (e.g., a wrapper script for compilation, execution, and / or program testing). For example, these steps may be executed outside of the LLVM compiler.
[0198] In some exemplary aspects, a PCB hardware description language (PHDL) simulator may be implemented to perform one or more roles of an assembler, encoder, and / or linker.
[0199] In some exemplary aspects, the compiler 200 may be configured to provide technical solutions to support robustness, which may enable the compilation of a wide range of loop selections in the presence of HW limitations. For example, the compiler 200 may be configured to support technical solutions that may not produce verification errors.
[0200] In some exemplary aspects, the compiler 200 may be configured to provide technical solutions to support programmability, which may provide the user with the ability to express code in multiple ways that can be correctly compiled to the VMP architecture.
[0201] In some exemplary aspects, the compiler 200 may be configured to provide technical solutions to support an improved user experience, which may allow the user to be able to debug and / or profile the code. For example, the improved user experience may provide informative error messages, reporting tools, and / or profiling tools.
[0202] In some exemplary aspects, the compiler 200 may be configured to provide technical solutions to support improved performance, e.g., to optimize VMP assembly code and / or iterator access, which may result in faster execution. For example, improved performance may be achieved through high-utilization computing units and the use of their complex CISC.
[0203] Reference Figure 3 , which schematically illustrates a vector processor 300 according to some exemplary aspects. For example, the vector processor 180 ( Figure 1 ) may implement one or more elements of the vector processor 300, and / or may perform one or more operations and / or functionality of the vector processor 300.
[0204] In some exemplary aspects, the vector processor 300 may include a vector microcode processor (VMP).
[0205] In some exemplary aspects, vector processor 300 may include a wide vector machine, e.g., supporting a very long instruction word (VLIW) architecture and / or a single instruction / multiple data (SIMD) architecture.
[0206] In some exemplary aspects, vector processor 300 may be configured to provide a technical solution to support high performance for short integer types, which may be common in, e.g., computer vision and / or deep learning algorithms.
[0207] In other aspects, vector processor 300 may include any other type of vector processor, and / or may be configured to support any other additional or alternative functionality.
[0208] In some exemplary aspects, as Figure 3 shown, vector processor 300 may include a vector processing block (vector processor) 310, a scalar processor 330, and a direct memory access (DMA) 340, e.g., as described below.
[0209] In some exemplary aspects, as Figure 3 shown, vector processing block 310 may be configured to process (e.g., efficiently process) image data and / or vector data. For example, vector processing block 310 may be configured to use vector calculation units, e.g., to accelerate calculations.
[0210] In some exemplary aspects, scalar processor 330 may be configured to perform scalar calculations. For example, scalar processor 330 may act as "glue logic" for a program that includes vector calculations. For example, some (e.g., even most) of the calculations of a program may be performed by vector processing block 310. However, several tasks (e.g., some basic tasks) (e.g., scalar calculations) may be performed by scalar processor 330.
[0211] In some exemplary aspects, DMA 340 may be configured to interface with one or more memory elements in a chip that includes vector processor 300.
[0212] In some exemplary aspects, DMA 340 may be configured to read inputs from main memory and / or write outputs to main memory.
[0213] In some exemplary aspects, scalar processor 330 and vector processing block 310 may use respective local memories to process data.
[0214] In some exemplary aspects, as Figure 3 shown, vector processor 300 may include an extractor and decoder 350, which may be configured to control scalar processor 330 and / or vector processing block 310.
[0215] In some exemplary aspects, the operation of scalar processor 330 and / or vector processing block 310 may be triggered by instructions stored in program memory 352.
[0216] In some exemplary aspects, DMA 340 may be configured to transfer data, for example, in parallel with the execution of program instructions in memory 352.
[0217] In some exemplary aspects, DMA 340 may be controlled by software, for example, via a configuration register and not by instructions, and may accordingly be considered a second execution “thread” in vector processor 300.
[0218] In some exemplary aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include one or more data processing units, for example, a set of data processing units, as described below, for example.
[0219] In some exemplary aspects, a data processing unit may include hardware configured to perform computations, such as an arithmetic logic unit (ALU).
[0220] In one instance, a data processing unit may be configured to add numbers and / or store numbers in memory.
[0221] In some exemplary aspects, a data processing unit may be controlled by commands encoded in program memory 352 and / or in a configuration register. For example, the configuration register may be memory-mapped and may be written to by memory store commands of scalar processor 330.
[0222] In some exemplary aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include a status configuration that includes a set of registers and memory, as described below, for example.
[0223] In some exemplary aspects, as Figure 3 shown, vector processor block 310 may include a set of vector memories 312 that may be configured to store data to be processed by vector processor block 310, for example.
[0224] In some exemplary aspects, as Figure 3 shown, vector processor block 310 may include a set of vector registers 314 that may be configured to be used in data processing performed by vector processor block 310, for example.
[0225] In some exemplary aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may be associated with a set of memory mappings.
[0226] In some exemplary aspects, the memory map may include a set of addresses accessible by a data processing unit, which can load data from / to and / or store data in registers and memory.
[0227] In some exemplary aspects, as Figure 3 shown, the vector processing block 310 may include a plurality of address generation units (AGUs) 320, which may include addresses accessible to them, for example, in one or more memories in the memory 312.
[0228] In some exemplary aspects, as Figure 3 shown, the vector processor block 310 may include a plurality of data processing units 316, for example, as described below.
[0229] In some exemplary aspects, the data processing unit 316 may be configured to process commands, for example, including several digits at a time. In one instance, the command may include 8 digits. In another instance, the command may include 4 digits, 16 digits, or any other counted number of digits.
[0230] In some exemplary aspects, two or more data processing units 316 may be used simultaneously. In one instance, the data processing unit 316 may process and execute multiple different commands in an entire single cycle, for example, 3 different commands, for example, including 8 digits.
[0231] In some exemplary aspects, the data processing unit 316 may be asymmetric. For example, the first and second data processing units 316 may support different commands. For example, addition may be performed by the first data processing unit 316, and / or multiplication may be performed by the second data processing unit 316. For example, both operations may be performed by one or more additional other data processing units 316.
[0232] In some exemplary aspects, the data processing unit 316 may be configured to support arithmetic operations for many combinations of input and output data types.
[0233] In some exemplary aspects, the data processing unit 316 may be configured to support one or more operations that may be less common. For example, the processing unit 316 may support operations working with a look-up table (LUT) of the vector processor 300 and / or any other operations.
[0234] In some exemplary aspects, the data processing unit 316 may be configured to support efficient computation of non-linear functions, histograms, and / or random data access, for example, which may help to implement algorithms such as image scaling, Hough transform, and / or any other algorithms.
[0235] In some exemplary aspects, the vector memory 312 may include, for example, a memory bank having a size of 16K or any other size, and the memory bank may be accessed in the same cycle.
[0236] In one instance, the maximum memory access size may be 64 bits. According to this instance, the peak throughput may be 256 bits, for example, 64 x 4 = 256. For example, a high memory bandwidth may be achieved to utilize the computing power of the data processing unit 316.
[0237] In one instance, two data processing units 316 may support 16 eight-bit multiply-accumulate operations (MACs) per cycle. According to this instance, the two data processing units 316 may not be useful, for example, in the case where the input numbers are not fetched at this speed, and / or in the absence of exactly 256 bits of input, for example, 16 x 8 x 2 = 256.
[0238] In some exemplary aspects, the AGU 320 may be configured to perform memory access operations, for example, load and store data from / to the vector memory 314.
[0239] In some exemplary aspects, the AGU 320 may be configured to calculate the addresses of input and output data items, for example, to process I / O in cases where, for example, the high bandwidth is not sufficient to utilize the data processing unit 316.
[0240] In some exemplary aspects, the AGU 320 may be configured to calculate the addresses of input and / or output data items, for example, before typing a vector command block (e.g., a loop), based on configuration registers written by the scalar processor 330.
[0241] For example, the AGU 320 may be configured to write an image base address pointer, width, height, and / or stride to configuration registers, for example, in order to iterate over an image.
[0242] In some exemplary aspects, the AGU 320 may be configured to handle addressing (e.g., all addressing), for example, to provide a technical solution where the data processing unit 316 may not have the burden of incrementing a pointer or counter in a loop and / or the burden of checking for an end-of-line condition, for example, to zero out a counter in a loop.
[0243] In some exemplary aspects, as Figure 3 shown, the AGU 320 may include 4 AGUs, and correspondingly, four memories 312 may be accessed in the same cycle. In other aspects, any other count of AGUs 32 may be implemented.
[0244] In some exemplary aspects, the AGU 320 may not be "bound" to the memory bank 312. For example, the AGU 320 (e.g., each AGU 320) may access the memory bank 312 (e.g., each memory bank 312), e.g., as long as two or more AGU 320s do not attempt to access the same memory bank 312 in the same cycle.
[0245] In some exemplary aspects, the vector register 314 may be configured to support communication between the data processing unit 316 and the AGU 320.
[0246] In one instance, the total number of vector registers 314 may be 28, and they may be partitioned into subsets, e.g., based on their functionality. For example, a first subset of the vector registers 314 may be used for input / output of, e.g., all data processing units 316 and / or the AGU 320; and / or a second subset of the vector registers 314 may not be used for output of some operations (e.g., most operations) and may be used for one or more other operations, e.g., to store loop-invariant inputs.
[0247] In some exemplary aspects, the data processing unit 316 (e.g., each data processing unit 316) may have one or more registers for hosting the output of the last executed operation, e.g., the output may be fed as input to other data processing units 316. For example, these registers may "bypass" the vector register 314 and may operate faster than writing these outputs to the first set of vector registers 314.
[0248] In some exemplary aspects, the fetcher and decoder 350 may be configured to support low-overhead vector loops, e.g., very low-overhead vector loops (also referred to as "zero-overhead vector loops"), e.g., where it may not be necessary to check the termination (exit) condition of the vector loop during its execution.
[0249] For example, when the AGU 320 finishes iterating over the configured memory region, e.g., the AGU 320 may signal the termination (exit) condition.
[0250] For example, when the AGU 320 signals the termination condition, e.g., the fetcher and decoder 350 may exit the loop.
[0251] For example, the scalar processor 330 may be utilized to configure loop parameters, e.g., the first and last instructions and / or the exit condition.
[0252] In one example, vector loops can be exploited, for example, with high memory bandwidth and / or inexpensive addressing to solve control and data flow problems, e.g., to provide a technical solution to allow a data processing unit 316 to process data with substantially no additional overhead.
[0253] In some exemplary aspects, a scalar processor 330 can be configured to provide one or more functions that can be complementary to the functions of the vector processing block 310. For example, most (e.g., the majority) of the work in a vector program can be performed by the data processing unit 316. For example, the scalar processor 330 can be utilized to "glue" together the various vector code blocks of a vector program.
[0254] In some exemplary aspects, the scalar processor 330 can be implemented separately from the vector processing block 310. In other aspects, the scalar processor 330 can be configured to share one or more components and / or functions with the vector processing block 310.
[0255] In some exemplary aspects, the scalar processor 330 can be configured to perform operations that may not be suitable for execution on the vector processing block 310.
[0256] For example, the scalar processor 330 can be utilized to execute 32-bit C programs. For example, the scalar processor 330 can be configured to support 1, 2, and / or 4-byte data types of C code and / or some or all of the arithmetic operators of C code.
[0257] For example, the scalar processor 330 can be configured to provide a technical solution to perform operations that cannot be executed on the vector processing block 310, e.g., without using a fully-on CPU.
[0258] In some exemplary aspects, the scalar processor 330 can include, for example, a scalar data memory 332 having a size of 16K or any other size, which can be configured to store data, e.g., variables used by the scalar portion of a program.
[0259] For example, the scalar processor 330 can store local and / or global variables declared by portable C code, and these variables can be allocated to the scalar data memory by a compiler (e.g., compiler 200( Figure 2 ))
[0260] In some exemplary aspects, as Figure 3 shown, the scalar processor 330 can include a set of vector registers 334 or can be associated therewith, and this set of vector registers can be used for data processing performed by the scalar processor 330.
[0261] In some exemplary aspects, the scalar processor 330 may be associated with a scalar memory map that may support the scalar processor 330 in accessing substantially all of the state of the vector processor 300. For example, the scalar processor 330 may configure the vector units and / or DMA channels via the scalar memory map.
[0262] In some exemplary aspects, the scalar processor 330 may not be permitted to access one or more block control registers that may be used by an external processor to run and debug vector programs.
[0263] In some exemplary aspects, the DMA 340 may be configured to communicate with one or more other components of the chip implementing the vector processor 300, for example, via the main memory. For example, the DMA 340 may be configured to transfer data blocks, e.g., large, contiguous data blocks, e.g., to support the scalar processor 330 and / or the vector processing block that may manipulate data stored in the local memory. For example, a vector program may be able to use the DMA 340 to read data from the main chip memory.
[0264] In some exemplary aspects, the DMA 340 may be configured to communicate with other elements of the chip via, for example, multiple DMA channels (e.g., 8 DMA channels or any other count of DMA channels). For example, a DMA channel (e.g., each DMA channel) may be able to transfer a rectangular patch from the local memory to the main chip memory, or vice versa. In other aspects, the DMA channels may transfer any other type of data block between the local memory and the main chip memory.
[0265] In some exemplary aspects, a rectangular patch may be defined by a base address pointer, width, height, and stride.
[0266] For example, at peak throughput, 8 bytes may be transferred per cycle; however, there may be overhead for each patch and / or for each row within a patch.
[0267] In some exemplary aspects, the DMA 340 may be configured to transfer data in parallel with computations via, for example, multiple DMA channels, e.g., as long as the commands being executed do not access the local memory involved in the transfer.
[0268] In one instance, since all channels may access the same memory bus, implementing a transfer using several channels may not save I / O cycles, for example, compared to the case when using a single channel. However, multiple DMA channels may be utilized to schedule several transfers and execute them in parallel with computations. For example, this may be advantageous compared to a single channel that may not permit scheduling a second transfer until the first transfer is complete.
[0269] In some exemplary aspects, the DMA 340 may be associated with a memory map that may support DMA channel access to vector memories and / or scalar data. For example, access to vector memories may occur in parallel with computations. For example, access to scalar data may generally not be allowed to be parallel, for example, because the scalar processor 330 may be involved in almost any reasonable program and may access its local variables during a transfer, which may result in memory contention with an active DMA channel.
[0270] In some exemplary aspects, the DMA 340 may be configured to provide technical solutions to support the parallelization of I / O and computations. For example, a program performing computations may not have to wait for I / O, for example, in cases where these computations can be run quickly by the vector processing block 310.
[0271] In some exemplary aspects, an external processor (e.g., a CPU) may be configured to initiate the execution of a program on the vector processor 300. For example, the vector processor 300 may remain idle, for example, as long as program execution has not been initiated.
[0272] In some exemplary aspects, an external processor may be configured to debug a program, for example, execute a single step at a time, stop when the program reaches a breakpoint, and / or inspect the contents of registers and memories storing program variables.
[0273] In some exemplary aspects, an external memory map may be implemented to support an external processor in controlling the vector processor 300 and / or debugging a program, for example, by writing to the control registers of the vector processor 300.
[0274] In some exemplary aspects, the external memory map may be implemented as a superset of the scalar memory map. For example, this implementation may make all registers and memories defined by the architecture of the vector processor 300 accessible to a debugger backend running on an external processor.
[0275] In some exemplary aspects, for example when the vector processor 300 terminates a program, the vector processor 300 may issue an interrupt signal.
[0276] In some exemplary aspects, the interrupt signal may be used, for example, to implement a driver to maintain a queue of programs scheduled for execution by the vector processor 300, and / or may be used, for example, to initiate a new program by an external processor when a previously executed program has completed.
[0277] Return reference Figure 1, in some exemplary aspects, the compiler 160 may be configured to generate the target code 115 based on one or more loops, which may be based on, for example, the source code 112, as described below, for example.
[0278] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115 based on one or more loops, which may be configured, for example, according to a loop execution scheme, as described below, for example.
[0279] In some exemplary aspects, the loop execution scheme may be configured to provide a technical solution to support one or more vector processing architectures, such as, for example, the VLIW architecture and / or any other architecture, as described below, for example.
[0280] In some exemplary aspects, the loop execution scheme may be configured to provide a technical solution to support the improvement of one or more types of loop nests (e.g., imperfect loop nests), as described below, for example.
[0281] In some exemplary aspects, the loop nest may at least include an outer loop and an inner loop, as described below, for example.
[0282] In some exemplary aspects, the loop nest may include an outer loop (e.g., the outermost loop), an inner loop (e.g., the innermost loop), and one or more nested loops (also referred to as "intermediate nested loops" or "intermediate loops"), which may be nested between the outer loop and the inner loop, as described below, for example.
[0283] In some exemplary aspects, the loop nest may include multiple loops nested at multiple nesting levels, as described below, for example.
[0284] In one instance, the multiple loops may include, for example, a first loop (e.g., the outer loop) at a first nesting level and a second loop (e.g., the inner loop) at a second nesting level.
[0285] In one instance, the multiple loops may include one or more intermediate loops, for example, at one or more intermediate nesting levels, for example, between the first nesting level and the second nesting level.
[0286] In one instance, the multiple loops may include three loops at three nesting levels. For example, the three loops may include a first loop at the first nesting level, e.g., the outermost loop; a second loop at the second nesting level, e.g., the intermediate loop; and a third loop at the third nesting level, e.g., the innermost loop. For example, the second loop may be nested in the first loop, and the third loop may be nested in the second loop.
[0287] In some exemplary aspects, it may be desirable to provide a technical solution to efficiently transform an imperfect loop nest into a perfect loop nest, e.g., to improve the performance of an executable program, e.g., when executed by a processor (e.g., a vector processor or any other target processor), e.g., as described below.
[0288] In some exemplary aspects, a perfect loop nest may be configured to include a loop nest where all computational operations of the loop nest reside in the innermost loop of the perfect loop nest.
[0289] For example, the outer loops of a perfect loop nest may not include any computational instructions.
[0290] For example, all computational instructions of a perfect loop nest may be in the innermost loop of the perfect loop nest.
[0291] In one instance, one or more processor architectures may need to use a perfect loop nest in a program and / or may benefit from using a perfect loop nest in a program.
[0292] In another instance, a perfect loop nest may be applicable to one or more (e.g., more) loop optimizations.
[0293] In another instance, for example when a perfect loop nest is not used, one or more scheduling schemes (e.g., modulo scheduling) that may be critical optimizations for a VLIW target may not be able to optimize code across loop levels / basic blocks.
[0294] In another instance, one or more processor architectures may only support perfect loop nests. For example, these architectures may rely on being able to perfect a loop nest into a perfect loop.
[0295] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115 based on one or more loops, which may be configured, for example, according to a loop execution scheme that may be configured to provide a technical solution to improve the performance of a program executed by a target processor (e.g., a vector processor), e.g., by transforming an imperfect loop nest into a perfect loop nest, e.g., as described below.
[0296] In some exemplary aspects, the loop execution scheme may be configured to provide a technical solution to improve the performance of a program executed by a target processor (e.g., a vector processor), e.g., by efficiently transforming an imperfect loop nest into a folded loop, e.g., as described below.
[0297] In some exemplary aspects, the folded loop of a perfect loop nest may be configured to include a single basic block loop that includes all nested loops of the perfect loop, e.g., as described below.
[0298] In some exemplary aspects, the execution of a single basic block loop may be preconfigured, e.g., along different dimensions, e.g., the dimension may correspond to an original loop in an original loop nest.
[0299] For example, the execution of a folded loop may be preconfigured, e.g., by control hardware (“HW controlled”), e.g., as described below.
[0300] In one instance, one or more processor architectures may only support folded loops. For example, these processor architectures may rely on the ability to fold and / or refine a loop nest into a single basic block and / or a perfect loop.
[0301] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115 based on one or more loops, which may be configured, e.g., according to a loop execution scheme, which may be configured to provide a technical solution to support transforming an imperfect loop nest into a perfect loop nest, e.g., to improve the performance of an executable program, e.g., as described below.
[0302] In some exemplary aspects, a loop execution scheme may be configured to provide a technical solution to efficiently transform an imperfect loop nest into a folded loop, e.g., by transforming a perfect loop nest into a folded loop, e.g., as described below.
[0303] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115 based on a compilation scheme, which may be configured to, e.g., provide a technical solution to support computing one or more predicates, which may be used to transform an imperfect loop nest into a perfect loop nest and / or transform an imperfect loop nest into a folded loop, e.g., as described below.
[0304] In some exemplary aspects, a predicate may be configured to indicate, identify, confirm, predict, and / or assert the start and / or end of a loop, e.g., the first iteration or the last iteration of a loop.
[0305] In some exemplary aspects, a predicate may be configured to identify the start and / or end of a loop, e.g., even without processing and / or maintaining induction variables, e.g., as described below.
[0306] In one instance, one or more predicates may be utilized to indicate the start and / or end of the execution of one or more inner loops nested within an original loop nest, e.g., as described below.
[0307] In one instance, it may be important to efficiently compute a predicate, e.g., in cases where loop refinement and / or loop folding rely on predication, e.g., as described below.
[0308] In another example, it may be important to efficiently compute predicates, e.g., to support a processor architecture that may not be able to efficiently compute induction variables, as described below.
[0309] In some exemplary aspects, a loop execution scheme may be configured to provide a technical solution to support a processor architecture that may not support predicated instructions for non-memory access operations.
[0310] In some exemplary aspects, a loop execution scheme may be configured to provide a technical solution to support computing predicates, e.g., to support transforming a loop nest into a perfect loop nest and / or a folded loop, as described below.
[0311] In some exemplary aspects, a loop execution scheme may be configured to provide a technical solution to support more efficient execution of a program, e.g., while avoiding the need to compute predicates based on induction variables, as described below.
[0312] In some exemplary aspects, a loop execution scheme may be configured to provide a technical solution to support one or more architectures that do not have predicated operations in hardware and / or have loop nests controlled by hardware, as described below.
[0313] In some exemplary aspects, compiler 160 may be configured to identify one or more loop nests based on source code, as described below.
[0314] In some exemplary aspects, compiler 160 may be configured to identify one or more of the loop nests in source code 112, e.g., in the case where the loop nest is included in source code 112.
[0315] In some exemplary aspects, compiler 160 may be configured to identify one or more of the loop nests in code (e.g., mid-end code that may be compiled from source code 112).
[0316] In some exemplary aspects, compiler 160 may be configured to transform one or more identified loop nests into one or more perfect loop nests, as described below.
[0317] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into target code 115, e.g., such that target code 115 may be based on one or more perfect loop nests, as described below.
[0318] In some exemplary aspects, the compiler 160 may be configured to transform one or more identified loop nests into one or more perfect loop nests, for example, according to a loop perfection scheme, as described below.
[0319] In some exemplary aspects, the compiler 160 may be configured to move one or more outer instructions from the outer loop level of a loop nest to the inner loop (e.g., the innermost loop) of the loop nest, for example, while using one or more predicates to guard the execution of the outer instructions moved to the inner loop, as described below.
[0320] In some exemplary aspects, the predicate may be configured to check and / or represent the state of an induction variable that counts the number of iterations of the inner loop, as described below.
[0321] In some exemplary aspects, the predicate (also referred to as a "loop start predicate") may be configured to identify when the induction variable may be equal to the start of the corresponding inner loop, e.g., for an instruction (also referred to as a "sunk instruction") moved from before the inner loop to the inner loop, as described below.
[0322] In some exemplary aspects, the predicate (also referred to as a "loop end predicate") may be configured to identify when the induction variable may be equal to the end of the corresponding inner loop, e.g., for an instruction (also referred to as a "hoisted instruction") moved from after the inner loop to the inner loop, as described below.
[0323] In some exemplary aspects, the compiler 160 may be configured to move all instructions from the outer loop level (nesting level) of a loop nest to the innermost loop of the loop nest, e.g., to transform the loop nest into a perfect loop nest, as described below.
[0324] In some exemplary aspects, the compiler 160 may be configured to identify one or more outer loop instructions of the outer loop of a loop nest that are external to the inner loop of the loop nest, as described below.
[0325] In some exemplary aspects, the compiler 160 may be configured to move the outer loop instructions to the inner loop of the loop nest, for example, based on a location-based criterion related to the position of the outer loop instructions relative to the inner loop, as described below.
[0326] In some exemplary aspects, the location-based criterion may be used to identify whether the outer loop instructions are before the inner loop ("pre-header instructions") or after the inner loop ("latch instructions"), as described below.
[0327] In some exemplary aspects, the compiler 160 may be configured to transform the outer loop instructions into conditional instructions within the inner loop of the loop nest, as described below.
[0328] For example, a conditional instruction may be configured based on such location-based criteria, as described below, for example.
[0329] In some exemplary aspects, a conditional instruction may be configured based on a predicate, for example, to indicate, identify, confirm, predict, and / or assert an iteration count of an inner loop, as described below, for example.
[0330] In some exemplary aspects, a conditional instruction may be configured as a conditional selection operation that may be based on a predicate of an iteration count of an inner loop, as described below, for example.
[0331] In some exemplary aspects, a conditional instruction may be configured to select between two values based on a predicate of an iteration count of an inner loop, as described below, for example.
[0332] In some exemplary aspects, an outer loop instruction may include an operation on a variable, and the conditional instruction may be configured to select between two values for the variable based on a predicate of an iteration count of an inner loop, as described below, for example.
[0333] In some exemplary aspects, an outer loop instruction may include an operation on a variable, and the conditional instruction may be configured to select between two operations on the variable based on a predicate of an iteration count of an inner loop, as described below, for example.
[0334] In some exemplary aspects, a predicate may be utilized to indicate, identify, confirm, predict, and / or assert whether the inner loop is in the first iteration of the inner loop or the last iteration of the inner loop, as described below, for example.
[0335] In some exemplary aspects, the compiler 160 may be configured to sink instructions, for example, by moving instructions that are to be performed before the first iteration of the inner loop into the inner loop, as described below, for example.
[0336] In some exemplary aspects, a pre-header instruction may be sunk, for example, by moving the pre-header instruction into the inner loop and transforming the pre-header instruction into a pre-header conditional instruction, as described below, for example.
[0337] In some exemplary aspects, a pre-header conditional instruction may include a condition for configuring a result of the pre-header conditional instruction based on a predicate on the inner loop, for example, as described below.
[0338] In some exemplary aspects, the compiler 160 may generate target code 115 based on the compiled code, which may be configured to configure a specific result of the pre-header conditional instruction, for example, when the predicate identifies that the execution of the inner loop is before the first iteration of the inner loop, as described below, for example.
[0339] In some exemplary aspects, the compiler 160 may be configured to promote instructions, for example, by moving instructions that are to be performed after the last iteration of an inner loop into the inner loop, as described below, for example.
[0340] In some exemplary aspects, a latch instruction may be promoted, for example, by moving the latch instruction into the inner loop and transforming the latch instruction into a latch conditional instruction, as described below, for example.
[0341] In some exemplary aspects, the latch conditional instruction may include a condition for configuring the result of the latch conditional instruction based on a predicate, for example, on an inner loop, as described below, for example.
[0342] In some exemplary aspects, the compiler 160 may generate the target code 115 based on the compiled code, which may be configured to configure a particular result of the latch conditional instruction, for example, when the predicate identifies that the execution of the inner loop is after the last iteration of the inner loop, as described below, for example.
[0343] In some exemplary aspects, the compiler 160 may be configured to perform promotion and / or sinking operations repeatedly and / or iteratively, for example, to iterate over all instructions in a loop nest, for example, until substantially all instructions are moved from an outer loop to the innermost loop, as described below, for example.
[0344] For example, a loop execution scheme may be configured to provide a technical solution to support the VMP iterator to work only on perfect loop nests. For example, the loop execution scheme may result in a situation where there are no instructions between two subsequent "for" statements that form a loop, as described below, for example.
[0345] In some exemplary aspects, the compiler 160 may be configured to compile the source code 112 into the target code 115 by transforming one or more loop nests into folded loops, for example, as described below, for example.
[0346] In some exemplary aspects, one or more loop nests may be transformed into folded loops, for example, to provide a technical solution to improve the performance of the execution of a program by a target processor (e.g., a vector processor and / or any other processor), as described below, for example.
[0347] In some exemplary aspects, the compiler 160 may be configured to transform one or more identified loop nests into folded loops according to a loop folding scheme, for example, as described below, for example.
[0348] In some exemplary aspects, the compiler 160 may be configured to apply the loop folding scheme based on the result of a loop perfection scheme, as described below.
[0349] In some exemplary aspects, the compiler 160 may be configured to apply a loop folding scheme, e.g., even without applying a loop peeling scheme, e.g., when the loop peeling scheme is unnecessary, e.g., when the input loop nest includes a perfect loop nest.
[0350] In one instance, the compiler 160 may be configured to apply a loop folding scheme, e.g., to provide object code 115 configured for execution by one or more processor architectures that support only single basic block loops.
[0351] In some exemplary aspects, the compiler 160 may be configured to fold multiple individual loops of a loop nest into a single loop, which may be configured to perform, e.g., substantially all iterations of the original loop nest, as described below.
[0352] In some exemplary aspects, the execution of the folded loop may be pre-configured, e.g., to configure the advancement along the dimensions of the original loop.
[0353] In some exemplary aspects, the loop folding scheme may be configured to provide a technical solution to support the execution of the object code 115 by one or more processor architectures, including processor architectures where the computation of induction variables may be computationally expensive.
[0354] In some exemplary aspects, the compiler 160 may be configured to identify, e.g., based on the source code 112, a perfect loop nest that includes multiple nested loops, as described below.
[0355] In some exemplary aspects, the multiple nested loops may correspond to respective multiple dimensions, as described below.
[0356] In some exemplary aspects, the dimensions of the nested loops may be executed during multiple iterations of the nested loops, as described below.
[0357] In some exemplary aspects, the compiler 160 may be configured to configure the folded loop based on the loop nest, e.g., by folding multiple loop nests into a single loop based on multiple dimensions, as described below.
[0358] In some exemplary aspects, the folded loop may be configured based on, e.g., the computation of induction variables from the loop nest, as described below.
[0359] In some exemplary aspects, the folded loop may be configured according to a predicate computation mechanism that may be configured to support the computation of loop start predicates and / or loop end predicates, as described below.
[0360] In some exemplary aspects, a predicate computation mechanism may be configured to provide a technical solution to support the computation of loop start predicates and / or loop end predicates, e.g., even without computing induction variables of one or more loops in a loop nest, e.g., as described below.
[0361] In some exemplary aspects, a predicate computation mechanism may be configured to provide a technical solution to support the computation of loop start predicates and / or loop end predicates, e.g., at a processor architecture that may not support the computation of induction variables, and / or at a processor architecture where computing induction variables may be computationally expensive, e.g., when the loop is fully pre-configured.
[0362] In some exemplary aspects, a predicate computation mechanism may include a mechanism (a "hardware mask reset") that may be configured to compute loop start predicates and / or loop end predicates, e.g., based on the configuration of the AGU, e.g., as described below.
[0363] In some exemplary aspects, the hardware mask reset may configure the AGU to set a mask to true, e.g., when a loop starts or ends, e.g., as described below.
[0364] In other aspects, the predicate computation mechanism may use any other additional or alternative mechanism to compute loop start predicates and / or loop end predicates to implement.
[0365] In some exemplary aspects, compiler 160 may be configured to identify a loop nest based on source code 112, e.g., as described below.
[0366] In some exemplary aspects, compiler 160 may be configured to identify a loop nest that includes, e.g., multiple loops, the multiple loops including, e.g., at least a first loop and a second loop nested in the first loop, e.g., as described below.
[0367] In some exemplary aspects, the first loop may include at least one first loop instruction outside of the second loop, and the second loop may include one or more second loop instructions, e.g., as described below.
[0368] In some exemplary aspects, compiler 160 may be configured to transform the loop nest into a transformed loop, e.g., as described below.
[0369] In some exemplary aspects, the transformed loop may include a conditional instruction that may be based on, e.g., the first loop instruction, e.g., as described below.
[0370] In some exemplary aspects, the conditional instruction may be based on, e.g., the state of a second loop predicate that may be configured to, e.g., identify the start or the end of the second loop, e.g., as described below.
[0371] In some exemplary aspects, a transformed loop may include one or more transformed loop instructions based on one or more second loop instructions, as described below, for example.
[0372] In some exemplary aspects, one or more transformed loop instructions may include at least one (e.g., some or all) of the one or more second loop instructions, as described below, for example.
[0373] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on, for example, the compilation of source code 112 such that the target code 115 may be based on a transformed loop, as described below, for example.
[0374] In some exemplary aspects, compiler 160 may be configured to generate target code 115 that is configured to be executed, for example, by a target vector processor (e.g., vector processor 180), as described below, for example.
[0375] In some exemplary aspects, compiler 160 may be configured to generate target code 115 that is configured to be executed, for example, by a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor (e.g., processor 180).
[0376] In other aspects, compiler 160 may be configured to generate target code 115 that is configured to be executed, for example, by any other suitable type of processor.
[0377] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on, for example, source code 112 that includes Open Computing Language (OpenCL) code.
[0378] In other aspects, compiler 160 may be configured to generate target code 115 based on, for example, source code 112 that includes any other suitable type of code.
[0379] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into target code 115 according to, for example, an LLVM-based compilation scheme.
[0380] In other aspects, compiler 160 may be configured to compile source code 112 into target code 115 according to any other suitable compilation scheme.
[0381] In some exemplary aspects, compiler 160 may be configured to generate target code 115 to include AGU instructions to configure the AGU of target processor 180 to execute target code 1115, as described below, for example.
[0382] In some exemplary aspects, the compiler 160 may be configured to generate AGU instructions that are configured to cause the AGU to set an AGU mask to true based on the start or end of a second loop, for example, as described below.
[0383] In some exemplary aspects, the compiler 160 may be configured to generate conditional instructions that will be based on the AGU mask, for example, as described below.
[0384] In some exemplary aspects, the compiler 160 may be configured to generate conditional instructions to include a select operation to select between a first value and a second value based on, for example, the state of a second loop predicate, for example, as described below.
[0385] In some exemplary aspects, the compiler 160 may be configured to identify that a first loop instruction includes a first loop operation on a variable and generate a conditional instruction that includes a select operation to select between a first operation on the variable and a second operation on the variable, for example, as described below.
[0386] In some exemplary aspects, the first operation on the variable may be based on, for example, the first loop operation on the variable, for example, as described below.
[0387] In some exemplary aspects, the second operation on the variable may be based on, for example, a second loop operation on the variable in a second loop, for example, as described below.
[0388] In other aspects, the compiler 160 may be configured to generate conditional instructions to include any other suitable conditional instructions and / or operations.
[0389] In some exemplary aspects, the compiler 160 may be configured to generate a transformed loop that includes at least one second loop predicate instruction that may be configured to identify the state of a second loop predicate, for example, as described below.
[0390] In some exemplary aspects, the second loop predicate instruction may include, for example, an induction variable (IV)-independent instruction that is independent of the IV of the second loop and the IV of the first loop, for example, as described below.
[0391] In some exemplary aspects, the compiler 160 may be configured to generate a second loop predicate instruction that may be configured to retrieve the AGU mask, for example, as described below.
[0392] In some exemplary aspects, the compiler 160 may configure conditional instructions that will be based on the retrieved AGU mask, for example, as described below.
[0393] In some exemplary aspects, the compiler 160 may be configured to generate at least one second loop predicate instruction to include, for example, at least one IV-based instruction, which may be based on the IV of the second loop, as described below, for example.
[0394] In some exemplary aspects, the compiler 160 may be configured to generate at least one IV-based instruction that will be separate from the conditional instruction, as described below, for example.
[0395] In some exemplary aspects, the compiler 160 may be configured to generate a conditional instruction to include at least one IV-based instruction, as described below, for example.
[0396] In some exemplary aspects, the compiler 160 may be configured to identify that a plurality of loops includes a third loop nested within a first loop, such that a second loop is nested within the third loop, and the first loop instruction is outside the third loop, as described below, for example.
[0397] In some exemplary aspects, the compiler 160 may be configured to generate a conditional instruction based on the states of a second loop predicate and a third loop predicate, as described below, for example.
[0398] In some exemplary aspects, the third loop predicate may be configured to identify the start or the end of the third loop, as described below, for example.
[0399] In some exemplary aspects, the compiler 160 may be configured to generate a transformed loop to include at least one third loop predicate instruction, which may be configured to identify the state of the third loop predicate, as described below, for example.
[0400] In some exemplary aspects, the compiler 160 may be configured to identify that the third loop includes third loop instructions outside the second loop, as described below, for example.
[0401] In some exemplary aspects, the compiler 160 may be configured to generate a transformed loop to include another conditional instruction based on the third loop instruction, as described below, for example.
[0402] In some exemplary aspects, the another conditional instruction may be based on the state of a specific predicate used to identify the start or the end of the second loop, as described below, for example.
[0403] In some exemplary aspects, the compiler 160 may be configured to configure the second loop predicate as the specific predicate, for example, based on a determination that the first loop instruction and the third loop instruction are both pre-header instructions or both latch instructions relative to the second loop, as described below, for example.
[0404] In some exemplary aspects, the compiler 160 may be configured to, for example, based on the determination that the first instruction of the first loop instruction or the third loop instruction is a pre-header instruction relative to the second loop and the second instruction of the first loop instruction or the third loop instruction is a latch instruction relative to the second loop, configure the second loop predicate or the first predicate of a specific predicate as a loop start predicate for identifying the start of the second loop, and configure the second predicate of the second loop predicate or the specific predicate as a loop end predicate for identifying the end of the second loop, as described below, for example.
[0405] In some exemplary aspects, the compiler 160 may be configured to identify that multiple loop nests are nested at multiple nesting levels, as described below, for example.
[0406] In some exemplary aspects, the multiple nesting levels may include one or more intermediate nesting levels between a first nesting level including a first loop and a second nesting level including a second loop, as described below, for example.
[0407] In some exemplary aspects, the compiler 160 may be configured to generate conditional instructions that will be respectively based on the state of the second loop predicate and the state of one or more intermediate loop predicates corresponding to one or more intermediate nesting levels, as described below, for example.
[0408] In some exemplary aspects, the second nesting level may be the innermost nesting level among the multiple nesting levels, as described below, for example.
[0409] In some exemplary aspects, the compiler 160 may be configured to identify that at least one first loop instruction includes a pre-header instruction that will be performed before the first iteration of the second loop, as described below, for example.
[0410] In some exemplary aspects, the compiler 160 may be configured to configure the second loop predicate to include a loop start predicate that is used to identify the start of the second loop, as described below, for example.
[0411] In some exemplary aspects, the compiler 160 may be configured to configure the conditional instruction based on the pre-header instruction to be before all of the transformed loop instructions in one or more transformed loop instructions based on one or more second loop instructions, as described below, for example.
[0412] In some exemplary aspects, the compiler 160 may be configured to identify that at least one first loop instruction includes a latch instruction that will be performed after the last iteration of the second loop, as described below, for example.
[0413] In some exemplary aspects, the compiler 160 may be configured to configure a second loop predicate to include a loop end predicate that is used to identify the end of the second loop, as described below, for example.
[0414] In some exemplary aspects, the compiler 160 may be configured to configure a conditional instruction based on a latch instruction to be after all transformed loop instructions in one or more transformed loop instructions based on one or more second loop instructions, as described below, for example.
[0415] In some exemplary aspects, the compiler 160 may be configured to generate a transformed loop that includes a perfect flat loop in which all computational operations of the loop nest are implemented in the innermost loop, as described below, for example.
[0416] In some exemplary aspects, the compiler 160 may be configured to generate a transformed loop that includes a fully folded loop that includes only a single basic block loop based on multiple loops, as described below, for example.
[0417] In some exemplary aspects, the compiler 160 may compile the source code 112 of a program to be executed by a target processor (e.g., processor 180), as described below.
[0418] For example, the compiler 160 may identify a loop nest based on the source code 112, where the loop nest includes an outer loop along a dimension based on the value half_height and an inner loop along a dimension based on the value half_width, as follows:
[0419]
[0420] Example (4)
[0421] For example, as shown in Example 4, the outer loop may include a pre-header instruction, e.g., unsigned short left = 0, which may reside before the header of the inner loop.
[0422] For example, as shown in Example 4, the pre-header instruction may be executed each time before the start of the execution of the inner loop.
[0423] In some exemplary aspects, the compiler 160 may be configured to sink a prefetch instruction into the inner loop, e.g., to generate a perfect loop nest, as described below, for example.
[0424] For example, the compiler 160 may be configured to move a pre-header instruction into the inner loop and transform the pre-header instruction into a conditional pre-header instruction, as follows:
[0425]
[0426]
[0427] Example (5)
[0428] In some exemplary aspects, as shown in Example 5, a pre-header instruction (unsigned short left = 0) can be transformed into a conditional pre-header instruction (unsigned short left = predicate? 0 : right).
[0429] In some exemplary aspects, as shown in Example 5, the conditional pre-header instruction can be based on a predicate that can indicate, identify, confirm, predict, and / or assert the start of an inner loop.
[0430] In some exemplary aspects, as shown in Example 5, the conditional pre-header instruction can be configured such that the result of the conditional pre-header instruction can be equivalent to the execution of the pre-header instruction (unsigned short left = 0), e.g., only when the predicate "xx == 0?" is true, e.g., only when the inner loop starts to execute.
[0431] In some exemplary aspects, the compiler 160 can be configured to generate the target code 115 based on the perfect loop nesting of Example 5.
[0432] In some exemplary aspects, the compiler 160 can be configured to transform the perfect loop nesting of Example 5 into a folded loop, e.g., along dimensions that can be based on, e.g., the value half_height and the value half_width, e.g., as follows:
[0433]
[0434] Example (6)
[0435] In some exemplary aspects, as shown in Example 6, the folded loop can include a single block that can execute, e.g., according to the perfect loop nesting of Example 5.
[0436] In some exemplary aspects, as shown in Example 6, the folded loop can include a conditional pre-header instruction (unsigned short left = predicate? 0 : right) that can be based on a predicate.
[0437] In some exemplary aspects, the induction variables "y" and "xx" of the loop of Example 5 may not be explicitly calculated in the folded loop of Example 6, e.g., because the folded loop may not include any inner loops to advance each of the separate induction variables "y" and "xx".
[0438] In some exemplary aspects, the compiler 160 may be configured to generate the folded loop of Instance 6, for example, in a manner that supports the calculation of a predicate (predicate=(is this the start of xx loop?)), for example, as follows:
[0439]
[0440] Instance (7)
[0441] In some exemplary aspects, as shown in Instance 7, the folded loop may calculate a predicate (predicate=(is this the start of xx loop?)) based on induction variables "y" and "xx", for example.
[0442] In some exemplary aspects, the compiler 160 may be configured to generate the folded loop of Instance 6, for example, in a manner that supports the calculation of a predicate (predicate=(is this the start of xx loop?)), for example, as follows:
[0443]
[0444]
[0445] Instance (8)
[0446] In some exemplary aspects, as shown in Instance 8, the folded loop may calculate a predicate (predicate=(is this the start of xx loop?)) based on induction variables "y" and "xx", for example.
[0447] In some exemplary aspects, the compiler 160 may be configured to generate the folded loop of Instance 6, for example, in a manner that supports the calculation of a predicate, for example, even without relying on induction variables "y" and "xx", for example, as described below.
[0448] In some exemplary aspects, the compiler 160 may be configured to generate the folded loop of Instance 6 based on a predicate calculation mechanism that may be configured to support the calculation of a loop start predicate and / or a loop end predicate, for example, even without maintaining and / or calculating induction variables "y" and / or "xx", for example, as described below.
[0449] In some exemplary aspects, the compiler 160 may be configured to generate the folded loop of Instance 6 based on a hardware mask reset mechanism that may be configured to calculate a loop start predicate and / or a loop end predicate based on the configuration of the AGU of the target processor 180, for example, as described below.
[0450] In some exemplary aspects, the compiler 160 may be configured to generate the folded loop of Instance 6, for example, by configuring the AGU of the target processor 180, for example, such that the AGU sets a mask to true when the loop starts or ends, for example, as described below.
[0451] In some exemplary aspects, the compiler 160 may generate the target code 115 to include AGU instructions to configure the AGU of the target processor 180, for example, before the execution of the loop, for example, as described below.
[0452] In some exemplary aspects, the compiler 160 may generate AGU instructions to configure the AGU to set a mask to true, for example, when the loop starts or ends, for example, in each iteration of each loop in a loop nest.
[0453] For example, the compiler 160 may be configured to generate the target code 115 based on the folded loop of Instance 6, for example, by configuring the target code 115 to include AGU instructions to be executed by the AGU of the target processor 180, for example, as follows:
[0454]
[0455]
[0456] Instance (9)
[0457] For example, according to Instance 9, the first dimension (dimension "0") of the AGU may be set, for example, based on the induction variable "xx".
[0458] For example, according to Instance 9, the second dimension (dimension "1") of the AGU may be set, for example, based on the induction variable "y".
[0459] For example, according to Instance 9, the AGU may be used to calculate the loop start predicate for xx = 0.
[0460] For example, according to Instance 9, the AGU instructions may be configured to set the first dimension (dimension 0) of the AGU to start from a value (e.g., 0) (which may be based on the start value of the loop on the variable "xx") and end at a value (e.g., half_width) (which may be based on the end value of the loop on the variable "xx").
[0461] In some exemplary aspects, the compiler 160 may compile the source code 112 of a program to be executed by a target processor (e.g., processor 180), as described below.
[0462] For example, the compiler 160 may identify a loop nest, e.g., based on the source code 112, that includes an outer loop along a dimension based on the value zcount, an intermediate (nested) loop along a dimension based on the value ycount, and an inner loop along a dimension based on the variable xcount, e.g., as follows:
[0463]
[0464] Instance (10)
[0465] In some exemplary aspects, the compiler 160 may be configured to transform the loop nest of instance 10 into a perfect loop nest, e.g., by performing a plurality of operations according to a loop perfection scheme, e.g., as described below.
[0466] In some exemplary aspects, the compiler 160 may be configured to perform a first operation to sink an outer instruction from an outer loop (e.g., the outermost loop) into an inner loop having a loop level that is one level lower than the loop level of the outer loop, e.g., as described below.
[0467] In some exemplary aspects, the compiler 160 may be configured to perform additional subsequent operations, e.g., where each subsequent operation may be configured to sink an outer instruction from an inner loop into the next inner loop having a loop level that is one level lower than the loop level of the inner loop, e.g., as described below.
[0468] In some exemplary aspects, the compiler 160 may be configured to perform additional subsequent operations, e.g., until the innermost loop level is reached, e.g., as described below.
[0469] In some exemplary aspects, as shown in instance 10, the loop nest may include a first preheader instruction, e.g., Sum = 0, that may be before the header of the outer loop.
[0470] In some exemplary aspects, as shown in instance 10, the first preheader instruction may be executed, e.g., before the first iteration of the outer loop.
[0471] In some exemplary aspects, the compiler 160 may be configured to sink the first preheader instruction into the outer loop, e.g., by: moving the first preheader instruction into the outer loop and transforming the first preheader instruction into a conditional preheader instruction, e.g., as follows:
[0472]
[0473] Instance (11)
[0474] In some exemplary aspects, as shown in Example 11, the first pre-header instruction sum = 0 can be transformed into a conditional pre-header instruction Sum = (z == 0? 0 : Sum), which can be based on the predicate z == 0, and the predicate can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of the outer loop.
[0475] In some exemplary aspects, as shown in Example 11, the outer loop can include two pre-header instructions, for example, the conditional pre-header instruction Sum = (z == 0? 0 : Sum) and a second pre-header instruction (e.g., unsigned short left = 0), and the second pre-header instruction can be located before the head of the nested loop.
[0476] In some exemplary aspects, as shown in Example 11, the outer loop can include a latch instruction, for example, Sum++, and the latch instruction can be after the nested loop.
[0477] In some exemplary aspects, the compiler 160 can be configured to sink two pre-header instructions into the nested loop, for example, by moving the two pre-header instructions into the nested loop and transforming the two pre-header instructions into two corresponding conditional pre-header instructions, as described below.
[0478] In some exemplary aspects, the compiler 160 can be configured to lift the latch instruction into the nested loop, for example, by moving the latch instruction into the nested loop and transforming the latch instruction into a conditional latch instruction.
[0479] For example, the compiler 160 can be configured to sink two pre-header instructions into the nested loop and lift the latch instruction into the nested loop, as follows:
[0480]
[0481] Example (12)
[0482] In some exemplary aspects, as shown in Example 12, the conditional pre-header instruction Sum = (z == 0? 0 : Sum) can be transformed into a conditional pre-header instruction Sum = (z == 0 && y == 0? 0 : Sum) based on the following predicates: the predicate z == 0, which can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of the outer loop; and the predicate y == 0, which can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of the nested loop.
[0483] In some exemplary aspects, as shown in Example 12, the second pre-header instruction unsigned short left = 0 can be transformed into a conditional pre-header instruction left = (y == 0? 0 : right) based on the predicate y == 0, and the predicate can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of a nested loop.
[0484] In some exemplary aspects, as shown in Example 12, the latch instruction Sum++ can be transformed into a conditional latch instruction, e.g., Sum = (y == ycount - 1? Sum + 1 : Sum), based on the predicate y == ycount, and the predicate can be configured to indicate, identify, confirm, predict, and / or assert the last iteration of a nested loop.
[0485] In some exemplary aspects, as shown in Example 12, a nested loop can include three outer instructions relative to the inner loop, e.g., including two pre-header instructions and one latch instruction.
[0486] For example, the two pre-header instructions can include the conditional pre-header instruction Sum = (z == 0 && y == 0? 0 : Sum), and the conditional pre-header instruction left = (y == 0? 0 : right).
[0487] For example, the latch instruction can include the conditional latch instruction Sum = (y == ycount - 1? Sum + 1 : Sum).
[0488] In some exemplary aspects, the compiler 160 can be configured to sink the two pre-header instructions into the inner loop, e.g., by moving the two pre-header instructions into the inner loop and transforming the two pre-header instructions into two conditional pre-header instructions, respectively.
[0489] In some exemplary aspects, the compiler 160 can be configured to lift the latch instruction into the inner loop, e.g., by moving the latch instruction into the inner loop and transforming the latch instruction into a conditional latch instruction.
[0490] For example, the compiler 160 can be configured to sink the two pre-header instructions into the inner loop and lift the latch instruction into the inner loop, e.g., as follows:
[0491]
[0492]
[0493] Example (13)
[0494] In some exemplary aspects, as shown in Example 13, a conditional pre-header instruction Sum = (z == 0 && y == 0? 0 : Sum) can be transformed into a conditional pre-header instruction (e.g., Sum = (z == 0 && y == 0 && xx == 0? 0 : Sum)) based on the following predicates: the predicate z == 0, which can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of an outer loop; the predicate y == 0, which can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of a nested loop; and the predicate xx == 0, which can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of an inner loop.
[0495] In some exemplary aspects, as shown in Example 13, a conditional pre-header instruction left = (y == 0? 0 : right) can be transformed into a conditional pre-header instruction (e.g., left = (y == 0 && xx == 0? 0 : right)) based on the following predicates: the predicate y == 0, which can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of a nested loop; and the predicate xx == 0, which can be configured to indicate, identify, confirm, predict, and / or assert the first iteration of an inner loop.
[0496] In some exemplary aspects, as shown in Example 13, a conditional latch instruction Sum = (y == ycount – 1? Sum + 1 : Sum) can be transformed into a conditional latch instruction (e.g., Sum = (y == ycount – 1 && xx == xcount – 1? Sum + 1 : Sum)) based on, for example, the following predicates: the predicate y == ycount – 1, which can be configured to indicate, identify, confirm, predict, and / or assert the last iteration of a nested loop; and the predicate xx == xcount – 1, which can be configured to indicate, identify, confirm, predict, and / or assert the last iteration of an inner loop.
[0497] In some exemplary aspects, as shown in Example 13, a loop nest can be transformed into a perfect loop nest, e.g., because all instructions can be performed in the innermost loop of the loop nest, e.g., after the "for statement" of the inner loop.
[0498] In some exemplary aspects, the compiler 160 can be configured to generate the target code 115 based on, for example, the code of Example 13.
[0499] In some exemplary aspects, the compiler 160 can be configured to fold the perfect loop nest of Example 13 into a folded loop according to, for example, a loop folding scheme, as described above.
[0500] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115, for example, based on the folded loop, which may be based on the code of instance 13.
[0501] Reference Figure 4 , which schematically illustrates a method of compiling code for a processor. For example, Figure 4 one or more operations of the method of Figure 1 may be performed by: a system, for example, system 100 ( Figure 1 ); a device, for example, device 102 ( Figure 1 ); a server, for example, server 170 ( Figure 1 ); and / or a compiler, for example, compiler 160 ( Figure 2 ) and / or compiler 200 (
[0502] In some exemplary aspects, as indicated at block 402, the method may include identifying the instructions of the outer loop of the loop nest. For example, the compiler 160 ( Figure 1 ) may generate a list (WorkList) of one or more identified instructions of the outer loop of the loop nest based on the source code 112 ( Figure 1 ), for example, as described above.
[0503] In some exemplary aspects, as indicated at block 404, the method may include initializing the currently analyzed loop level (CurrLoopLevel) of the loop nest, for example, based on the total number of loop dimensions (loop levels) of the loop nest. For example, the compiler 160 ( Figure 1 ) may initialize the currently analyzed loop level of the loop nest to the highest loop level, for example, the highest loop level corresponding to the outermost loop of the loop nest.
[0504] In some exemplary aspects, as indicated at block 406, the method may include iterating over the loop levels of the loop nest, for example, from the highest loop level to the lowest loop level until the currently analyzed loop level includes the loop level of the innermost loop of the loop nest. For example, the compiler 160 ( Figure 1 ) may iterate over the loop levels of the loop nest.
[0505] In some exemplary aspects, as indicated at block 408, the method may include determining the loop nest iteration to be terminated when the currently analyzed loop level includes the level of the innermost loop of the loop nest. For example, the compiler 160 ( Figure 1 ) may terminate the loop nest iteration to create a perfect loop nest of instance 13, for example, when the loop level includes the innermost loop of the loop nest of instance 10, as described above.
[0506] In some exemplary aspects, as indicated at block 409, the method may include determining whether the currently analyzed loop includes an outer instruction that is to be moved to an inner loop. For example, compiler 160( Figure 1 ) may determine whether the currently analyzed loop includes an outer instruction from the WorkList, as described above, for example.
[0507] In some exemplary aspects, as indicated at block 410, the method may include designating one or more identified outer instructions of the currently analyzed loop that are to be moved to an inner loop. For example, compiler 160( Figure 1 ) may designate one or more identified outer instructions in the WorkList that are to be moved to an inner loop, as described above, for example.
[0508] In some exemplary aspects, as indicated at block 412, the method may include determining whether the identified outer instructions include a preheader instruction or a latch instruction. For example, compiler 160( Figure 1 ) may determine whether the identified outer instructions include a preheader instruction or a latch instruction, as described above, for example.
[0509] In some exemplary aspects, as indicated at block 414, the method may include sinking the preheader instruction into an inner loop having a loop level that is one level lower than the loop level of the outer loop. For example, compiler 160( Figure 1 ) may sink the preheader instruction into the inner loop, as described above, for example.
[0510] In some exemplary aspects, as indicated at block 415, the method may include promoting the latch instruction into an inner loop having a loop level that is one level lower than the loop level of the outer loop. For example, compiler 160( Figure 1 ) may promote the latch instruction into the inner loop, as described above, for example.
[0511] In some exemplary aspects, as indicated by arrow 416, the method may include returning to repeat the operations beginning at block 409 with respect to the next identified outer instruction of the outer loop, for example, until the last outer instruction of the outer loop is reached.
[0512] In some exemplary aspects, as indicated at block 418, the method may include returning to repeat the operations beginning at block 406 for the next outer loop (the next outer loop having a loop level that is one level lower than the loop level of the currently analyzed outer loop), for example, until the innermost loop of the loop nest is reached.
[0513] In some exemplary aspects, Figure 4 one or more operations of the method may be implemented to configure select instructions at a loop level (e.g., each loop level) based on start / end predicates for the loop level, for example.
[0514] For example, these select instructions can be optimized by paying attention to how the sink values are used, for example.
[0515] For example, if the sink values are not used in one or more loop levels, one or more selects / predications can be optimized away.
[0516] In some exemplary aspects, a compiler (e.g., compiler 160( Figure 1 )) can be configured to compile the code of a loop according to one or more operations of a loop compilation algorithm (e.g., based on the code of a loop of the source code 112( Figure 1 ), for example, as described below.
[0517] In some exemplary aspects, a loop compilation algorithm can be configured, for example, relative to code using Phi instructions according to an LLVM-based compilation scheme, for example, as described below. For example, Phi instructions can be utilized to implement φ nodes in a static single assignment (SSA) graph representing a function according to an LLVM-based compilation scheme. In other aspects, a loop compilation algorithm can be configured relative to code using any other additional or alternative type of instructions and / or according to any other suitable compilation scheme.
[0518] In some exemplary aspects, a loop compilation algorithm can include one or more of the following operations:
[0519] 1. Aggregate initial instructions for sinking / lifting:
[0520] · Initialize a list to all store instructions in the outermost loop
[0521] · Add all store operands to the list, for example, while excluding address calculations for AGUifiable stores
[0522] · Remove all Phi instructions (Phis) from the list, for example, because they will be sunk as users (via the backedge)
[0523] 2. For each loop level, starting from the outermost loop and going to the innermost loop:
[0524] · For each preheader instruction, sink (2.1) the instruction. If the instruction also requires predication (2.2), add it to the predication list
[0525] · For each latch instruction, lift (2.3) the instruction. If the instruction also requires predication (2.4), add it to the predication list
[0526] · Add predication (2.5) for each instruction in the predication list
[0527] 2.1 - Sinking Instruction
[0528] The sinking instruction can be completed by moving the instruction from the current loop - level pre - header to the next inner - loop - level pre - header. If the instruction is a store, we can convert it to a masked_store, e.g., based on the FirstIter predicate. If the instruction also feeds a PHI, a first - iteration selection (2.6) can be added based on the level of the sinking instruction.
[0529] 2.2 - Sinking Requires Predication
[0530] The sunk instruction requires predication, e.g., if its value is used (and subsequently stored) in a latch. This is because we may need to ensure that its value remains valid (and not overwritten) until the latch. The operation of adding predication to the sunk instruction is also called “rotation”.
[0531] 2.3 - Hoisting Instruction
[0532] The hoisting instruction can be completed by moving the instruction from the current loop - level latch to the next inner - loop - level latch.
[0533] If the instruction is a store, we can convert it to a masked_store, e.g., based on the LastIter predicate
[0534] 2.4 - Hoisting Requires Predication
[0535] The hoisted instruction may require predication, e.g., if it feeds a PHI.
[0536] Such an instruction may not be hoisted as is and may (e.g., should) only be “executed” in the latch.
[0537] 2.5 - Adding Predication
[0538] Adding predication to the sunk instruction (or “rotating it”) can be done, for example, by adding a first - iteration selection (2.6) between the instruction value and its value from the previous iteration.
[0539] Adding predication to the hoisted instruction can be done by adding a last - iteration selection (2.6) between the instruction value and the PHI it feeds.
[0540] 2.6 - First Iteration / Last Iteration Selection
[0541] The first - iteration selection can be a selection that predicates whether the loop has just entered a certain loop level.
[0542] The last iteration selection can be a selection that predicates whether the loop has just exited a certain loop level.
[0543] There are several ways to generate such predicates, for example:
[0544] · Generate the IV and compare it with the initial / final value, as described above
[0545] · Use hardware support, such as a MaskReset component, as in the algorithm (1) described above
[0546] In one instance, one or more (e.g., some or all) of the operations in the operation of algorithm 1 can be applied to the following loop:
[0547] L2:
[0548] %init.val = 5
[0549] L1:
[0550] %sum = phi(%init.val, %sum.inc)
[0551] L0: ...
[0553] %sum.inc = %sum + 1
[0554] store %sum.inc
[0555] Instance (14)
[0556] In one instance, the loop of instance 14 can be folded using induction generation, e.g., by applying sinking / lifting at L2. For example, when sinking (2.1) %init.val, it may be noted that it feeds a Phi in L1. Accordingly, a first iteration selection can be added.
[0557] For example, the store operation can be masked when lifting it (2.3).
[0558] For example, the following loop nest can be determined after, e.g., L2 sinking / lifting:
[0559] L2:
[0560] L1:
[0561] %sum = phi(undef, %sum.inc)
[0562] %mask1 = cmp %induction.L1 == 0 / / L1_start
[0563] %select1 = select(%mask1, %init.val, %sum)
[0564] L0: ...
[0566] %sum.inc = %select1 + 1
[0567] %store_mask = cmp %induction.L2 == %count.L2 / / L2_end
[0568] masked_store %store_mask, %sum.inc
[0569] Example (15)
[0570] For example, when sinking (2.1) L1, it can be recognized that %select1 feeds an instruction in the latch (%sum.inc), so it can be added to the predicated list according to (2.2), for example.
[0571] For example, when lifting (2.3) L1, it can be recognized that %sum.inc feeds PHI (%sum), so it can be added to the predicated list according to (2.4).
[0572] For example, the following perfect loop can be determined after L1 sinking / lifting, for example:
[0573] L2:
[0574] L1:
[0575] L0:
[0576] %sum = phi(undef, %sum.inc)
[0577] %mask1 = cmp %induction.L1 == 0 / / L1_start
[0578] %select1 = select(%mask1, %init.val, %sum) ...
[0580] %sum.inc = %select1 + 1
[0581] %store_mask = cmp %induction.L2 == %count.L2 / / L2_end
[0582] masked_store %store_mask, %sum.inc
[0583] Instance (16)
[0584] For example, according to the loop of Instance 16, all instructions can be in the innermost loop. For example, predicatization can be added to two instructions, for example, as follows:
[0585] L2:
[0586] L1:
[0587] L0:
[0588] %sum = phi(undef, %sum.inc)
[0589] %prev.select2 = phi(undef, %select2)
[0590] %prev.select3 = phi(undef, %select3)
[0591] %mask1 = cmp %induction.L1 == 0 / / L1_start
[0592] %select1 = select(%mask1, %init.val, %sum)
[0593] %mask2 = cmp %induction.L0 == 0 / / L0_start
[0594] %select2 = select(%mask2, %select1, %prev.select2) ...
[0596] %sum.inc = %select2 + 1
[0597] %mask3 = cmp %induction.L1 == %count.L1 / / L1_end
[0598] %select3 = select(%mask3, %sum.inc, %prev.select3)
[0599] %store_mask = cmp %induction.L2 == %count.L2 / / L2_end
[0600] masked_store %store_mask, %select3
[0601] Instance (17)
[0602] For example, induction can be generated as follows:
[0603] %induction.L1.init = phi(undef, %induction.L1)
[0604] %induction.L2.init = phi(undef, %induction.L2)
[0605] %induction.L3.init = phi(undef, %induction.L3)
[0606] %add1 = add %induction.L1.init, 1
[0607] %pred.L1 = cmp %induction.L1.init == %count.L1
[0608] %induction.L1 = select %pred.L1, 0, %add1
[0609] %add2 = add %induction.L2.init, 1
[0610] %induction.L2.init2 = select %pred.L1, %add2, %induction.L2.init
[0611] %pred.L2 = cmp %induction.L2.init2 == %count.L2
[0612] %induction.L2 = select %pred.L2, 0, %induction.L2.init2
[0613] %add3 = add %induction.L3.init, 1
[0614] %induction.L3.init3 = select %pred.L2, %add3, %induction.L3.init
[0615] %pred.L3 = cmp %induction.L3.init3 == %count.L3
[0616] %induction.L3 = select %pred.L3, 0, %induction.L3.init2
[0617] Example (18)
[0618] In one example, the loop of Example 17 can be folded, for example, using a mask reset mechanism as described below.
[0619] For example, sinking / lifting can be implemented at L2. For example, when sinking (2.1)%init.val, it can be recognized that it feeds a Phi in L1, so a first iteration mask reset selection can be added.
[0620] For example, a store operation can be masked, for example, when lifting it (2.3).
[0621] For example, after sinking / lifting at L2, the following loop can be reached:
[0622] L2:
[0623] L1:
[0624] %sum = phi(undef, %sum.inc)
[0625] %mask1 = mask.reset / / L1_start
[0626] %select1 = select(%mask1, %init.val, %sum)
[0627] L0: ...
[0629] %sum.inc = %select1 + 1
[0630] %store_mask = mask.reset / / L2_end
[0631] masked_store %store_mask, %sum.inc
[0632] Example (19)
[0633] For example, when sinking (2.1) L1, it can be recognized that %select1 feeds an instruction in the latch (%sum.inc), so it can be added to the predicated list, for example, according to (2.2).
[0634] For example, when lifting (2.3) L1, it can be recognized that %sum.inc feeds a PHI (%sum), so it can be added to the predicated list according to (2.4).
[0635] For example, the following perfect loop can be determined, for example, after sinking / lifting at L1:
[0636] L2:
[0637] L1:
[0638] L0:
[0639] %sum = phi(undef, %sum.inc)
[0640] %mask1 = mask.reset / / L1_start
[0641] %select1 = select(%mask1, %init.val, %sum) ...
[0643] %sum.inc = %select1 + 1
[0644] %store_mask = mask.reset / / L2_end
[0645] masked_store %store_mask, %sum.inc
[0646] Example (20)
[0647] For example, all instructions in the loop of Example 20 can be in the innermost loop.
[0648] For example, predication can be added to two instructions, for example, as follows:
[0649] L2:
[0650] L1:
[0651] L0:
[0652] %sum = phi(undef, %sum.inc)
[0653] %prev.select2 = phi(undef, %select2)
[0654] %prev.select3 = phi(undef, %select3)
[0655] %mask1 = mask.reset / / L1_start
[0656] %select1 = select(%mask1, %init.val, %sum)
[0657] %mask2 = mask.reset / / L0_start
[0658] %select2 = select(%mask2, %select1, %prev.select2) ...
[0660] %sum.inc = %select2 + 1
[0661] %mask3 = mask.reset / / L1_end
[0662] %select3 = select(%mask3, %sum.inc, %prev.select3)
[0663] %store_mask = mask.reset / / L2_end
[0664] masked_store %store_mask, %select3
[0665] Example (21)
[0666] Reference Figure 5 , which schematically illustrates a method for compiling code for a processor. For example, Figure 5 one or more operations of the method may be performed by: a system, such as system 100( Figure 1 ); a device, such as device 102( Figure 1 ); a server, such as server 170( Figure 1 ); and / or a compiler, such as compiler 160( Figure 1 ) and / or compiler 200( Figure 2 ).
[0667] In some exemplary aspects, as indicated at block 502, the method may include identifying a loop nest based on source code, the loop nest including a plurality of loops, the plurality of loops including at least a first loop and a second loop nested within the first loop. For example, the first loop may include at least one first loop instruction outside the second loop, and the second loop may include one or more second loop instructions. For example, compiler 160( Figure 1 ) may be configured to identify the loop nest, for example, based on source code 112( Figure 1 ), as described above.
[0668] In some exemplary aspects, as indicated at block 504, the method may include transforming a loop nest into a transformed loop. For example, the transformed loop may include a conditional instruction based on a first loop instruction. For example, the conditional instruction may be based on the state of a second loop predicate. For example, the second loop predicate may identify the start or the end of a second loop. For example, the transformed loop may include one or more transformed loop instructions based on one or more second loop instructions. For example, compiler 160( Figure 1 ) may be configured to transform a loop nest into a transformed loop, e.g., as described above.
[0669] In some exemplary aspects, as indicated at block 506, the method may include generating target code based on compilation of source code, where the target code is based on the transformed loop. For example, compiler 160( Figure 1 ) may be configured to generate target code 115( Figure 1 ) based on the transformed loop, e.g., as described above.
[0670] Referring Figure 6 , which schematically illustrates a manufactured product 600 in accordance with some exemplary aspects. Product 600 may include one or more tangible computer-readable (“machine-readable”) non-transitory storage media 602, which may include, for example, computer-executable instructions implemented by logic 604 that are operable to cause at least one computer processor to be capable of implementing one or more operations at device 102( Figure 1 ), server 170( Figure 1 ) and / or compiler 160( Figure 1 ) to cause device 102( Figure 1 ), server 170( Figure 1 ) and / or compiler 160( Figure 1 ) to perform, trigger, and / or implement one or more operations and / or functionality, and / or to perform, trigger, and / or implement one or more operations and / or functionality described in reference Figures 1 to 5 and / or one or more operations described herein. The phrases “non-transitory machine-readable medium” and “computer-readable non-transitory storage medium” are intended to include all computer-readable media, with the sole exception of transitory propagating signals.
[0671] In some exemplary aspects, product 600 and / or machine-readable storage medium 602 may include one or more types of computer-readable storage media capable of storing data, including volatile memory, non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. For example, machine-readable storage medium 602 may include RAM, DRAM, double data rate DRAM (DDR-DRAM), SDRAM, static RAM (SRAM), ROM, programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory (e.g., NOR or NAND flash memory), content addressable memory (CAM), polymer memory, phase change memory, ferroelectric memory, silicon oxide nitride oxide (SONOS) memory, disk, hard disk drive, etc. The computer-readable storage medium may include any suitable medium involved in downloading or transferring a computer program from a remote computer to a requesting computer via a communication link (e.g., a modem, radio, or network connection), where the computer program is carried by a data signal embedded in a carrier wave or other propagated medium.
[0672] In some exemplary aspects, logic 604 may include instructions, data, and / or code that, if executed by a machine, may cause the machine to perform the methods, processes, and / or operations described herein. The machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, etc., and may be implemented using any suitable combination of hardware, software, firmware, etc.
[0673] In some exemplary aspects, logic 604 may include or may be implemented as software, software modules, applications, programs, subroutines, instructions, instruction sets, computing code, words, values, symbols, etc. The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. The instructions may be implemented according to a predefined computer language, manner, or syntax for instructing a processor to perform a particular function. The instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language, machine code, etc.
[0674] Example
[0675] The following examples relate to further aspects.
[0676] Example 1 includes a product that includes one or more tangible computer-readable non-transitory storage media that contain computer-executable instructions that are operable to cause at least one processor to cause a compiler, when executed by the at least one processor, to: identify a loop nest based on source code, the loop nest including a plurality of loops, the plurality of loops including at least a first loop and a second loop nested within the first loop, wherein the first loop includes at least one first loop instruction outside of the second loop, and wherein the second loop includes one or more second loop instructions; transform the loop nest into a transformed loop, the transformed loop including: a conditional instruction based on the first loop instruction, the conditional instruction based on the state of a second loop predicate, wherein the second loop predicate is used to identify the start or the end of the second loop; and one or more transformed loop instructions based on the one or more second loop instructions; and generate target code based on the compilation of the source code, wherein the target code is based on the transformed loop.
[0677] Example 2 includes the subject matter of Example 1, and optionally wherein the transformed loop includes at least one second loop predicate instruction configured to identify the state of the second loop predicate.
[0678] Example 3 includes the subject matter of Example 2, and optionally wherein the second loop predicate instruction includes an induction variable (IV)-independent instruction that is independent of the IV of the second loop and the IV of the first loop.
[0679] Example 4 includes the subject matter of Example 2 or 3, and optionally wherein the target code includes one or more address generation unit (AGU) instructions to configure the AGU of the target processor to execute the target code, the AGU instructions being configured to cause the AGU to set an AGU mask to true based on the start or the end of the second loop, wherein the second loop predicate instruction is configured to retrieve the AGU mask, and wherein the conditional instruction is based on the AGU mask.
[0680] Example 5 includes the subject matter of Example 2, and optionally wherein at least one second loop predicate instruction includes at least one induction variable (IV)-based instruction that is based on the IV of the second loop.
[0681] Example 6 includes the subject matter of Example 5, and optionally wherein at least one IV-based instruction is separate from the conditional instruction.
[0682] Example 7 includes the subject matter of Example 5, and optionally wherein the conditional instruction includes at least one IV-based instruction.
[0683] Example 8 includes the subject matter according to any one of Examples 1 to 7, and optionally wherein a plurality of loops include a third loop nested within a first loop, a second loop nested within the third loop, the first loop instructions being outside the third loop, wherein the conditional instruction is based on the state of a second loop predicate and the state of a third loop predicate, and wherein the third loop predicate is used to identify the start or the end of the third loop.
[0684] Example 9 includes the subject matter according to Example 8, and optionally wherein the transformed loop includes at least one third loop predicate instruction configured to identify the state of the third loop predicate.
[0685] Example 10 includes the subject matter according to Example 8 or 9, and optionally wherein the third loop includes third loop instructions outside the second loop, and wherein the transformed loop includes another conditional instruction based on the third loop instructions, and wherein the another conditional instruction is based on the state of a specific predicate used to identify the start or the end of the second loop.
[0686] Example 11 includes the subject matter according to Example 10, and optionally wherein the instruction, when executed, causes the compiler to configure the second loop predicate as the specific predicate based on a determination that the first loop instructions and the third loop instructions are both preheader instructions or both latch instructions relative to the second loop.
[0687] Example 12 includes the subject matter according to Example 10, and optionally wherein the instruction, when executed, causes the compiler to configure a first predicate of the second loop predicate or the specific predicate as a loop start predicate for identifying the start of the second loop and a second predicate of the second loop predicate or the specific predicate as a loop end predicate for identifying the end of the second loop based on a determination that a first instruction of the first loop instructions or the third loop instructions is a preheader instruction relative to the second loop and a second instruction of the first loop instructions or the third loop instructions is a latch instruction relative to the second loop.
[0688] Example 13 includes the subject matter according to any one of Examples 1 to 12, and optionally wherein a plurality of loops are nested at a plurality of nesting levels, and wherein the plurality of nesting levels includes one or more intermediate nesting levels between a first nesting level including the first loop and a second nesting level including the second loop, and wherein the conditional instruction is based on the state of the second loop predicate and the states of one or more intermediate loop predicates corresponding to the one or more intermediate nesting levels, respectively.
[0689] Example 14 includes the subject matter according to Example 13, and optionally wherein the second nesting level is the innermost nesting level among the plurality of nesting levels.
[0690] Example 15 includes the subject matter according to any one of Examples 1 to 14, and optionally wherein at least one first loop instruction includes a pre-header instruction to be performed before the first iteration of a second loop, wherein the second loop predicate includes a loop start predicate for identifying the start of the second loop.
[0691] Example 16 includes the subject matter according to Example 15, and optionally wherein a conditional instruction is before all of the transformed loop instructions in one or more transformed loop instructions based on one or more second loop instructions.
[0692] Example 17 includes the subject matter according to any one of Examples 1 to 16, and optionally wherein at least one first loop instruction includes a latch instruction to be performed after the last iteration of a second loop, wherein the second loop predicate includes a loop end predicate for identifying the end of the second loop.
[0693] Example 18 includes the subject matter according to Example 17, and optionally wherein a conditional instruction is after all of the transformed loop instructions in one or more transformed loop instructions based on one or more second loop instructions.
[0694] Example 19 includes the subject matter according to any one of Examples 1 to 18, and optionally wherein the transformed loop includes a perfect flat loop in which all computational operations of the loop nest are implemented in the innermost loop.
[0695] Example 20 includes the subject matter according to any one of Examples 1 to 19, and optionally wherein the transformed loop includes a fully folded loop that includes only a single basic block loop based on a plurality of loops.
[0696] Example 21 includes the subject matter according to any one of Examples 1 to 20, and optionally wherein the target code includes one or more address generation unit (AGU) instructions to configure the AGU of the target processor to execute the target code, the AGU instructions being configured to cause the AGU to set an AGU mask to true based on the start or the end of a second loop, wherein the conditional instruction is based on the AGU mask.
[0697] Example 22 includes the subject matter according to any one of Examples 1 to 21, and optionally wherein the first loop instruction includes a first loop operation on a variable, wherein the conditional instruction includes a selection operation to select between a first operation on the variable and a second operation on the variable, wherein the first operation on the variable is based on the first loop operation on the variable.
[0698] Example 23 includes the subject matter according to Example 22, and optionally wherein the second operation on the variable is based on a second loop operation on the variable in a second loop.
[0699] Example 24 includes the subject matter according to any one of Examples 1 to 23, and optionally wherein the conditional instruction includes a selection operation to select between a first value and a second value based on the state of a second loop predicate.
[0700] Example 25 includes the subject matter according to any one of Examples 1 to 24, and optionally wherein one or more transformed loop instructions include at least one second loop instruction of one or more second loop instructions.
[0701] Example 26 includes the subject matter according to any one of Examples 1 to 25, and optionally wherein the source code includes Open Computing Language (OpenCL) code.
[0702] Example 27 includes the subject matter according to any one of Examples 1 to 26, and optionally wherein the computer-executable instructions, when executed, cause a compiler to compile the source code into target code according to an LLVM-based compilation scheme.
[0703] Example 28 includes the subject matter according to any one of Examples 1 to 27, and optionally wherein the target code is configured to be executed by a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor.
[0704] Example 29 includes the subject matter according to any one of Examples 1 to 28, and optionally wherein the target code is configured to be executed by a target vector processor.
[0705] Example 30 includes a compiler configured to perform any one of the operations described in any one of Examples 1 to 29.
[0706] Example 31 includes a computing device configured to perform any one of the operations described in any one of Examples 1 to 29.
[0707] Example 32 includes a computing system comprising: at least one memory for storing instructions; and at least one processor for retrieving instructions from the memory and executing the instructions to cause the computing system to perform any one of the operations described in any one of Examples 1 to 29.
[0708] Example 33 includes a computing system comprising: a compiler for generating target code according to any one of the operations described in any one of Examples 1 to 29; and a processor for executing the target code.
[0709] Example 34 includes an apparatus comprising means for performing any one of the operations described in any one of Examples 1 to 29.
[0710] Example 35 includes an apparatus that includes: a memory interface; and processing circuitry configured to perform any one of the operations described in any of Examples 1 to 29.
[0711] Example 36 includes a method that includes any one of the operations described in any of Examples 1 to 29.
[0712] Functions, operations, components, and / or features described herein with reference to one or more aspects may be combined with or utilized in combination with one or more other functions, operations, components, and / or features described herein with reference to one or more other aspects, or vice versa.
[0713] Although certain features have been illustrated and described herein, many modifications, substitutions, changes, and equivalents will occur to those skilled in the art. Accordingly, it is to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the present disclosure.
Claims
1. A product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions that are operable, when executed by at least one processor, to cause the at least one processor to enable a compiler to: Identify loop nests based on source code, the loop nests including a plurality of loops, the plurality of loops including at least a first loop and a second loop nested within the first loop, wherein the first loop includes at least one first loop instruction outside the second loop, and wherein the second loop includes one or more second loop instructions; Transform the loop nests into transformed loops, the transformed loops including: Based on a conditional instruction of the first loop instruction, the conditional instruction being based on a status of a second loop predicate, wherein the second loop predicate is used to identify a start or an end of the second loop; and One or more transformed loop instructions based on the one or more second loop instructions; and Generate object code based on compilation of the source code, wherein the object code is based on the transformed loops.
2. The product according to claim 1, wherein the transformed loops include at least one second loop predicate instruction configured to identify the status of the second loop predicate.
3. The product according to claim 2, wherein the second loop predicate instruction includes an induction variable (IV)-independent instruction that is independent of the IV of the second loop and the IV of the first loop.
4. The product according to claim 2, wherein the object code includes one or more address generation unit (AGU) instructions to configure the AGU of a target processor to execute the object code, the AGU instructions being configured to set an AGU mask to true based on a start or an end of the second loop, wherein the second loop predicate instruction is configured to retrieve the AGU mask, and wherein the conditional instruction is based on the AGU mask.
5. The product according to claim 2, wherein the at least one second loop predicate instruction includes at least one induction variable (IV)-based instruction that is based on the IV of the second loop.
6. The product according to claim 5, wherein the at least one IV-based instruction is separate from the conditional instruction.
7. The product according to claim 5, wherein the conditional instruction includes the at least one IV-based instruction.
8. The product according to claim 1, wherein the plurality of loops includes a third loop nested within the first loop, the second loop being nested within the third loop, the first loop instruction being outside the third loop, wherein the conditional instruction is based on the status of the second loop predicate and the status of a third loop predicate, wherein the third loop predicate is used to identify a start or an end of the third loop.
9. The product according to claim 8, wherein the transformed loop includes at least one third loop predicate instruction configured to identify the state of the third loop predicate.
10. The product according to claim 8, wherein the third loop includes third loop instructions outside the second loop, wherein the transformed loop includes another conditional instruction based on the third loop instructions, and wherein the another conditional instruction is based on the state of a specific predicate for identifying the start or the end of the second loop.
11. The product according to claim 10, wherein when executed, the instruction causes the compiler to configure the second loop predicate as the specific predicate based on a determination that both the first loop instruction and the third loop instruction are pre-header instructions or both are latch instructions with respect to the second loop.
12. The product according to claim 10, wherein when executed, the instruction causes the compiler to configure the first predicate of the second loop predicate or the specific predicate as a loop start predicate for identifying the start of the second loop and to configure the second predicate of the second loop predicate or the specific predicate as a loop end predicate for identifying the end of the second loop based on a determination that a first instruction of the first loop instruction or the third loop instruction is a pre-header instruction with respect to the second loop and a second instruction of the first loop instruction or the third loop instruction is a latch instruction with respect to the second loop.
13. The product according to claim 1, wherein the plurality of loops are nested at a plurality of nesting levels, wherein the plurality of nesting levels include one or more intermediate nesting levels between a first nesting level including the first loop and a second nesting level including the second loop, and wherein the conditional instructions are respectively based on the state of the second loop predicate and the state of one or more intermediate loop predicates corresponding to the one or more intermediate nesting levels.
14. The product according to claim 13, wherein the second nesting level is the innermost nesting level among the plurality of nesting levels.
15. The product according to any one of claims 1 to 14, wherein the at least one first loop instruction includes a pre-header instruction to be performed before the first iteration of the second loop, and wherein the second loop predicate includes a loop start predicate for identifying the start of the second loop.
16. The product according to claim 15, wherein the conditional instruction is before all of the transformed loop instructions in the one or more transformed loop instructions based on the one or more second loop instructions.
17. The product according to any one of claims 1 to 14, wherein the at least one first loop instruction includes a latch instruction to be performed after the last iteration of the second loop, and wherein the second loop predicate includes a loop end predicate for identifying the end of the second loop.
18. The product according to claim 17, wherein the conditional instruction is after all of the transformed loop instructions in the one or more transformed loop instructions based on the one or more second loop instructions.
19. The product according to any one of claims 1 to 14, wherein the transformed loop includes a perfect flat loop, in which all computational operations of the loop nest are implemented in the innermost loop.
20. The product according to any one of claims 1 to 14, wherein the transformed loop includes a fully folded loop, and the fully folded loop includes only a single basic block loop based on the plurality of loops.
21. The product according to any one of claims 1 to 14, wherein the target code includes one or more address generation unit (AGU) instructions to configure the AGU of the target processor to execute the target code, and the AGU instructions are configured to set the AGU mask to true based on the start or the end of the second loop, and wherein the conditional instruction is based on the AGU mask.
22. The product according to any one of claims 1 to 14, wherein the first loop instruction includes a first loop operation on a variable, and the conditional instruction includes a selection operation to select between a first operation on the variable and a second operation on the variable, and wherein the first operation on the variable is based on the first loop operation on the variable.
23. The product according to any one of claims 1 to 14, wherein the conditional instruction includes a selection operation to select between a first value and a second value based on the state of the second loop predicate.
24. The product according to any one of claims 1 to 14, wherein the one or more transformed loop instructions include at least one second loop instruction among the one or more second loop instructions.
25. The product according to any one of claims 1 to 14, wherein the source code includes Open Computing Language (OpenCL) code.
26. The product according to any one of claims 1 to 14, wherein the computer-executable instructions, when executed, cause the compiler to compile the source code into the target code according to an LLVM-based compilation scheme.
27. The product according to any one of claims 1 to 14, wherein the target code is configured to be executed by a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor.
28. The product according to any one of claims 1 to 14, wherein the target code is configured to be executed by a target vector processor.
29. A computing system, which comprises: at least one memory for storing instructions; and at least one processor for retrieving the instructions from the memory and for executing the instructions to cause the computing system to: Identifying a loop nest based on source code, the loop nest including a plurality of loops, the plurality of loops including at least a first loop and a second loop nested in the first loop, wherein the first loop includes at least one first loop instruction outside the second loop, and wherein the second loop includes one or more second loop instructions; Transforming the loop nest into a transformed loop, the transformed loop including: A conditional instruction based on the first loop instruction, the conditional instruction being based on the state of a second loop predicate, wherein the second loop predicate is used to identify the start or the end of the second loop; and One or more transformed loop instructions based on the one or more second loop instructions; and Generating target code based on compilation of the source code, wherein the target code is based on the transformed loop.
30. The computing system according to claim 29, comprising a target processor for executing the target code.
31. A method, which comprises: Identifying a loop nest based on source code, the loop nest including a plurality of loops, the plurality of loops including at least a first loop and a second loop nested in the first loop, wherein the first loop includes at least one first loop instruction outside the second loop, and wherein the second loop includes one or more second loop instructions; Transforming the loop nest into a transformed loop, the transformed loop including: A conditional instruction based on the first loop instruction, the conditional instruction being based on the state of a second loop predicate, wherein the second loop predicate is used to identify the start or the end of the second loop; and One or more transformed loop instructions based on the one or more second loop instructions; and Generating target code based on compilation of the source code, wherein the target code is based on the transformed loop.
32. The method according to claim 31, wherein the transformed loop includes at least one second loop predicate instruction configured to identify the state of the second loop predicate.