Apparatus, system and method for compiling code for processor
By designing a compiler that integrates the front-end, mid-end and back-end to generate optimized object code, it solves the problem that existing compilers are difficult to support efficient processing of functional requirements in a vector processor environment, and realizes code generation of high-performance images and vector processing.
Patent Information
- Application Number
- CN202380071501.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-27
AI Technical Summary
Existing compilers are difficult to effectively support efficient processing of functional requirements, especially in vector processor environments.
A compiler is designed to generate optimized object code through a combination of front-end, mid-end and back-end, and supports high-performance image and vector processing of vector processors. Specific measures include lexical analysis, grammatical analysis, semantic analysis, automatic vectorization analysis, register allocation and instruction scheduling, etc.
It realizes efficient code generation in a vector processor environment, improves processing performance and code quality, and meets the needs of high-performance images and vector processing.
Smart Images

Figure CN120051760A_ABST
Abstract
Description
[0001] Cross-reference
[0002] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 415,303, filed on October 12, 2022, entitled "APPARATUS, SYSTEM, AND METHOD OF COMPILING CODE FOR A PROCESSOR", the entire disclosure of which is incorporated herein by reference. BACKGROUND OF THE DISCLOSURE
[0003] A compiler can be configured to compile source code into object code configured to be executed by a processor.
[0004] There is a need to provide technical solutions to support efficient processing functionality. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity of presentation. Additionally, reference numerals may be repeated in the figures to indicate corresponding or similar elements. The drawings are listed below.
[0006] Figure 1 is a schematic block diagram illustration of a system in accordance with some exemplary aspects.
[0007] Figure 2 is a schematic illustration of a compiler in accordance with some exemplary aspects.
[0008] Figure 3 is a schematic illustration of a vector processor in accordance with some exemplary aspects.
[0009] Figure 4 is a schematic flowchart illustration of a method of compiling code for a processor in accordance with some exemplary aspects.
[0010] Figure 5 is a schematic flowchart illustration of a method of compiling code for a processor in accordance with some exemplary aspects.
[0011] Figure 6 is a schematic illustration of a product in accordance with some exemplary aspects. DETAILED DESCRIPTION
[0012] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of some aspects. However, one of ordinary skill in the art will understand that some aspects may be practiced without these specific details. In other instances, well-known methods, procedures, components, units, and / or circuits have not been described in detail so as not to obscure the discussion.
[0013] Some portions of the detailed description that follows are presented in terms of algorithms and symbolic representations of operations on data bits or binary digital signals within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to convey the substance of their work to others skilled in the art.
[0014] An algorithm is here, and generally, conceived to be a self-consistent sequence of acts or operations leading to a desired result. These include physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc. It should be understood, however, that all of these terms and like terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities.
[0015] Discussions herein using terms such as, for example, "processing", "computing", "calculating", "determining", "establishing", "analyzing", "checking", etc. can refer to operations and / or processes of a computer, computing platform, computing system, or other electronic computing device that manipulates data represented as physical (e.g., electronic) quantities within a computer's registers and / or memory and / or transforms that data into other data similarly represented as physical quantities within a computer's registers and / or memory or other information storage media capable of storing instructions for performing the operations and / or processes.
[0016] As used herein, the terms "plurality" and "a variety" include, for example, "a plurality" or "two or more". For example, "a plurality of items" includes two or more items.
[0017] References to "an aspect", "one aspect", "exemplary aspect", "aspects", etc. indicate that the aspect so described may include a particular feature, structure, or characteristic, but not every aspect necessarily includes the particular feature, structure, or characteristic. Moreover, repeated use of the phrase "in one aspect" does not necessarily refer to the same aspect, although it may.
[0018] As used herein, unless otherwise specified, use of the ordinal adjectives "first", "second", "third", etc. to describe a common object merely indicates that different instances of the same object are being referred to and is not intended to imply that the objects so described must be in a given sequence, whether in terms of time, space, rank, or any other manner.
[0019] For example, some aspects may be captured in the form of purely hardware aspects, purely software aspects, or aspects that include both hardware and software elements. Some aspects may be implemented in software, including but not limited to firmware, resident software, microcode, etc.
[0020] In addition, some aspects may be captured in the form of a computer program product that is accessible from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For example, a computer-usable or computer-readable medium may be or may include any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0021] In some exemplary aspects, the medium may be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium.
[0022] In some exemplary aspects, a data processing system suitable for storing and / or executing program code may include, for example, at least one processor that is directly or indirectly coupled to memory elements via a system bus. The memory elements may include, for example, local memory, mass storage devices, and cache memory that are employed during the actual execution of program code, and the cache memory may provide temporary storage of at least some program code to reduce the number of times code must be retrieved from the mass storage device during execution.
[0023] In some exemplary aspects, input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) may be directly or indirectly coupled to the system via an intermediate I / O controller. In some exemplary aspects, a network adapter may be coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices, for example, via an intermediate private or public network. In some exemplary aspects, modems, cable modems, and Ethernet cards are exemplary instances of network adapter types. Other suitable components may be used.
[0024] Some aspects may be used in conjunction with a variety of devices and systems, such as, for example, computing devices, computers, mobile computers, non-mobile computers, server computers, etc.
[0025] As used herein, the term "circuitry" may refer to, be part of, or include an application specific integrated circuit (ASIC), an integrated circuit, an electronic circuit, a processor (shared, dedicated, or grouped), and / or a memory (shared, dedicated, or grouped) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable hardware components that provide the described functionality. In some aspects, some of the functionality associated with the circuitry may be implemented by one or more software or firmware modules. In some aspects, the circuitry may include logic that is at least partially operable in hardware.
[0026] The term "logic" may refer to, for example, computing logic embedded in circuitry of a computing device and / or computing logic stored in a memory of a computing device. For example, the logic may be accessed by a processor of the computing device to execute the computing logic for performing computing functions and / or operations. In one instance, the logic may be embedded in various types of memories and / or firmware, such as silicon blocks of various chips and / or processors. The logic may be included in and / or implemented as part of various circuitry, such as processor circuitry, control circuitry, and / or the like. In one instance, the logic may be embedded in volatile and / or non-volatile memories, including random access memory, read-only memory, programmable memory, magnetic memory, flash memory, persistent memory, and the like. The logic may be executed by one or more processors using a memory (e.g., registers, caches, buffers, and / or the like) coupled to the one or more processors, e.g., to execute the logic as needed.
[0027] Now referring to Figure 1 , which schematically illustrates a block diagram of a system 100 according to some exemplary aspects.
[0028] As Figure 1 shown, in some exemplary aspects, system 100 may include a computing device 102.
[0029] In some exemplary aspects, device 102 may be implemented using suitable hardware components and / or software components, such as a processor, a controller, a memory unit, a storage unit, an input unit, an output unit, a communication unit, an operating system, an application, and the like.
[0030] In some exemplary aspects, device 102 may include, for example, a computer, a mobile computing device, a non-mobile computing device, a laptop computer, a notebook computer, a tablet computer, a handheld computer, a personal computer (PC), and the like.
[0031] In some exemplary aspects, device 102 may include, for example, one or more of the following: a processor 191, an input unit 192, an output unit 193, a memory unit 194, and / or a storage unit 195. Device 102 may optionally include other suitable hardware components and / or software components. In some exemplary aspects, some or all components of one or more devices in device 102 may be enclosed in a common housing or package and may be interconnected or operably associated using one or more wired or wireless links. In other aspects, components of one or more devices in device 102 may be distributed among multiple or separate devices.
[0032] In some exemplary aspects, processor 191 may include, for example, a central processing unit (CPU), a digital signal processor (DSP), one or more processor cores, a single-core processor, a dual-core processor, a multi-core processor, a microprocessor, a host processor, a controller, multiple processors or controllers, a chip, a microchip, one or more circuits, circuitry, a logic unit, an integrated circuit (IC), an application-specific IC (ASIC), or any other suitable general-purpose or specific processor or controller. Processor 191 may execute, for example, instructions of an operating system (OS) of device 102 and / or instructions of one or more suitable applications.
[0033] In some exemplary aspects, input unit 192 may include, for example, a keyboard, a keypad, a mouse, a touch screen, a touchpad, a trackball, a stylus, a microphone, or other suitable pointing or input devices. Output unit 193 may include, for example, a monitor, a screen, a touch screen, a flat panel display, a light-emitting diode (LED) display unit, a liquid crystal display (LCD) display unit, a plasma display unit, one or more audio speakers or headphones, or other suitable output devices.
[0034] In some exemplary aspects, memory unit 194 includes, for example, a random access memory (RAM), a read-only memory (ROM), a dynamic RAM (DRAM), a synchronous DRAM (SD-RAM), a flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short-term memory unit, a long-term memory unit, or other suitable memory units. Storage unit 195 may include, for example, a hard disk drive, a solid-state drive (SSD), or other suitable removable or non-removable storage units. Memory unit 194 and / or storage unit 195 may store, for example, data processed by device 102.
[0035] In some exemplary aspects, device 102 may be configured to communicate with one or more other devices via at least one network 103 (e.g., a wireless and / or wired network).
[0036] In some exemplary aspects, network 103 may include a wired network, a local area network (LAN), a wireless network, a wireless LAN (WLAN) network, a radio network, a cellular network, a WiFi network, an IR network, a Bluetooth (BT) network, and the like.
[0037] In some exemplary aspects, device 102 may be configured to perform and / or execute one or more operations, modules, processes, procedures, and / or the like, as described herein, for example.
[0038] In some exemplary aspects, device 102 may include a compiler 160 that may be configured to generate object code 115 based on source code 112, for example, as described below.
[0039] In some exemplary aspects, compiler 160 may be configured to translate source code 112 into object code 115, for example, as described below.
[0040] In some exemplary aspects, compiler 160 may include or may be implemented as software, software modules, applications, programs, subroutines, instructions, instruction sets, computing code, words, values, symbols, and / or the like.
[0041] In some exemplary aspects, source code 112 may include computer code written in a source language.
[0042] In some exemplary aspects, the source language may include a programming language. For example, the source language may include a high-level programming language, such as, for example, the C language, the C++ language, and / or the like.
[0043] In some exemplary aspects, object code 115 may include computer code written in a target language.
[0044] In some exemplary aspects, the target language may include a low-level language, such as, for example, assembly language, object code, machine code, and the like.
[0045] In some exemplary aspects, object code 115 may include one or more object files that may create and / or form an executable program.
[0046] In some exemplary aspects, the executable program may be configured to be executed on a target computer. For example, the target computer may include specific computer hardware, a specific machine, and / or a specific operating system.
[0047] In some exemplary aspects, the executable program may be configured to be executed on processor 180, for example, as described below.
[0048] In some exemplary aspects, the processor 180 may include a vector processor 180, for example, as described below. In other aspects, the processor 180 may include any other type of processor.
[0049] Some exemplary aspects are described herein with respect to a compiler (e.g., compiler 160) that is configured to compile source code 112 into target code 115 that is configured to be executed by a vector processor 180, for example, as described below. In other aspects, the compiler (e.g., compiler 160) is configured to compile source code 112 into target code 115 that is configured to be executed by any other type of processor 180.
[0050] In some exemplary aspects, the processor 180 may be implemented as part of the device 102.
[0051] In other aspects, the processor 180 may be implemented as part of any other device that is, for example, separate from the device 102.
[0052] In some exemplary aspects, the vector processor 180 (also referred to as an "array processor") may include processors that can be configured to process an entire vector in one instruction, for example, as described below.
[0053] In other aspects, the executable program may be configured to be executed on any other additional or alternative type of processor.
[0054] In some exemplary aspects, the vector processor 180 may be designed to support high-performance image and / or vector processing. For example, the vector processor 180 may be configured to process 1 / 2 / 3 / 4D arrays of fixed-point data and / or floating-point arrays very quickly and / or efficiently, for example.
[0055] In some exemplary aspects, the vector processor 180 may be configured to process arbitrary data, such as a structure with a pointer to a structure. For example, the vector processor 180 may include a scalar processor to compute non-vector data, assuming the non-vector data is minimal, for example.
[0056] In some exemplary aspects, the compiler 160 may be implemented as a local application to be executed by the device 102. For example, the memory unit 194 and / or the storage unit 195 may store the instructions resulting from the compiler 160, and / or the processor 191 may be configured to execute the instructions generated in the compiler 160 and / or perform one or more calculations and / or processes of the compiler 160, for example, as described below.
[0057] In other aspects, the compiler 160 may include a remote application to be executed by any suitable computing system (e.g., server 170).
[0058] In some exemplary aspects, server 170 may include at least a remote server, a network-based server, a cloud server, and / or any other server.
[0059] In some exemplary aspects, server 170 may include a suitable memory and / or storage unit 174 and a suitable processor 171, the memory and / or storage unit having instructions generated in compiler 160 stored thereon, the processor for executing the instructions, as described below, for example.
[0060] In some exemplary aspects, compiler 160 may include a combination of remote applications and local applications.
[0061] In one instance, compiler 160 may be downloaded and / or received by a user of device 102 from another computing system (e.g., server 170) such that compiler 160 may be executed locally by the user of device 102. For example, the instructions may be received and stored temporarily in the memory of device 102 or any suitable short-term memory or buffer, for example, before being executed by processor 191 of device 102.
[0062] In another instance, compiler 160 may include a client module to be executed locally by device 102 and a server module to be executed by server 170. For example, the client module may include and / or may be implemented as a local application, a web application, a website, a web client, such as a HyperText Markup Language (HTML) web application, etc.
[0063] For example, one or more first operations of compiler 160 may be executed locally by device 102, and / or one or more second operations of compiler 160 may be executed remotely by server 170.
[0064] In other aspects, compiler 160 may include any other suitable computing arrangement and / or scheme, or may be implemented by any other suitable computing arrangement and / or scheme.
[0065] In some exemplary aspects, system 100 may include an interface 110 (e.g., a user interface) to interface between a user of device 102 and one or more elements of system 100 (e.g., compiler 160).
[0066] In some exemplary aspects, interface 110 may be implemented using any suitable hardware components and / or software components, such as a processor, a controller, a memory unit, a storage unit, an input unit, an output unit, a communication unit, an operating system, and / or an application.
[0067] In some aspects, interface 110 may be implemented as part of any suitable module, system, device, or component of system 100.
[0068] In other aspects, interface 110 may be implemented as a separate element of system 100.
[0069] In some exemplary aspects, interface 110 may be implemented as part of device 102. For example, interface 110 may be associated with and / or included as part of device 102.
[0070] In one instance, interface 110 may be implemented as part of, for example, middleware and / or any suitable application of device 102. For example, interface 110 may be implemented as part of compiler 160 and / or part of the OS of device 102.
[0071] In some exemplary aspects, interface 110 may be implemented as part of server 170. For example, interface 110 may be associated with and / or included as part of server 170.
[0072] In one instance, interface 110 may include or be part of: web-based applications, websites, web pages, plugins, ActiveX controls, rich content components (e.g., Flash or Shockwave components), etc.
[0073] In some exemplary aspects, interface 110 may be associated with and / or may include: for example, gateway (GW) 113 and / or application programming interface (API) 114, for example, to transfer information and / or communication between elements of system 100 and / or to one or more other parties (e.g., internal or external parties), users, applications, and / or systems.
[0074] In some aspects, interface 110 may include any suitable graphical user interface (GUI) 116 and / or any other suitable interface.
[0075] In some exemplary aspects, interface 110 may be configured to receive source code 112 from a user of, for example, device 102 via, for example, GUI 116 and / or API 114.
[0076] In some exemplary aspects, interface 110 may be configured to transfer source code 112 to, for example, compiler 160, for example, to generate object code 115, as described below.
[0077] Reference Figure 2 , which schematically illustrates compiler 200 according to some exemplary aspects. For example, compiler 160 ( Figure 1 ) may implement one or more elements of compiler 200, and / or may perform one or more operations and / or functionality of compiler 200.
[0078] In some exemplary aspects, asFigure 2 As shown, the compiler 200 can be configured to generate the target code 233, for example, by compiling the source code 212 in the source language.
[0079] In some exemplary aspects, as Figure 2 shown, the compiler 200 can include a front end 210 that is configured to receive and analyze the source code 212 in the source language.
[0080] In some exemplary aspects, the front end 210 can be configured to generate the intermediate code 213, for example, based on the source code 212.
[0081] In some exemplary aspects, the intermediate code 213 can include a lower-level representation of the source code 212.
[0082] In some exemplary aspects, the front end 210 can be configured to perform, for example, lexical analysis, syntax analysis, semantic analysis, and / or any other additional or alternative type of analysis on the source code 212.
[0083] In some exemplary aspects, the front end 210 can be configured to use the analysis results of the source code 212 to identify errors and / or problems. For example, the front end 210 can be configured to generate error messages, for example, including error and / or warning messages, for example, the error message can identify the location in the source code 212, for example, the location where the error or problem is detected.
[0084] In some exemplary aspects, as Figure 2 shown, the compiler 200 can include a middle end 220 that is configured to receive and process the intermediate code 213 and generate the adjusted (e.g., optimized) intermediate code 223.
[0085] In some exemplary aspects, the middle end 220 can be configured to perform one or more adjustments (e.g., optimizations) on the intermediate code 213, for example, to generate the adjusted intermediate code 223.
[0086] In some exemplary aspects, the middle end 220 can be configured to perform one or more optimizations on the intermediate code 213, for example, independent of the type of target computer used to execute the target code 233.
[0087] In some exemplary aspects, the middle end 220 can be implemented to support the use of the optimized intermediate code 223, for example, for different machine types.
[0088] In some exemplary aspects, the middle end 220 can be configured to optimize the intermediate representation of the intermediate code 223, for example, to improve the performance and / or quality of the resulting target code 233.
[0089] In some exemplary aspects, one or more optimizations to the intermediate code 213 may include, for example, inlining expansion, dead code elimination, constant propagation, loop transformation, parallelization, and / or the like.
[0090] In some exemplary aspects, as Figure 2 shown, the compiler 200 may include a backend 230 that is configured to receive and process the adjusted intermediate code 213 and to generate target code 233 based on the adjusted intermediate code 213.
[0091] In some exemplary aspects, the backend 230 may be configured to perform one or more operations and / or processes that may be specific to the target computer for executing the target code 233. For example, the backend 230 may be configured to process the optimized intermediate code 213 by applying analysis, transformation, and / or optimization operations to the adjusted intermediate code 213, and these operations may be configured, for example, based on the target computer for executing the target code 233.
[0092] In some exemplary aspects, one or more analysis, transformation, and / or optimization operations applied to the adjusted intermediate code 213 may include, for example, resource and storage decisions, such as, register allocation, instruction scheduling, and / or the like.
[0093] In some exemplary aspects, the target code 233 may include target - related assembly code that may be specific to the target computer for executing the target code 233 and / or the target operating system of the target computer.
[0094] In some exemplary aspects, the target code 233 may include target - related assembly code for a processor (e.g., vector processor 180( Figure 1 ))).
[0095] In some exemplary aspects, the compiler 200 may include a vector microcode processor (VMP) Open Computing Language (OpenCL) compiler, as described below, for example. In other aspects, the compiler 200 may include any other type of vector processor compiler or may be implemented as part of any other type of vector processor compiler.
[0096] In some exemplary aspects, the VMP OpenCL compiler may include a low - level virtual machine (LLVM) - based compiler that may be configured according to an LLVM - based compilation scheme, for example, to degrade OpenCL C code to VMP accelerator assembly code, for example, suitable for execution by the vector processor 180( Figure 1 )).
[0097] In some exemplary aspects, compiler 200 may include one or more techniques that may be required to compile code into a format suitable for the VMP architecture, e.g., in addition to the open-source LLVM compiler passes.
[0098] In some exemplary aspects, FE 210 may be configured to parse OpenCL C code and translate it, e.g., via an Abstract Syntax Tree (AST), into, e.g., LLVM Intermediate Representation (IR).
[0099] In some exemplary aspects, compiler 200 may include a dedicated API, e.g., to detect the correct pattern for compiler pattern matching, e.g., a pattern suitable for VMP. For example, VMP may be configured as a Complex Instruction Set Computer (CISC) machine that implements a very complex Instruction Set Architecture (ISA) that may be difficult to target from standard C code. Accordingly, compiler pattern matching may not easily detect the correct pattern, and for such cases, the compiler may require a dedicated API.
[0100] In some exemplary aspects, FE 210 may implement one or more vendor extensions built-in that may be targeted at the VMP-specific ISA, e.g., in addition to the standard OpenCL built-ins that may be optimized for VMP machines.
[0101] In some exemplary aspects, FE 210 may be configured to implement OpenCL constructs and / or work-item functions.
[0102] In some exemplary aspects, ME 220 may be configured to process LLVM IR code that may be generic and target-independent, e.g., although it may include one or more hooks for a specific target architecture.
[0103] In some exemplary aspects, ME 220 may perform one or more custom passes, e.g., to support the VMP architecture, e.g., as described below.
[0104] In some exemplary aspects, ME 220 may be configured to perform one or more operations for control flow graph (CFG) linearization analysis, e.g., as described below.
[0105] In some exemplary aspects, CFG linearization analysis may be configured to linearize code, e.g., in cases where VMP vector code does not support standard control flow, e.g., by converting if statements to select patterns.
[0106] In one instance, ME 220 may receive a given code, e.g., as follows:
[0107]
[0108] According to this example, ME 220 can be configured to apply CFG linearization analysis to a given code, for example, as follows:
[0109] tmpA = A + 5;
[0110] tmpB = B * 2;
[0111] mask = x > 0;
[0112] A = Select mask, tmpA, A
[0113] B = Select not mask, tmpB, B
[0114] Example (1)
[0115] In some exemplary aspects, ME 220 can be configured to perform one or more operations of auto-vectorization analysis, for example, as described below.
[0116] In some exemplary aspects, auto-vectorization analysis can be configured to vectorize (e.g., auto-vectorize) a given code, for example, to utilize the vector capabilities of the VMP.
[0117] In some exemplary aspects, ME 220 can be configured to perform auto-vectorization analysis, for example, to vectorize the code into a scalar form. For example, some or all operations of auto-vectorization analysis may not be performed if the code is already provided in a vectorized form.
[0118] In some exemplary aspects, for example, in some use cases and / or scenarios, the compiler may not always be able to auto-vectorize the code, for example, due to data dependencies between loop iterations.
[0119] In one example, ME 220 can receive a given code, for example, as follows:
[0120]
[0121] According to this example, ME 220 can be configured to perform CFG auto-vectorization analysis by applying a first transformation, for example, as follows:
[0122]
[0123] For example, ME 220 can be configured to perform CFG auto-vectorization analysis by applying a second transformation, for example, after the first transformation, for example, as follows:
[0124]
[0125] In some exemplary aspects, the ME 220 may be configured to perform one or more operations of Scratchpad Memory Loop Access Analysis (SPMLAA), for example, as described below.
[0126] In some exemplary aspects, SPMLAA may define a Processing Block (PB), for example, which should later be outlined and compiled for the VMP.
[0127] In some exemplary aspects, the Processing Block may include an acceleration loop, which may be executed by the vector unit of the VMP.
[0128] In some exemplary aspects, a PB (e.g., each PB) may include memory references. For example, some or all of the memory accesses may refer to local memory banks.
[0129] In some exemplary aspects, the VMP may enable access to the memory banks through an AGU (e.g., the AGU 320 as described below with reference to Figure 3 and a Scatter-Gather unit (SG).
[0130] In some exemplary aspects, the AGU may be pre-configured, for example, before the loop execution. For example, the loop trip count may be calculated, for example, before running the Processing Block.
[0131] In some exemplary aspects, image references may be created at this stage, for example, some or all of the image references, and then the stride and offset may be calculated, for example, the per-dimension stride and offset for each reference.
[0132] In some exemplary aspects, the ME 220 may be configured to perform one or more operations of AGU Planner Analysis, for example, as described below.
[0133] In some exemplary aspects, the AGU Planner Analysis may include iterator specification, which may cover image references from the entire Processing Block, for example, all image references.
[0134] In some exemplary aspects, the iterator may cover single references or groups of references.
[0135] In some exemplary aspects, one or more memory references may be merged through shuffle instructions and / or reuse the same access, and / or save values read from previous iterations.
[0136] In some exemplary aspects, other memory references, for example, without a linear access pattern, may be processed using a Scatter-Gather (SG) unit, which may have a performance penalty, for example, because it may need to maintain indices and / or masks.
[0137] In some exemplary aspects, the scheduling may be configured as an arrangement of iterators in a processing block. For example, the processing block may theoretically have multiple schedules, for example.
[0138] In some exemplary aspects, the AGU scheduler analysis may be configured to build all possible schedules for all PBs and select a combination, for example, the best combination, from all valid combinations, for example.
[0139] In some exemplary aspects, the total number of iterators in a valid combination may be limited, for example, not exceeding the number of available AGUs on the VMP.
[0140] In some exemplary aspects, one or more parameters may be defined for an iterator (e.g., for each iterator), for example, including stride, width, and / or base, for example, as part of the AGU scheduler analysis. For example, the minimum-maximum range of an iterator may be defined in dimensions, for example, for each dimension, for example, as part of the AGU scheduler analysis.
[0141] In some exemplary aspects, the AGU scheduler analysis may be configured to track and evaluate memory references to an image, for example, each memory reference, for example, to understand its access pattern.
[0142] In one instance, according to Instance 2a / 2b, image "a" as the base address can be accessed with a 32-byte stride for 64 iterations.
[0143] In some exemplary aspects, LLVM may include Scalar Evaluation Analysis (SCEV), which may compute the access pattern, for example, to understand each image reference.
[0144] In some exemplary aspects, ME 220 may utilize the masking capabilities of the AGU, for example, to avoid maintaining induction variables, which may incur a performance penalty.
[0145] In some exemplary aspects, ME 220 may be configured to perform one or more operations of rewrite analysis, for example, as described below.
[0146] In some exemplary aspects, the rewrite analysis may be configured to transform the code of the processing block, for example, when setting iterators and / or modifying memory access instructions.
[0147] In some exemplary aspects, the setting of iterators (e.g., all iterators) may be implemented in IR in a target-specific intrinsic function. For example, the setting of iterators may reside in the pre-header of the outermost loop.
[0148] In some exemplary aspects, the rewrite analysis may include loop perfection analysis, for example, as described below.
[0149] In some exemplary aspects, code may be compiled with the goal that substantially all computations should be performed within the innermost loop.
[0150] For example, loop peeling analysis may promote instructions, e.g., to move operations that would occur after the last iteration of a loop into the loop.
[0151] For example, loop peeling analysis may sink instructions, e.g., to move operations that would occur before the first iteration of a loop into the loop.
[0152] For example, loop peeling analysis may promote instructions and / or sink instructions such that substantially all instructions are moved from outer loops to the innermost loop.
[0153] For example, loop peeling analysis may be configured to provide a technical solution to support VMP iterators, e.g., to work only on perfectly nested loops.
[0154] For example, loop peeling analysis may result in a situation where there are no instructions between "for" statements that form a loop, e.g., to support a VMP iterator that cannot simulate such a situation.
[0155] In some exemplary aspects, loop peeling analysis may be configured to fold nested loops into a single folded loop.
[0156] In one instance, ME 220 may receive a given code, e.g., as follows:
[0157]
[0158] According to this instance, ME 220 may be configured to perform loop peeling analysis to fold the nested loops in the code into a single folded loop, e.g., as follows:
[0159]
[0160] In some exemplary aspects, ME 220 may be configured to perform one or more operations of vector loop demarcation analysis, e.g., as described below.
[0161] In some exemplary aspects, vector loop demarcation analysis may be configured to partition code between a scalar subsystem and a vector subsystem, e.g., between the vector processing block 310 ( Figure 3 ) and the scalar processor 330 ( Figure 3 ) as described with reference to Figure 3 .
[0162] In some exemplary aspects, the VMP accelerator may include scalar and / or vector subsystems, as described below, for example. For example, each of the subsystems may have different computing units / processors. Accordingly, scalar code may be compiled on a scalar compiler (e.g., the SSC compiler), and / or accelerated vector code may run on the VMP vector processor.
[0163] In some exemplary aspects, vector loop delineation analysis may be configured to create separate functions for the loop bodies of accelerated vector code. For example, these functions may be tagged for the VMP and / or may proceed to the VMP backend, for example, while the rest of the code may be compiled by the SSC compiler.
[0164] In some exemplary aspects, one or more parts of a vector loop (e.g., configuration of vector units and / or initialization of vector registers) may be performed by scalar units. However, these parts may be performed at a later stage, for example, by backpatching the scalar code, for example, because the scalar code may still be in LLVM IR before being processed by the SSC compiler.
[0165] In some exemplary aspects, the BE 230 may be configured to translate LLVM IR into machine instructions. For example, the BE 230 may not be target-agnostic and may be familiar with target-specific architectures and optimizations, for example, as compared to the ME 220 which may be agnostic to target-specific architectures.
[0166] In some exemplary aspects, the BE 230 may be configured to perform one or more analyses that may be specific to the target machine (e.g., the VMP machine) to which the code is being lowered, for example, although the BE 230 may use the general LLVM.
[0167] In some exemplary aspects, the BE 230 may be configured to perform one or more operations of instruction lowering analysis, as described below, for example.
[0168] In some exemplary aspects, instruction lowering analysis may be configured to translate LLVM IR into target-specific instruction machine IR (MIR), for example, by translating the LLVM IR into a directed acyclic graph (DAG).
[0169] In some exemplary aspects, the DAG may undergo an instruction legalization process, for example, based on data types and / or VMP instructions, which may be supported by the VMP HW.
[0170] In some exemplary aspects, instruction lowering analysis may be configured to perform a pattern matching process, for example, after the instruction legalization process, to lower the nodes (e.g., each node) in the DAG into, for example, VMP-specific machine instructions.
[0171] In some exemplary aspects, instruction demotion analysis may be configured to generate MIR, for example, after a pattern matching process.
[0172] In some exemplary aspects, instruction demotion analysis may be configured to demote instructions according to a machine application binary interface (ABI) and / or calling convention.
[0173] In some exemplary aspects, BE 230 may be configured to perform one or more operations of unit balance analysis, for example, as described below.
[0174] In some exemplary aspects, unit balance analysis may be configured to balance instructions among VMP computing units, for example, among the data processing units 316( Figure 3 as described below. Figure 3 )
[0175] In some exemplary aspects, unit balance analysis may be familiar with some or all available arithmetic transformations, and / or may perform transformations according to an optimal algorithm.
[0176] In some exemplary aspects, BE 230 may be configured to perform one or more operations of modulo scheduler (pipeliner) analysis, for example, as described below.
[0177] In some exemplary aspects, the pipeliner may be configured to schedule instructions according to one or more constraints (such as data dependencies, resource bottlenecks, and / or any other constraints), for example, using the swing modulo scheduling (SMS) heuristic and / or any other additional and / or alternative heuristics.
[0178] In some exemplary aspects, the pipeliner may be configured to schedule, for example, a set of very long instruction word (VLIW) instructions (such as the initiation interval (II)) that the program will iterate over during a steady state.
[0179] In some exemplary aspects, a performance metric may be measured, and the performance metric may be based on the number of cycles executable by a typical loop, for example, as follows:
[0180] (Input data size in bytes)*II / (Bytes consumed / produced per iteration)
[0181] In some exemplary aspects, the pipeliner may attempt to minimize the II as much as possible, for example, to improve performance.
[0182] In some exemplary aspects, the pipeliner may be configured to calculate the minimum II and schedule accordingly. For example, if the pipeliner scheduling fails, the pipeliner may attempt to increase the II and retry the scheduling, for example, until a predefined II threshold is violated.
[0183] In some exemplary aspects, BE 230 may be configured to perform one or more operations of register allocation analysis, e.g., as described below.
[0184] In some exemplary aspects, register allocation analysis may be configured to attempt to assign registers in an efficient (e.g., optimal) manner.
[0185] In some exemplary aspects, register allocation analysis may assign values to bypass vector registers, general-purpose vector registers, and / or scalar registers.
[0186] In some exemplary aspects, values may include private variables, constants, and / or values rotated across iterations.
[0187] In some exemplary aspects, register allocation analysis may implement an optimal heuristic that is suitable for one or more VMP register file (regfile) constraints. For example, in some use cases, register allocation analysis may not use standard LLVM register allocation.
[0188] In some exemplary aspects, in some cases, register allocation analysis may fail, which may mean that the loop cannot be compiled. Accordingly, register allocation analysis may implement a retry mechanism that may return to the modulo scheduler and may attempt to reschedule the loop, e.g., with an increased startup interval. For example, in many cases, increasing the startup interval may reduce register shortages, and / or may support compilation of vector loops.
[0189] In some exemplary aspects, BE 230 may be configured to perform one or more operations of SSC configuration analysis, e.g., as described below.
[0190] In some exemplary aspects, SSC configuration analysis may be configured to set the configuration for executing the kernel, e.g., AGU configuration.
[0191] In some exemplary aspects, SSC configuration analysis may be performed at a later stage, e.g., due to the configuration computed after legalization, register allocation analysis, and / or modulo scheduling analysis.
[0192] In some exemplary aspects, SSC configuration analysis may include a zero-overhead loop (ZOL) mechanism in vector loops. For example, the ZOL mechanism may configure the loop trip count based on the access pattern of memory references in the loop, e.g., to avoid executing instructions that check the loop exit condition on every iteration.
[0193] In some exemplary aspects, the VMP compilation flow may include one or more (e.g., a few) steps that may be invoked during the compilation flow in a test library (e.g., a wrapper script for compilation, execution, and / or program testing). For example, these steps may be executed outside of the LLVM compiler.
[0194] In some exemplary aspects, a PCB Hardware Description Language (PHDL) simulator may be implemented to perform one or more roles of an assembler, an encoder, and / or a linker.
[0195] In some exemplary aspects, the compiler 200 may be configured to provide technical solutions to support robustness, which may enable the compilation of a wide range of loop selections in the presence of HW limitations. For example, the compiler 200 may be configured to support technical solutions that may not generate verification errors.
[0196] In some exemplary aspects, the compiler 200 may be configured to provide technical solutions to support programmability, which may provide the user with the ability to express code in multiple ways that can be correctly compiled to the VMP architecture.
[0197] In some exemplary aspects, the compiler 200 may be configured to provide technical solutions to support an improved user experience, which may allow the user to be able to debug and / or profile the code. For example, the improved user experience may provide informative error messages, reporting tools, and / or profiling tools.
[0198] In some exemplary aspects, the compiler 200 may be configured to provide technical solutions to support improved performance, e.g., to optimize VMP assembly code and / or iterator access, which may result in faster execution. For example, improved performance may be achieved through high utilization of computational units and the use of their complex CISC.
[0199] Reference Figure 3 , which schematically illustrates a vector processor 300 according to some exemplary aspects. For example, the vector processor 180 ( Figure 1 ) may implement one or more elements of the vector processor 300, and / or may perform one or more operations and / or functionality of the vector processor 300.
[0200] In some exemplary aspects, the vector processor 300 may include a Vector Microcode Processor (VMP).
[0201] In some exemplary aspects, the vector processor 300 may include a wide vector machine, e.g., supporting a Very Long Instruction Word (VLIW) architecture and / or a Single Instruction / Multiple Data (SIMD) architecture.
[0202] In some exemplary aspects, the vector processor 300 may be configured to provide technical solutions to support high performance for short integer types, which may be common in, for example, computer vision and / or deep learning algorithms.
[0203] In other respects, the vector processor 300 may include any other type of vector processor, and / or may be configured to support any other additional or alternative functionality.
[0204] In some exemplary aspects, as Figure 3 shown, the vector processor 300 may include a vector processing block (vector processor) 310, a scalar processor 330, and a direct memory access (DMA) 340, for example, as described below.
[0205] In some exemplary aspects, as Figure 3 shown, the vector processing block 310 may be configured to process (e.g., efficiently process) image data and / or vector data. For example, the vector processing block 310 may be configured to use vector calculation units, for example, to accelerate calculations.
[0206] In some exemplary aspects, the scalar processor 330 may be configured to perform scalar calculations. For example, the scalar processor 330 may act as "glue logic" for a program that includes vector calculations. For example, some (e.g., even most) of the calculations of a program may be performed by the vector processing block 310. However, several tasks (e.g., some basic tasks) (e.g., scalar calculations) may be performed by the scalar processor 330.
[0207] In some exemplary aspects, the DMA 340 may be configured to interface with one or more memory elements in a chip that includes the vector processor 300.
[0208] In some exemplary aspects, the DMA 340 may be configured to read inputs from the main memory and / or write outputs to the main memory.
[0209] In some exemplary aspects, the scalar processor 330 and the vector processing block 310 may use respective local memories to process data.
[0210] In some exemplary aspects, as Figure 3 shown, the vector processor 300 may include an extractor and decoder 350, which may be configured to control the scalar processor 330 and / or the vector processing block 310.
[0211] In some exemplary aspects, the operations of the scalar processor 330 and / or the vector processing block 310 may be triggered by instructions stored in the program memory 352.
[0212] In some exemplary aspects, the DMA 340 may be configured to transfer data, for example, in parallel with the execution of program instructions in the memory 352.
[0213] In some exemplary aspects, the DMA 340 can be controlled by software, e.g., via configuration registers rather than instructions, and can accordingly be considered a second execution “thread” in the vector processor 300.
[0214] In some exemplary aspects, the scalar processor 330, the vector processing block 310, and / or the DMA 340 can include one or more data processing units, e.g., a set of data processing units, e.g., as described below.
[0215] In some exemplary aspects, the data processing unit can include hardware configured to perform computations, e.g., an arithmetic logic unit (ALU).
[0216] In one instance, the data processing unit can be configured to add numbers and / or store numbers in memory.
[0217] In some exemplary aspects, the data processing unit can be controlled by commands, e.g., encoded in the program memory 352 and / or configuration registers. For example, the configuration registers can be memory-mapped and can be written to by memory store commands of the scalar processor 330.
[0218] In some exemplary aspects, the scalar processor 330, the vector processing block 310, and / or the DMA 340 can include a status configuration that includes a set of registers and memory, e.g., as described below.
[0219] In some exemplary aspects, as Figure 3 shown, the vector processor block 310 can include a set of vector memories 312 that can be configured to store, e.g., data to be processed by the vector processor block 310.
[0220] In some exemplary aspects, as Figure 3 shown, the vector processor block 310 can include a set of vector registers 314 that can be configured to be used, e.g., in data processing performed by the vector processor block 310.
[0221] In some exemplary aspects, the scalar processor 330, the vector processing block 310, and / or the DMA 340 can be associated with a set of memory mappings.
[0222] In some exemplary aspects, the memory mapping can include a set of addresses accessible by the data processing unit that can load data from / to and / or store data in registers and memory.
[0223] In some exemplary aspects, as Figure 3As shown, the vector processing block 310 may include a plurality of address generation units (AGUs) 320, which may include addresses they can access, for example, in one or more memories in the memory 312.
[0224] In some exemplary aspects, as Figure 3 shown, the vector processor block 310 may include a plurality of data processing units 316, for example, as described below.
[0225] In some exemplary aspects, the data processing unit 316 may be configured to process commands, for example, including several digits at a time. In one instance, the command may include 8 digits. In another instance, the command may include 4 digits, 16 digits, or any other counted number of digits.
[0226] In some exemplary aspects, two or more data processing units 316 may be used simultaneously. In one instance, the data processing unit 316 may process and execute multiple different commands in an entire single cycle, for example, 3 different commands, for example, including 8 digits.
[0227] In some exemplary aspects, the data processing unit 316 may be asymmetric. For example, the first and second data processing units 316 may support different commands. For example, addition may be performed by the first data processing unit 316, and / or multiplication may be performed by the second data processing unit 316. For example, both operations may be performed by one or more additional other data processing units 316.
[0228] In some exemplary aspects, the data processing unit 316 may be configured to support arithmetic operations for many combinations of input and output data types.
[0229] In some exemplary aspects, the data processing unit 316 may be configured to support one or more operations that may be less common. For example, the processing unit 316 may support operations working with the look-up table (LUT) of the vector processor 300 and / or any other operations.
[0230] In some exemplary aspects, the data processing unit 316 may be configured to support efficient computations of non-linear functions, histograms, and / or random data access, for example, which may help to implement algorithms like image scaling, Hough transform, and / or any other algorithms.
[0231] In some exemplary aspects, the vector memory 312 may include, for example, memory banks having a size of 16K or any other size, which may be accessed in the same cycle.
[0232] In one example, the maximum memory access size can be 64 bits. According to this example, the peak throughput can be 256 bits, e.g., 64 x 4 = 256. For example, a high memory bandwidth can be achieved to utilize the computing power of the data processing unit 316.
[0233] In one example, two data processing units 316 can support 16 eight-bit multiply-accumulate operations (MACs) per cycle. According to this example, the two data processing units 316 may not be useful, e.g., in cases where the input numbers are not fetched at that rate, and / or where there is no exactly 256-bit input, e.g., 16 x 8 x 2 = 256.
[0234] In some exemplary aspects, the AGU 320 can be configured to perform memory access operations, e.g., load and store data from / to the vector memory 314.
[0235] In some exemplary aspects, the AGU 320 can be configured to calculate the addresses of input and output data items, e.g., to handle I / O in cases where, for example, the high bandwidth is insufficient to utilize the data processing unit 316.
[0236] In some exemplary aspects, the AGU 320 can be configured to calculate the addresses of input and / or output data items, e.g., before typing a vector command block (e.g., a loop), based on configuration registers written by the scalar processor 330.
[0237] For example, the AGU 320 can be configured to write an image base address pointer, width, height, and / or stride to configuration registers, e.g., to iterate over an image.
[0238] In some exemplary aspects, the AGU 320 can be configured to handle addressing (e.g., all addressing), e.g., to provide a technical solution where the data processing unit 316 may not have the burden of incrementing a pointer or counter in a loop and / or the burden of checking for end-of-line conditions, e.g., to zero out a counter in a loop.
[0239] In some exemplary aspects, as Figure 3 shown, the AGU 320 can include 4 AGUs and, correspondingly, can access four memories 312 in the same cycle. In other aspects, any other count of AGUs 32 can be implemented.
[0240] In some exemplary aspects, the AGU 320 may not be "tied" to a memory bank 312. For example, the AGU 320 (e.g., each AGU 320) can access a memory bank 312 (e.g., each memory bank 312), provided that two or more AGUs 320 do not attempt to access the same memory bank 312 in the same cycle.
[0241] In some exemplary aspects, the vector register 314 may be configured to support communication between the data processing unit 316 and the AGU 320.
[0242] In one instance, the total number of vector registers 314 may be 28, and they may be divided into several subsets, for example, based on their functions. For example, a first subset of the vector registers 314 may be used for input / output of all data processing units 316 and / or the AGU 320; and / or a second subset of the vector registers 314 may not be used for the output of some operations (e.g., most operations) and may be used for one or more other operations, for example, to store loop-invariant inputs.
[0243] In some exemplary aspects, the data processing unit 316 (e.g., each data processing unit 316) may have one or more registers for hosting the output of the last executed operation, and for example, this output may be fed as input to other data processing units 316. For example, these registers may "bypass" the vector register 314 and may work faster than writing these outputs to the first set of vector registers 314.
[0244] In some exemplary aspects, the fetcher and decoder 350 may be configured to support low-overhead vector loops, for example, very low-overhead vector loops (also referred to as "zero-overhead vector loops"), for example, where it may not be necessary to check the termination (exit) condition of the vector loop during the execution of the vector loop.
[0245] For example, when the AGU 320 finishes iterating over the configured memory region, the AGU 320 may signal the termination (exit) condition.
[0246] For example, when the AGU 320 signals the termination condition, the fetcher and decoder 350 may exit the loop.
[0247] For example, the scalar processor 330 may be utilized to configure loop parameters, such as the first and last instructions and / or the exit condition.
[0248] In one instance, vector loops may be utilized, for example, with high memory bandwidth and / or inexpensive addressing to, for example, solve control and data flow problems, for example, providing a technical solution to allow the data processing unit 316 to process data with substantially no additional overhead.
[0249] In some exemplary aspects, scalar processor 330 may be configured to provide one or more functions that may be complementary to the functions of vector processing block 310. For example, most (e.g., the majority) of the work in a vector program may be performed by data processing unit 316. For example, scalar processor 330 may be utilized to, for example, "glue" together the various vector code blocks of a vector program.
[0250] In some exemplary aspects, scalar processor 330 may be implemented separately from vector processing block 310. In other aspects, scalar processor 330 may be configured to share one or more components and / or functions with vector processing block 310.
[0251] In some exemplary aspects, scalar processor 330 may be configured to perform operations that may not be suitable for execution on vector processing block 310.
[0252] For example, scalar processor 330 may be utilized to execute 32-bit C programs. For example, scalar processor 330 may be configured to support 1, 2, and / or 4-byte data types of C code and / or some or all of the arithmetic operators of C code.
[0253] For example, scalar processor 330 may be configured to provide a technical solution to perform operations that cannot be executed on vector processing block 310, for example, without using the fully open CPU.
[0254] In some exemplary aspects, scalar processor 330 may include, for example, a scalar data memory 332 having a size of 16K or any other size, and the scalar data memory may be configured to store data, for example, variables used by the scalar portion of a program.
[0255] For example, scalar processor 330 may store local and / or global variables declared by portable C code, and these variables may be allocated to the scalar data memory by a compiler (e.g., compiler 200( Figure 2 ))
[0256] In some exemplary aspects, as Figure 3 shown, scalar processor 330 may include a set of vector registers 334 or may be associated with them, and the set of vector registers may be used for data processing performed by scalar processor 330.
[0257] In some exemplary aspects, scalar processor 330 may be associated with a scalar memory map, and the scalar memory map may support scalar processor 330 to access substantially all the states of vector processor 300. For example, scalar processor 330 may configure the vector unit and / or DMA channels via the scalar memory map.
[0258] In some exemplary aspects, the scalar processor 330 may not be permitted access to one or more block control registers that may be used by an external processor to run and debug vector programs.
[0259] In some exemplary aspects, the DMA 340 may be configured to communicate, for example, via the main memory with one or more other components of the chip implementing the vector processor 300. For example, the DMA 340 may be configured to transfer data blocks, e.g., large, contiguous data blocks, e.g., to support the scalar processor 330 and / or the vector processing block, which may manipulate data stored in local memory. For example, a vector program may be able to use the DMA 340 to read data from the main chip memory.
[0260] In some exemplary aspects, the DMA 340 may be configured to communicate, for example, via a plurality of DMA channels (e.g., 8 DMA channels or any other counted number of DMA channels) with other elements of the chip. For example, a DMA channel (e.g., each DMA channel) may be able to transfer a rectangular patch from local memory to the main chip memory, or vice versa. In other aspects, the DMA channels may transfer any other type of data block between local memory and the main chip memory.
[0261] In some exemplary aspects, a rectangular patch may be defined by a base address pointer, a width, a height, and a stride.
[0262] For example, at peak throughput, 8 bytes may be transferred per cycle; however, there may be overhead for each patch and / or for each row in the patch.
[0263] In some exemplary aspects, the DMA 340 may be configured to transfer data in parallel with computations, for example, via a plurality of DMA channels, e.g., as long as the commands being executed do not access the local memory involved in the transfer.
[0264] In one instance, since all channels may access the same memory bus, using several channels to effect a transfer may not save I / O cycles, for example, compared to the case of using a single channel. However, multiple DMA channels may be utilized to schedule several transfers and execute them in parallel with computations. For example, this may be advantageous compared to a single channel, which may not allow scheduling of a second transfer until the first transfer is complete.
[0265] In some exemplary aspects, the DMA 340 may be associated with a memory map that may support DMA channel access to vector memories and / or scalar data. For example, access to vector memories may occur in parallel with computations. For example, access to scalar data may not typically be allowed in parallel, e.g., because the scalar processor 330 may be involved in almost any reasonable program and may access its local variables during a transfer, which may lead to memory contention with active DMA channels.
[0266] In some exemplary aspects, the DMA 340 may be configured to provide a technical solution to support the parallelization of I / O and computations. For example, a program performing computations may not have to wait for I / O, e.g., in cases where these computations can be run quickly by the vector processing block 310.
[0267] In some exemplary aspects, an external processor (e.g., a CPU) may be configured to initiate the execution of a program on the vector processor 300. For example, the vector processor 300 may remain idle, e.g., as long as program execution has not been initiated.
[0268] In some exemplary aspects, an external processor may be configured to debug a program, e.g., execute a single step at a time, stop when the program reaches a breakpoint, and / or inspect the contents of registers and memories storing program variables.
[0269] In some exemplary aspects, an external memory map may be implemented to support an external processor in controlling the vector processor 300 and / or debugging a program, e.g., by writing to the control registers of the vector processor 300.
[0270] In some exemplary aspects, the external memory map may be implemented as a superset of the scalar memory map. For example, this implementation may make all registers and memories defined by the architecture of the vector processor 300 accessible to a debugger backend running on an external processor.
[0271] In some exemplary aspects, the vector processor 300 may issue an interrupt signal, e.g., when the vector processor 300 terminates a program.
[0272] In some exemplary aspects, the interrupt signal may be used, e.g., to implement a driver to maintain a queue of programs scheduled for execution by the vector processor 300, and / or may be used, e.g., to initiate a new program by an external processor when a previously executed program has completed.
[0273] Return reference Figure 1, in some exemplary aspects, the compiler 160 may be configured to generate target code 115 that is configured to utilize the registers of a processor (e.g., a vector processor, e.g., vector processor 180) according to a register allocation scheme, e.g., as described below.
[0274] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to utilize a reduced number of allocated registers for executing a program by a processor (e.g., a vector processor), e.g., as described below.
[0275] In one instance, the register allocation scheme may be configured to provide a technical solution to utilize a reduced number of allocated registers that can be allocated from multiple vector registers 314 ( Figure 3 ) for executing a program by vector processor 300 ( Figure 3 ), e.g., as described below.
[0276] In some exemplary aspects, a compiler (e.g., compiler 160) may be configured to generate target code (e.g., target code 115) that is configured to utilize the registers of a vector processor (e.g., vector processor 189) according to a register allocation scheme, e.g., as described below.
[0277] In other aspects, a compiler (e.g., compiler 160) may be configured to generate target code (e.g., target code 115) that is configured to utilize the registers of any other suitable type of processor (e.g., any other suitable type of processor) according to a register allocation scheme, e.g., as described below.
[0278] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to improve (e.g., optimize) the allocation of registers used to execute an executable program, e.g., as described below.
[0279] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to support an improved allocation (e.g., an efficient allocation, e.g., an optimized allocation) of registers used to execute an executable program, e.g., as described below.
[0280] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to utilize a reduced number (e.g., an optimized number, e.g., a minimum number) of allocated registers to execute an executable program, e.g., as described below.
[0281] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to improve the performance of an executable program, e.g., by reducing the number of allocated registers used to execute the executable program, e.g., as described below.
[0282] In some exemplary aspects, it may be desirable to provide a technical solution to efficiently allocate vector registers of a vector processor for the execution of a program, e.g., in order to reduce the number of allocated vector registers, as described below, for example.
[0283] For example, the number of physical registers implemented by a chip including a processor (e.g., a vector processor or any other processor) may be limited, e.g., according to the design and / or layout of the chip. Accordingly, the number of vector registers implemented by the vector processor may be limited by the number of physical registers on the chip implementing the vector processor.
[0284] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce the number of allocated registers, e.g., for a CPU (e.g., a vector processor) with limited storage capacity (e.g., a limited register pool) and / or for a processor (e.g., a vector processor) with limited or no support for memory overflow / fill operations (e.g., for storing live values).
[0285] For example, a processor without fill / overflow capabilities and / or with limited storage capacity may instead be forced to use computational resources.
[0286] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce the number of allocated registers, e.g., to avoid or even eliminate the use of such additional computational resources.
[0287] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce a register shortage for the execution of a program by a processor (e.g., a vector processor or any other processor).
[0288] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce a register shortage, e.g., by reducing the number of allocated registers for the execution of a program, as described below, for example.
[0289] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to support the efficient execution of a program (e.g., a complex program that may be sensitive to register shortages). For some programs, for example, the register shortage problem may be a bottleneck, and accordingly, the register shortage may affect performance.
[0290] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce the number of allocated registers for the execution of a program, e.g., while providing a proper allocation of registers (e.g., vector registers) for the execution of the program, as described below, for example.
[0291] In some exemplary aspects, the compiler 160 may be configured to: process a given instruction schedule, for example, based on the source code 112, and generate the target code 115, which may be configured to utilize a reduced number of allocated registers, for example, for successful register allocation, as described below, for example.
[0292] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115, which is configured to utilize a data processing unit (e.g., ALU) of a processor (e.g., vector processor) to store variables (e.g., live values) of an executable program, as described below, for example.
[0293] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115, which is configured to utilize a data processing unit of a vector processor to store one or more variables (e.g., live values) of an executable program, for example, instead of storing these variables in one or more registers, as described below, for example.
[0294] In one instance, the compiler 160 may be configured to generate the target code 115, which is configured to utilize one or more data processing units in the data processing unit 316 ( Figure 3 ) to store (e.g., temporarily store) one or more variables (e.g., live values) of an executable program, for example, instead of storing one or more of these variables in one or more vector registers 314 ( Figure 3 ), as described below, for example.
[0295] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115, which is configured to utilize the internal state of the ALU, for example, to store live values in the ALU, for example, instead of storing these live values in physical registers, as described below, for example.
[0296] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115, which is configured to utilize the latency of instructions by the ALU, for example, for storing live values in the ALU, as described below, for example.
[0297] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115, which is configured to store the live values of variables in the ALU, for example, by live range splitting, as described below, for example.
[0298] In some exemplary aspects, active range splitting may include splitting the active range (interval) of the active values of a variable into, for example, a first active interval (e.g., where the active values are stored by the ALU) and a second active interval (e.g., where the active values are stored in a physical vector register), as described below, for example.
[0299] In some exemplary aspects, the active range (interval) of the active values of a variable may be split more than once, for example, to provide more than two active intervals. For example, the count of active intervals may be increased, for example, to provide more cycles during which the register is available, as described below, for example.
[0300] In some exemplary aspects, the active range of a variable may include a range of cycles of an executable program, for example, between a first cycle (e.g., which includes the first use and / or generation of the variable) and a second cycle (e.g., which includes a second use after, for example, the first use of the variable), as described below, for example.
[0301] In some exemplary aspects, the compiler 160 may be configured to identify one or more cycles (“unused variable cycles”) in the active range of a variable during which the variable is active and unused, as described below, for example.
[0302] In some exemplary aspects, the compiler 160 may be configured to allocate one or more no-op instructions in the target code 115 that will be applied to the variable, for example, to store the active values of the variable by the ALU, as described below, for example.
[0303] In some exemplary aspects, the compiler 160 may be configured to allocate one or more no-op instructions in the target code that will be executed by the ALU, such that the active values of the variable can be stored by the ALU that executes the one or more no-op instructions, as described below, for example.
[0304] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115 that includes one or more no-op instructions that may be configured to execute instructions using the latency of the ALU, such that the active values of the variable can be temporarily stored by the ALU, as described below, for example.
[0305] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115 that includes one or more no-op instructions that may be configured to provide a technical solution to temporarily store the active values of the variable by the ALU that executes the one or more no-op instructions, for example, instead of storing the active values of the variable in a vector register, as described below, for example.
[0306] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 that includes one or more no-op instructions to be applied to a variable, e.g., that may not even substantially affect the throughput of the executable program, as described below, for example.
[0307] In some exemplary aspects, a no-op instruction may include an instruction that may be configured to cause the ALU to perform a sequence of operations that may be executed by the ALU and that may cause the active value of the variable to be maintained during one or more cycles, as described below, for example.
[0308] In some exemplary aspects, a no-op instruction may include an instruction that may be configured to cause the ALU to perform a sequence of operations, e.g., the sequence of operations includes a load operation to load the active value of the variable from a vector register, an operation to be applied to the active value of the variable (e.g., without changing the active value of the variable), and a store operation to store the active value of the variable back into the same register, as described below, for example.
[0309] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 that includes one or more no-op instructions that, when executed by the ALU, may cause the ALU to execute the sequence of operations of the no-op instruction over a plurality of cycles (latency cycles), which may be based on, e.g., the latency of the ALU used to execute the instruction, as described below, for example.
[0310] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 that includes one or more no-op instructions that, when executed by the ALU, may cause the ALU to store the active value of the variable for the duration of the latency cycle.
[0311] In one instance, one or more no-op instructions may include an add-zero instruction.
[0312] In another instance, one or more no-op instructions may include a shift-zero instruction, e.g., a left shift-zero instruction or a right shift-zero instruction.
[0313] In another instance, one or more no-op instructions may include a multiply-one instruction.
[0314] In another instance, one or more no-op instructions may include any other additional and / or alternative instructions that may be configured to cause the ALU to store the active value of the variable of the ALU when executed by the ALU, e.g., for the duration of one or more latency cycles.
[0315] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into object code 115, which may be configured for execution by target processor 180, for example, as described below.
[0316] In some exemplary aspects, compiler 160 may be configured to compile source code 112 that includes OpenCL code, for example, as described below. In other aspects, compiler 160 may be configured to compile any other type of source code 112.
[0317] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into object code 115, for example, according to an LLVM-based compilation scheme. In other aspects, any other additional or alternative compilation scheme may be utilized.
[0318] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into object code 115, which may be configured for execution by, for example, a VLIW SIMD target processor.
[0319] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into object code 115, which may be configured for execution by, for example, a target vector processor.
[0320] In other aspects, compiler 160 may be configured to compile source code 112 into object code 115, which may be configured for execution by any other additional or alternative type of processor.
[0321] In some exemplary aspects, object code 115 may be configured for execution by target processor 180 over a plurality of execution cycles, for example, the plurality of execution cycles including a first execution cycle, a second execution cycle after the first execution cycle, and a third execution cycle after the second execution cycle, for example, as described below.
[0322] In some exemplary aspects, compiler 160 may be configured to generate object code 115 that includes, for example, one or more no-op instructions, which may be configured to keep a first variable live, for example, such that the value of the first variable (e.g., live value) will be available, for example, in a register of the target processor during the first execution cycle and during the third execution cycle, for example, as described below.
[0323] In some exemplary aspects, compiler 160 may configure object code 115 to include, for example, a first instruction that will be applied to the value of a first variable (e.g., live value) in a register during the first execution cycle, for example, as described below.
[0324] In some exemplary aspects, the compiler 160 may configure the target code 115 to include, for example, a second instruction that will be applied to the value of a second variable in a register during a second execution cycle, as described below, for example.
[0325] In some exemplary aspects, the compiler 160 may configure the target code 115 to include, for example, a third instruction that will be applied to the value of a first variable (e.g., an active value) in a register during a third execution cycle, as described below, for example.
[0326] In some exemplary aspects, the compiler 160 may be configured to output the target code 115, for example, in a form executable by the target processor 180.
[0327] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115, which is configured for execution by, for example, a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor (e.g., processor 180).
[0328] In other aspects, the compiler 160 may be configured to generate the target code 115, which is configured for execution by, for example, any other suitable type of processor.
[0329] In some exemplary aspects, the compiler 160 may be configured to generate the target code 115, for example, based on the source code 112 that includes Open Computing Language (OpenCL) code.
[0330] In other aspects, the compiler 160 may be configured to generate the target code 115, for example, based on the source code 112 that includes any other suitable type of code.
[0331] In some exemplary aspects, the compiler 160 may be configured to compile the source code 112 into the target code 115, for example, according to a low-level virtual machine (LLVM)-based compilation scheme.
[0332] In other aspects, the compiler 160 may be configured to compile the source code 112 into the target code 115 according to any other suitable compilation scheme.
[0333] In some exemplary aspects, the compiler 160 may be configured to configure one or more no-op instructions, for example, to cause the target processor 180 to apply one or more operations to the value of the first variable (e.g., an active value), and the one or more operations may be configured to, for example, maintain the value of the first variable (e.g., an active value) unchanged between a first execution cycle and a third execution cycle, as described below, for example.
[0334] In some exemplary aspects, the compiler 160 may be configured to configure one or more no-op instructions, e.g., to cause a data processing unit of the target processor 180 (e.g., the data processing unit 316 of the vector processor 300 ( Figure 3 )) to temporarily maintain the value of a first variable (e.g., an active value) within the data processing unit, e.g., between a first execution cycle and a third execution cycle, as described below. Figure 3 ))
[0335] In some exemplary aspects, the compiler 160 may be configured to configure one or more no-op instructions to include, e.g., an add-zero instruction, as described below.
[0336] In some exemplary aspects, the compiler 160 may be configured to configure one or more no-op instructions to include, e.g., a shift-zero instruction, as described below.
[0337] In some exemplary aspects, the compiler 160 may be configured to configure one or more no-op instructions to include, e.g., a multiply-one instruction, as described below.
[0338] In other aspects, the compiler 160 may be configured to configure one or more no-op instructions to include any other suitable additional or alternative type of no-op instruction.
[0339] In some exemplary aspects, the compiler 160 may be configured to configure one or more no-op instructions, e.g., to cause a data processing unit of the target processor 180 (e.g., the data processing unit 316 of the vector processor 300 ( Figure 3 )) to initiate loading the value of a first variable (e.g., an active value) from a register in a first execution cycle, as described below. Figure 3 ))
[0340] In some exemplary aspects, the compiler 160 may be configured to configure one or more no-op instructions, e.g., to cause a data processing unit of the target processor 180 (e.g., the data processing unit 316 of the vector processor 300 ( Figure 3 )) to initiate storing the value of a first variable (e.g., an active value) in a register in a later execution cycle, e.g., before a third execution cycle, as described below. Figure 3 ))
[0341] In some exemplary aspects, the compiler 160 may be configured to generate one or more no-op instructions to include a plurality of no-op instructions, as described below.
[0342] In some exemplary aspects, the compiler 160 may be configured to generate a first-in-order no-op instruction among the plurality of no-op instructions, the first-in-order no-op instruction may be configured, for example, to cause a data processing unit (e.g., the vector processor 300 ( Figure 3 ) of the data processing unit 316 ( Figure 3 )): For example, in a first execution cycle, the value of a first variable (e.g., an active value) is loaded from a register; and a first-in-order operation is applied to the value of the first variable (e.g., the active value), for example, the first-in-order operation is configured to maintain the value of the first variable (e.g., the active value) unchanged, for example, as described below.
[0343] In some exemplary aspects, compiler 160 may be configured to generate a last no-op instruction in order among the plurality of no-op instructions, the last no-op instruction in order may be configured to, for example, cause a data processing unit (e.g., vector processor 300) of target processor 180 to execute the operation of the processor. Figure 3 ) data processing unit 316 ( Figure 3 )): Apply the last operation in order to the output of the previous no-op instruction, which last operation in order can be configured to, for example, maintain the value of a first variable (e.g., an active value) unchanged; and store the value of the first variable (e.g., an active value) in a register, for example, as described below.
[0344] In some exemplary aspects, compiler 160 may be configured to generate target code 115 that may be configured, for example, such that execution of a last-in-order no-op instruction of one or more no-op instructions will be performed by a data processing unit (e.g., vector processor 300) of target processor 180. Figure 3 ) of the data processing unit 316 ( Figure 3 ))For example, the last no-op execution cycle in the order before, for example, the third execution cycle is started, for example, as described below.
[0345] In some exemplary aspects, compiler 160 may be configured to generate target code 115 that may be configured, for example, such that the distance between the last no-op execution cycle and the third execution cycle in order may be based on, for example, a data processing unit of target processor 180 (e.g., vector processor 300 ( Figure 3 ) of the data processing unit 316 ( Figure 3 ))'s delay, for example, as described below.
[0346] In some exemplary aspects, compiler 160 may be configured, for example, based on a data processing unit (e.g., vector processor 300 ( Figure 3 ) of the data processing unit 316 ( Figure 3configure one or more no-op instructions based on the latency of ( )), for example, as described below.
[0347] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 that may be configured such that, for example, a first instruction will be executed by a first data processing unit of a target processor 180 (e.g., a vector processor 300 ( Figure 3 )'s first data processing unit 316 ( Figure 3 ), for example, as described below.
[0348] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 that may be configured such that, for example, one or more no-op instructions will be executed by a second data processing unit of a target processor 180 (e.g., a vector processor 300 ( Figure 3 )'s first data processing unit 316 ( Figure 3 ), for example, as described below.
[0349] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 that may be configured such that the execution of the first no-op instruction in order among one or more no-op instructions will be initiated by the second data processing unit in a first execution cycle, for example, as described below.
[0350] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 that may be configured such that, for example, a second instruction will be executed by a second data processing unit, for example, as described below.
[0351] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 that may be configured such that, for example, a third instruction will be executed by a first data processing unit, for example, as described below.
[0352] In some exemplary aspects, the first data processing unit of the target processor 180 may include a first ALU, and / or the second data processing unit of the target processor 180 may include a second ALU, for example, as described below. In other aspects, any other additional or alternative type of data processing unit may be implemented.
[0353] In some exemplary aspects, the compiler 160 may be configured to identify a plurality of active ranges corresponding to a respective plurality of variables based on the source code 112, for example, as described below.
[0354] In some exemplary aspects, the compiler 160 may be configured to identify, for example, a recognized variable having an active range that includes one or more unused execution cycles during which the recognized variable is not used, as described below, for example.
[0355] In some exemplary aspects, the compiler 160 may be configured to identify the recognized variable as a first variable, for example, based on a determination that a count of consecutive unused execution cycles in the active range of the recognized variable is greater than a predefined threshold, as described below, for example.
[0356] In some exemplary aspects, the compiler 160 may be configured to configure one or more no-op instructions, for example, to keep the first variable active during one or more unused execution cycles, as described below, for example.
[0357] In some exemplary aspects, the compiler 160 may be configured to allocate a register to store the value of a second variable, for example, during an unused execution cycle among one or more unused execution cycles, as described below, for example.
[0358] In some exemplary aspects, the compiler 160 may be configured to generate target code 115 based on a VLIW instruction that can be utilized to invoke multiple different instructions during the same cycle. For example, the multiple instructions may include a no-op instruction that can be executed in parallel with other instructions, such that performance may not be affected by the addition of the no-op instruction.
[0359] In some exemplary aspects, the compiler 160 may be configured to receive source code 112 that includes an instruction schedule for a program to be executed, as described below, for example.
[0360] In some exemplary aspects, the compiler 160 may be configured to identify the active range of one or more variables in a program to be executed, as described below, for example.
[0361] In some exemplary aspects, the compiler 160 may be configured to identify one or more cycles (“unused cycles”) in the active range of a variable.
[0362] For example, an unused cycle may include a cycle during which the variable is active but not used, as described below, for example.
[0363] In some exemplary aspects, the compiler 160 may be configured to determine a recognized data processing unit (e.g., an ALU) that has a latency when performing a no-op instruction and that is available during one or more unused cycles, as described below, for example.
[0364] In some exemplary aspects, the compiler 160 may be configured to insert one or more no-op instructions to be applied to variables by the identified data processing unit during one or more unused cycles, e.g., as described below.
[0365] In some exemplary aspects, having the identified data processing unit perform no-op operations on active variables may cause the identified data processing unit to temporarily store the variables during one or more unused cycles, e.g., according to the latency of the identified data processing unit, e.g., as described below.
[0366] In some exemplary aspects, using the identified data processing unit to store variables during one or more unused cycles may provide a technical solution to free registers (e.g., vector registers) during one or more unused cycles, e.g., as described below.
[0367] In some exemplary aspects, e.g., idle registers (e.g., idle vector registers) may be utilized, e.g., to store one or more other active values of other variables during unused cycles. As a result, the number of registers (e.g., vector registers) required for program execution may be reduced.
[0368] In some exemplary aspects, the compiler 160 may be configured to generate no-op instructions that may be configured to utilize instructions with different latencies, e.g., to provide a technical solution to free one or more register cycles, e.g., when needed.
[0369] In some exemplary aspects, the compiler 160 may be configured to generate target code 115, e.g., that includes one or more no-op instructions, e.g., to allocate register vectors in an efficient manner, e.g., as described below.
[0370] In some exemplary aspects, the compiler 160 may receive the source code 112 of the program to be executed.
[0371] In some exemplary aspects, the source code 112 may include an instruction schedule that may be configured to compute an expression based on a first variable represented as a and a second variable represented as b, e.g., as follows:
[0372] (a·2 + b·3)·a (1)
[0373] In some exemplary aspects, the expression result of expression 1 may be stored at a result memory address represented as "res_add".
[0374] In some exemplary aspects, variable a may be stored at a memory address represented as a_add, and variable b may be stored at a memory address represented as b_add.
[0375] In some exemplary aspects, the compiler 160 may be configured to compile source code 112 for a processor (e.g., a vector processor) that includes a first ALU and a second ALU that may be configured to perform addition and multiplication instructions with a latency of, for example, 2 cycles. For example, a processor (e.g., a vector processor) may have a memory unit with a latency of 1 cycle for memory access operations (e.g., a store operation to store a value into the memory unit or a load operation to load a value from the memory unit).
[0376] In some exemplary aspects, the execution of Expression 1 may be implemented according to the following instruction schedule:
[0377]
[0378]
[0379] Table (1)
[0380] As shown in Table 1, one or more first operations may be performed by the first ALU, and one or more second operations may be performed by the second ALU.
[0381] As shown in Table 1, the second ALU may not perform any operations during cycles 1 and 3.
[0382] As shown in Table 1, the execution of Expression 1 may be performed over 9 cycles.
[0383] In one example, three registers represented as R0, R1, and R2 may be utilized to execute the instructions of Table 1, for example, as follows:
[0384] Loop R0 R1 R2 0 Don't care Don't care Don't care 1 a Don't care Don't care 2 a Don't care b 3 a a_mul_2 Don't care 4 a a_mul_2 b_mul_3 5 a Don't care Don't care 6 a Sum Don't care 7 Don't care Don't care Don't care 8 Result Don't care Don't care
[0385] Table (2)
[0386] As shown in Table 2, it may be necessary for register R0 to store variable a during cycles 1 - 6 and to store the expression result during cycle 8.
[0387] As shown in Table 2, it may be necessary for register R2 to store the multiplication result of the product of variable b and 3, represented as b_mul_3, during cycle 4.
[0388] As shown in Table 2, register R2 may be needed during cycles 2 and 4, while register R1 may be needed to store value a_mul_2 during cycles 3 and 4 and to store the sum of values during cycle 6.
[0389] In some exemplary aspects, the compiler 160 may be configured to identify the active ranges of variables in the instruction schedule according to Table 1, for example, as follows:
[0390]
[0391] Table (3)
[0392] As shown in Table 3, the active range of variable a can be between cycle 1 and cycle 6, and variable a may not be used during some of these cycles, for example, cycles 2 - 5, as shown in Table 1.
[0393] In some exemplary aspects, the compiler 160 can identify that variable a is not used during one or more unused cycles (e.g., cycles 2 - 5) within the active range between cycle 1 and cycle 6.
[0394] In some exemplary aspects, the compiler 160 can identify that the second ALU (ALU2) is idle during cycle 1 and cycle 3.
[0395] In some exemplary aspects, the compiler 160 can insert a first no - op instruction represented as a_1 with a 2 - cycle delay, which will be performed by the second ALU (ALU2) at cycle 1, for example, in order to release register R0 during cycle 2, for example, as described below.
[0396] In some exemplary aspects, the compiler 160 can insert a second no - op instruction represented as a_2 with a 2 - cycle delay, which will be performed by the second ALU (ALU2) at cycle 3, for example, in order to release register R0 during cycle 4, for example, as described below.
[0397] In some exemplary aspects, the compiler 160 can specify that register R0 stores the multiplication result b_mul_3, for example, during cycle 4. For example, the idle register R0 during cycle 4 can be used to store the multiplication result b_mul_3, for example, instead of register R2. Accordingly, register R2 may become redundant, for example, because according to Table 2, no other operations require the use of register R2.
[0398] In some exemplary aspects, the compiler 160 can determine the target code 115 based on the updated instruction schedule including the first and second no - op instructions, for example, as follows:
[0399]
[0400] Table (4)
[0401] In some exemplary aspects, the variables in the updated instruction schedule can have an updated active range, which can be different from the active range in Table 3, for example, as follows.
[0402]
[0403]
[0404] Table (5)
[0405] As shown in Table 5, variable a can be assigned to register R0 in cycles 1, 3, 5, and 6.
[0406] As shown in Table 5, variable a may not be assigned to register R0 in cycles 2 and 4. For example, because variable a can be temporarily stored by the second ALU, for example, when executing a no-op instruction.
[0407] As shown in Table 5, the multiplication result b_mul_3 can be assigned to register R0 in cycle 4. As a result, register R2 may become redundant.
[0408] In some exemplary aspects, the updated instruction scheduling can be configured to allocate two vector registers, for example, registers R0 and R1, and may not require a third register R2, for example, as follows.
[0409] Loop R0 R1 R2 0 Don't care Don't care 1 a Don't care 2 b Don't care 3 a a_mul_2 4 b_mul_3 a_mul_2 5 a Don't care 6 a sum 7 Don't care Don't care 8 Result Don't care
[0410] Table (6)
[0411] As shown in Table 6, compiler 160 can specify that register R0 stores variable a in cycles 1, 3, 5, and 6; stores the multiplication result b_mul_3 in cycle 4, and stores the result in cycle 8.
[0412] As shown in Table 6, register R2 may become redundant.
[0413] In some exemplary aspects, the number of cycles of a program executed according to the instruction set of Table 1 can be equal to the number of cycles of a program executed according to the instruction set of Table 4, for example, 9 cycles. However, the number of allocated registers can be reduced by the updated instruction scheduling, for example, from three registers to two registers. Accordingly, the instruction set of Table 4 can be implemented to provide a technical solution with improved (e.g., optimized) performance, for example, as follows:
[0414]
[0415]
[0416] Table (7)
[0417] In some exemplary aspects, as shown in Table 7, a register allocation scheme can be implemented to provide a technical solution to save vector registers for use, for example, without affecting performance.
[0418] For example, in some use cases and / or scenarios, scheduling for a program may result in unsuccessful register allocation, e.g., due to a limited number of registers.
[0419] In one instance, attempting to schedule the execution of Expression 1 according to the instruction schedule of Table 1 may result in unsuccessful register allocation, e.g., if only two registers are available. One option to address such a situation may be to relax the instruction schedule, e.g., attempt to reduce the number of active variables sharing the same execution cycle. However, this option may result in a performance degradation.
[0420] In some exemplary aspects, execution of Expression 1 according to the above register allocation scheme (e.g., using the instruction schedule of Table 4) may provide a technical solution to support successful register allocation, e.g., even if only two registers are available, e.g., while avoiding the performance degradation caused by the instruction schedule of Table 1.
[0421] Reference Figure 4 , which schematically illustrates a method for compiling code for a processor. For example, Figure 4 one or more operations of the method of Figure 1 may be performed by: a system, e.g., System 100 ( Figure 1 ); a device, e.g., Device 102 ( Figure 1 ); a server, e.g., Server 170 ( Figure 1 ); and / or a compiler, e.g., Compiler 160 ( Figure 2 ) and / or Compiler 200 (
[0422] In some exemplary aspects, as indicated at block 402, the method may include observing an active range, e.g., an active range above a predefined threshold, e.g., a relatively large active range, where the variable "x" may not be used for a period of time (e.g., a relatively long period of time). For example, Compiler 160 ( Figure 1 ) may identify one or more unused cycles in the active range, e.g., as described above.
[0423] In some exemplary aspects, as indicated at block 404, the method may include checking whether there is an idle ALU available to execute a no-op operation during the unused cycle. For example, Compiler 160 ( Figure 1 ) may check to identify the data processing unit of the processor available during one or more unused cycles, e.g., as described above.
[0424] In some exemplary aspects, as indicated at block 406, the method may include inserting one or more no-op instructions that operate on a variable "x", the one or more no-op instructions being configured to return the variable "x" to the same register after one or more delay cycles during an unused period. For example, compiler 160( Figure 1 ) may generate target code 115 based on one or more no-op instructions that will be executed during the active range of the variable, as described above, for example.
[0425] In some exemplary aspects, for example, for each variable in the source code, one or more operations of the Figure 4 method may be repeated, for example, to generate target code 115 according to a register allocation scheme that may be configured to reduce the number of registers required for the execution of the source code. For example, reducing the number of required registers may provide a technical solution to support efficient use of CPUs with a limited number of registers and / or that do not support memory overflow / fill.
[0426] Referring to Figure 5 , which schematically illustrates a method for compiling code for a processor. For example, Figure 5 one or more operations of the Figure 1 method may be performed by: a system, such as system 100( Figure 1 ); a device, such as device 102( Figure 1 ); a server, such as server 170( Figure 1 ); and / or a compiler, such as compiler 160( Figure 2 ) and / or compiler 200(
[0427] In some exemplary aspects, as indicated at block 502, the method may include compiling source code into target code. For example, the target code may be configured for execution by a target processor in a plurality of execution cycles, for example, the plurality of execution cycles including a first execution cycle, a second execution cycle after the first execution cycle, and a third execution cycle after the second execution cycle. For example, the target code may include one or more no-operation (no-op) instructions that may be configured to, for example, maintain a first variable active, such that the value of the first variable (e.g., the active value) will be available in a register of the target processor during the first execution cycle and during the third execution cycle. For example, the target code may include a first instruction that will be applied to the value of the first variable (e.g., the active value) in a register during the first execution cycle, a second instruction that will be applied to the value of a second variable in a register during the second execution cycle, and / or a third instruction that will be applied to the value of the first variable (e.g., the active value) in a register during the third execution cycle. For example, compiler 160( Figure 1) can be configured to generate target code 115 that includes one or more no - no - op instructions, e.g., as described above.
[0428] In some exemplary aspects, as indicated at block 504, the method can include outputting the target code. For example, compiler 160( Figure 1 ) can be configured to output target code 115( Figure 1 ), e.g., as described above.
[0429] Refer to Figure 6 , which schematically illustrates a manufactured product 600 in accordance with some exemplary aspects. Product 600 can include one or more tangible computer - readable ("machine - readable") non - transitory storage media 602, which can include, for example, computer - executable instructions implemented by logic 604 that are operable to cause at least one computer processor to be able to perform one or more operations at device 102( Figure 1 ), server 170( Figure 1 ) and / or compiler 160( Figure 1 ) to perform, trigger, and / or implement one or more operations and / or functionality, and / or perform, trigger, and / or implement one or more operations and / or functionality described in reference Figure 1 and / or one or more operations described herein. The phrases "non - transitory machine - readable medium" and "computer - readable non - transitory storage medium" are intended to include all computer - readable media, with the sole exception of transitory propagated signals. Figure 1 and / or compiler 160( Figure 1 ) to perform, trigger, and / or implement one or more operations and / or functionality, and / or perform, trigger, and / or implement one or more operations and / or functionality described in reference Figures 1 to 5 and / or one or more operations described herein. The phrases "non - transitory machine - readable medium" and "computer - readable non - transitory storage medium" are intended to include all computer - readable media, with the sole exception of transitory propagated signals.
[0430] In some exemplary aspects, product 600 and / or machine-readable storage medium 602 may include one or more types of computer-readable storage media capable of storing data, including volatile memory, non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. For example, machine-readable storage medium 602 may include RAM, DRAM, double data rate DRAM (DDR-DRAM), SDRAM, static RAM (SRAM), ROM, programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory (e.g., NOR or NAND flash memory), content addressable memory (CAM), polymer memory, phase change memory, ferroelectric memory, silicon oxide nitride oxide (SONOS) memory, disk, hard disk drive, etc. The computer-readable storage medium may include any suitable medium involved in downloading or transferring a computer program from a remote computer to a requesting computer via a communication link (e.g., a modem, radio, or network connection), where the computer program is carried by a data signal embedded in a carrier wave or other propagation medium.
[0431] In some exemplary aspects, logic 604 may include instructions, data, and / or code that, if executed by a machine, may cause the machine to perform methods, processes, and / or operations as described herein. The machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, etc., and may be implemented using any suitable combination of hardware, software, firmware, etc.
[0432] In some exemplary aspects, logic 604 may include or may be implemented as software, software modules, applications, programs, subroutines, instructions, instruction sets, computing code, words, values, symbols, etc. The instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. The instructions may be implemented according to a predefined computer language, manner, or syntax for instructing a processor to perform a specific function. The instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language, machine code, etc.
[0433] Example
[0434] The following examples relate to further aspects.
[0435] Example 1 includes a product that includes one or more tangible computer-readable non-transitory storage media that contain computer-executable instructions that are operable to cause at least one processor, when executed by the at least one processor, to enable a compiler: to compile source code into target code that is configured to be executed by a target processor over a plurality of execution cycles that include a first execution cycle, a second execution cycle after the first execution cycle, and a third execution cycle after the second execution cycle, where the target code includes one or more no-op instructions that are configured to keep a first variable live such that the value of the first variable will be available in a register of the target processor during the first execution cycle and during the third execution cycle, where the target code includes a first instruction that will be applied to the value of the first variable in the register during the first execution cycle, a second instruction that will be applied to the value of a second variable in the register during the second execution cycle, and a third instruction that will be applied to the value of the first variable in the register during the third execution cycle; and to output the target code.
[0436] Example 2 includes the subject matter of Example 1, and optionally where one or more no-op instructions are configured to cause the target processor to apply one or more operations to the value of the first variable that are configured to keep the value of the first variable unchanged between the first execution cycle and the third execution cycle.
[0437] Example 3 includes the subject matter of Example 1 or 2, and optionally where one or more no-op instructions are configured to cause a data processing unit of the target processor to temporarily keep the value of the first variable within the data processing unit between the first execution cycle and the third execution cycle.
[0438] Example 4 includes the subject matter of any one of Examples 1 to 3, and optionally where one or more no-op instructions are configured to cause a data processing unit of the target processor: to initiate loading the value of the first variable from a register during the first execution cycle, and to initiate storing the value of the first variable in the register during a later execution cycle before the third execution cycle.
[0439] Example 5 includes the subject matter according to any one of Examples 1 to 4, and optionally one or more of the no-op instructions include a plurality of no-op instructions, the plurality of no-op instructions including a first no-op instruction in order and a last no-op instruction in order, wherein the first no-op instruction in order is configured to cause a data processing unit of a target processor to: load a value of a first variable from a register in a first execution cycle, and apply a first operation in order configured to maintain the value of the first variable unchanged to the value of the first variable, wherein the last no-op instruction in order is configured to cause the data processing unit of the target processor to: apply a last operation in order configured to maintain the value of the first variable unchanged to an output of a previous no-op instruction, and store the value of the first variable in a register.
[0440] Example 6 includes the subject matter according to any one of Examples 1 to 5, and optionally the target code is configured such that execution of the last no-op instruction in order among one or more no-op instructions will be initiated by the data processing unit of the processor in the last no-op execution cycle before a third execution cycle, wherein the distance between the last no-op execution cycle and the third execution cycle is based on the latency of the data processing unit.
[0441] Example 7 includes the subject matter according to any one of Examples 1 to 6, and optionally the instruction, when executed, causes the compiler to configure one or more no-op instructions based on the latency of the data processing unit of the target processor.
[0442] Example 8 includes the subject matter according to any one of Examples 1 to 7, and optionally the target code is configured such that a first instruction will be executed by a first data processing unit of the target processor, and one or more no-op instructions will be executed by a second data processing unit of the target processor.
[0443] Example 9 includes the subject matter according to Example 8, and optionally the target code is configured such that execution of the first no-op instruction in order among one or more no-op instructions will be initiated by the second data processing unit in a first execution cycle.
[0444] Example 10 includes the subject matter according to Example 8 or 9, and optionally the target code is configured such that a second instruction will be executed by the second data processing unit.
[0445] Example 11 includes the subject matter according to any one of Examples 8 to 10, and optionally the target code is configured such that a third instruction will be executed by the first data processing unit.
[0446] Example 12 includes the subject matter according to any one of Examples 8 to 11, and optionally wherein the first data processing unit includes a first arithmetic logic unit (ALU), and the second data processing unit includes a second ALU.
[0447] Example 13 includes the subject matter according to any one of Examples 1 to 12, and optionally wherein the instruction, when executed, causes the compiler to: identify a plurality of active ranges corresponding to a respective plurality of variables based on the source code; and identify the variables with the identified active ranges as first variables, the active ranges including one or more unused execution cycles in which the identified variables are not used.
[0448] Example 14 includes the subject matter according to Example 13, and optionally wherein the instruction, when executed, causes the compiler to identify the identified variables as first variables based on a determination that a count of consecutive unused execution cycles in the active ranges of the identified variables is greater than a predefined threshold.
[0449] Example 15 includes the subject matter according to Example 13 or 14, and optionally wherein the instruction, when executed, causes the compiler to configure one or more no-op instructions to keep the first variables active during one or more unused execution cycles.
[0450] Example 16 includes the subject matter according to any one of Examples 13 to 15, and optionally wherein the instruction, when executed, causes the compiler to allocate registers to store the values of second variables during unused execution cycles among one or more unused execution cycles.
[0451] Example 17 includes the subject matter according to any one of Examples 1 to 16, and optionally wherein one or more of the no-op instructions includes at least one of an add-zero instruction, a shift-zero instruction, or a multiply-one instruction.
[0452] Example 18 includes the subject matter according to any one of Examples 1 to 17, and optionally wherein the source code includes Open Computing Language (OpenCL) code.
[0453] Example 19 includes the subject matter according to any one of Examples 1 to 18, and optionally wherein the instruction, when executed, causes the compiler to compile the source code into target code according to an LLVM-based compilation scheme.
[0454] Example 20 includes the subject matter according to any one of Examples 1 to 19, and optionally wherein the target code is configured to be executed by a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor.
[0455] Example 21 includes the subject matter according to any one of Examples 1 to 20, and optionally wherein the target code is configured to be executed by a target vector processor.
[0456] Example 22 includes a compiler configured to perform any one of the operations described in any one of Examples 1 to 21.
[0457] Example 23 includes a computing device configured to perform any one of the operations described in any one of Examples 1 to 21.
[0458] Example 24 includes a computing system comprising: at least one memory for storing instructions; and at least one processor for retrieving instructions from the memory and executing the instructions to cause the computing system to perform any one of the operations described in any one of Examples 1 to 21.
[0459] Example 25 includes a computing system comprising: a compiler for generating target code according to any one of the operations described in any one of Examples 1 to 21; and a processor for executing the target code.
[0460] Example 26 includes an apparatus comprising means for performing any one of the operations described in any one of Examples 1 to 21.
[0461] Example 27 includes an apparatus comprising: a memory interface; and processing circuitry configured to: perform any one of the operations described in any one of Examples 1 to 21.
[0462] Example 28 includes a method comprising any one of the operations described in any one of Examples 1 to 21.
[0463] The functions, operations, components, and / or features described herein with reference to one or more aspects may be combined with or utilized in combination with one or more other functions, operations, components, and / or features described herein with reference to one or more other aspects, or vice versa.
[0464] Although certain features have been illustrated and described herein, many modifications, substitutions, changes, and equivalents may occur to those skilled in the art. Accordingly, it is to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of this disclosure.
Claims
1. A product, comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions, the computer-executable instructions being operable to cause the at least one processor to enable a compiler when executed by the at least one processor: Compile source code into object code, the object code being configured for execution by a target processor over a plurality of execution cycles, the plurality of execution cycles including a first execution cycle, a second execution cycle after the first execution cycle, and a third execution cycle after the second execution cycle, wherein the object code includes one or more no-op (no-operation) instructions, the one or more no-op instructions being configured to keep a first variable live such that the value of the first variable will be available in a register of the target processor during the first execution cycle and during the third execution cycle, wherein the object code includes a first instruction to be applied to the value of the first variable in the register during the first execution cycle, a second instruction to be applied to the value of a second variable in the register during the second execution cycle, and a third instruction to be applied to the value of the first variable in the register during the third execution cycle; and Output the object code.
2. The product according to claim 1, wherein the one or more no-op instructions are configured to cause the target processor to apply one or more operations to the value of the first variable, the one or more operations being configured to keep the value of the first variable unchanged between the first execution cycle and the third execution cycle.
3. The product according to claim 1, wherein the one or more no-op instructions are configured to cause a data processing unit of the target processor to temporarily keep the value of the first variable within the data processing unit between the first execution cycle and the third execution cycle.
4. The product according to claim 1, wherein the one or more no-op instructions are configured to cause the data processing unit of the target processor to: initiate loading the value of the first variable from the register during the first execution cycle, and initiate storing the value of the first variable in the register during a later execution cycle before the third execution cycle.
5. The product according to claim 1, wherein the one or more no-op instructions comprise a plurality of no-op instructions, the plurality of no-op instructions comprising a first no-op instruction in order and a last no-op instruction in order, wherein the first no-op instruction in order is configured to cause a data processing unit of the target processor to: load the value of the first variable from the register in the first execution cycle, and apply a first operation in order configured to maintain the value of the first variable unchanged to the value of the first variable, wherein the last no-op instruction in order is configured to cause the data processing unit of the target processor to: apply a last operation in order configured to maintain the value of the first variable unchanged to the output of a previous no-op instruction, and store the value of the first variable in the register.
6. The product according to claim 1, wherein the target code is configured such that the execution of the last no-op instruction in order among the one or more no-op instructions will be initiated by the data processing unit of the processor in the last no-op execution cycle before the third execution cycle, wherein the distance between the last no-op execution cycle and the third execution cycle is based on the latency of the data processing unit.
7. The product according to claim 1, wherein the instruction, when executed, causes the compiler to configure the one or more no-op instructions based on the latency of the data processing unit of the target processor.
8. The product according to claim 1, wherein the target code is configured such that the first instruction will be executed by a first data processing unit of the target processor, and the one or more no-op instructions will be executed by a second data processing unit of the target processor.
9. The product according to claim 8, wherein the target code is configured such that the execution of the first no-op instruction in order among the one or more no-op instructions will be initiated by the second data processing unit in the first execution cycle.
10. The product according to claim 8, wherein the target code is configured such that the second instruction will be executed by the second data processing unit.
11. The product according to claim 8, wherein the target code is configured such that the third instruction will be executed by the first data processing unit.
12. The product according to claim 8, wherein the first data processing unit comprises a first arithmetic logic unit (ALU), and the second data processing unit comprises a second ALU.
13. The product according to any one of claims 1 to 12, wherein the instruction, when executed, causes the compiler to: identify a plurality of active ranges corresponding to a respective plurality of variables based on the source code; and identify the variables with identified active ranges as the first variable, the active ranges including one or more unused execution cycles in which the identified variables are not used.
14. The product according to claim 13, wherein the instruction, when executed, causes the compiler to identify the identified variable as the first variable based on a determination that a count of consecutive unused execution cycles in the active range of the identified variable is greater than a predefined threshold.
15. The product according to claim 13, wherein the instruction, when executed, causes the compiler to configure the one or more no-op instructions to keep the first variable active during the one or more unused execution cycles.
16. The product according to claim 13, wherein the instruction, when executed, causes the compiler to allocate the register to store the value of the second variable during an unused execution cycle among the one or more unused execution cycles.
17. The product according to any one of claims 1 to 12, wherein the one or more no-op instructions include at least one of an add-zero instruction, a shift-zero instruction, or a multiply-one instruction.
18. The product according to any one of claims 1 to 12, wherein the source code includes Open Computing Language (OpenCL) code.
19. The product according to any one of claims 1 to 12, wherein the instruction, when executed, causes the compiler to compile the source code into the target code according to an LLVM-based compilation scheme.
20. The product according to any one of claims 1 to 12, wherein the target code is configured to be executed by a Very Long Instruction Word (VLIW) single instruction / multiple data (SIMD) target processor.
21. The product according to any one of claims 1 to 12, wherein the target code is configured to be executed by a target vector processor.
22. A computing system, which comprises: at least one memory for storing instructions; and at least one processor for retrieving the instructions from the memory and for executing the instructions to cause the computing system to: compile source code into target code, the target code being configured to be executed by a target processor in a plurality of execution cycles, the plurality of execution cycles including a first execution cycle, a second execution cycle after the first execution cycle, and a third execution cycle after the second execution cycle, wherein the target code includes one or more no-operation (no-op) instructions, the one or more no-op instructions being configured to keep a first variable active such that the value of the first variable will be available in a register of the target processor in the first execution cycle and in the third execution cycle, wherein the target code includes a first instruction to be applied to the value of the first variable in the register in the first execution cycle, a second instruction to be applied to the value of a second variable in the register in the second execution cycle, and a third instruction to be applied to the value of the first variable in the register in the third execution cycle; and output the target code.
23. The computing system according to claim 22, wherein the one or more no-op instructions are configured to cause the target processor to apply one or more operations to the value of the first variable, and the one or more operations are configured to keep the value of the first variable unchanged between the first execution cycle and the third execution cycle.
24. The computing system according to claim 22, wherein the one or more no-op instructions are configured to cause a data processing unit of the target processor to: initiate loading the value of the first variable from the register in the first execution cycle, and initiate storing the value of the first variable in the register in a later execution cycle before the third execution cycle.
25. The computing system according to claim 22, comprising the target processor.
26. A method, which comprises: compiling source code into target code, the target code being configured for execution by a target processor in a plurality of execution cycles, the plurality of execution cycles including a first execution cycle, a second execution cycle after the first execution cycle, and a third execution cycle after the second execution cycle, wherein the target code includes one or more no-operation (no-op) instructions, the one or more no-op instructions being configured to keep a first variable active such that the value of the first variable will be available in a register of the target processor in the first execution cycle and in the third execution cycle, wherein the target code includes a first instruction to be applied to the value of the first variable in the register in the first execution cycle, a second instruction to be applied to the value of a second variable in the register in the second execution cycle, and a third instruction to be applied to the value of the first variable in the register in the third execution cycle; and outputting the target code.
27. The method according to claim 26, wherein the one or more no-op instructions include at least one of an add-zero instruction, a shift-zero instruction, or a multiply-one instruction.