Apparatus, system and method for compiling code for processor

By applying vector processor optimization technology in the compiler, the problems of complex loop nesting and low efficiency in the existing technology are solved, and efficient computing performance is achieved.

CN120019360APending Publication Date: 2025-05-16MOBILEYE VISION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072158.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-10-12
Publication Date
2025-05-16

Smart Images

  • Figure CN120019360A_ABST
    Figure CN120019360A_ABST
Patent Text Reader

Abstract

For example, the disclosure provides a compiler that may be configured to identify a loop nest based on source code to be compiled into target code to be executed by a target processor, the loop nest including a plurality of loops including at least a first loop and a second loop nested in the first loop, the first loop comprises at least one first loop instruction outside the second loop; and generating address generation unit (AGU) configuration code to configure an AGU of the target processor based on the first loop instruction, where the AGU configuration code is to configure a first dimension of the AGU based on the first loop, and configuring a second dimension of the AGU based on the second loop to configure a memory access operation to be performed at a start of the second loop or at an end of the second loop based on the first loop instruction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 415,308, filed on October 12, 2022, entitled “APPARATUS, SYSTEM, AND METHOD OF VECTOR PROCESSING,” the entire disclosure of which is incorporated herein by reference. Background Art

[0003] The compiler may be configured to compile source code into object code configured for execution by the processor.

[0004] There is a need to provide technical solutions to support efficient processing functionality. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity of presentation. In addition, reference numerals may be repeated in the drawings to indicate corresponding or similar elements. The drawings are listed below.

[0006] Figure 1 is a schematic block diagram illustration of a system according to some exemplary aspects.

[0007] Figure 2 is a schematic illustration of a compiler according to some exemplary aspects.

[0008] Figure 3 is a schematic illustration of a vector processor according to some exemplary aspects.

[0009] Figure 4 is a schematic illustration of an implementation scheme for performing a latch-store operation in a loop nest according to some exemplary aspects.

[0010] Figure 5 is a schematic illustration of an implementation scheme for performing a pre-header load or store operation in a loop nest according to some exemplary aspects.

[0011] Figure 6 is a schematic flow chart illustration of a method of compiling code for a processor according to some exemplary aspects.

[0012] Figure 7 is a schematic illustration of a product according to some exemplary aspects. DETAILED DESCRIPTION

[0013] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of some aspects. However, it will be appreciated by those of ordinary skill in the art that some aspects may be practiced without these specific details. In other cases, well-known methods, procedures, components, units and / or circuits are not described in detail to avoid obscuring the discussion.

[0014] Some portions of the following detailed description are presented in terms of algorithms and symbolic representations of operations on data bits or binary digital signals within a computer memory. These algorithmic descriptions and representations may be techniques used by those skilled in the data processing arts to convey the substance of their work to others skilled in the art.

[0015] An algorithm is here and generally considered to be a self-consistent sequence of acts or operations leading to a desired result. These include physical manipulations of physical quantities. Typically, but not necessarily, these quantities capture forms of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for common sense reasons, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. It should be understood, however, that all of these terms and similar terms should be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities.

[0016] Discussions herein utilizing terms such as, for example, "process," "compute," "calculate," "determine," "create," "analyze," "verify," and the like may refer to the operations and / or processes of a computer, computing platform, computing system, or other electronic computing device that manipulates data represented as physical (e.g., electronic) quantities within the computer's registers and / or memories and / or transforms that data into other data similarly represented as physical quantities within the computer's registers and / or memories or other information storage media that may store instructions for performing operations and / or processes.

[0017] As used herein, the terms “plurality” and “a plurality” include, for example, “a plurality” or “two or more.” For example, “a plurality of items” includes two or more items.

[0018] References to "one aspect," "an aspect," "exemplary aspect," "various aspects," etc. indicate that the aspects so described may include particular features, structures, or characteristics, but not every aspect necessarily includes the particular features, structures, or characteristics. Furthermore, repeated use of the phrase "in one aspect" does not necessarily refer to the same aspect, although it may.

[0019] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe common objects merely indicates that different instances of the same object are being referenced and is not intended to imply that the objects so described must be in a given sequence in time, space, ranking, or in any other manner.

[0020] For example, some aspects may capture the form of entirely hardware aspects, entirely software aspects, or aspects including both hardware and software elements.Some aspects may be implemented in software, including but not limited to firmware, resident software, microcode, etc.

[0021] Furthermore, some aspects may be captured in the form of a computer program product that can be accessed from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For example, a computer-usable or computer-readable medium can be or can include any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with an instruction execution system, device, or apparatus.

[0022] In some exemplary aspects, the medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium.

[0023] In some exemplary aspects, a data processing system suitable for storing and / or executing program code may include at least one processor coupled directly or indirectly to a memory element, for example, via a system bus. The memory element may include, for example, local memory employed during actual execution of the program code, a mass storage device, and a cache memory that may provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from a mass storage device during execution.

[0024] In some exemplary aspects, input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system directly or through an intermediate I / O controller. In some exemplary aspects, a network adapter can be coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices, such as through an intermediate private or public network. In some exemplary aspects, modems, cable modems, and Ethernet cards are exemplary examples of network adapter types. Other suitable components can be used.

[0025] Some aspects may be used in connection with various devices and systems, such as computing devices, computers, mobile computers, non-mobile computers, server computers, and the like.

[0026] As used herein, the term "circuitry" may refer to, be a part of, or include an application specific integrated circuit (ASIC), an integrated circuit, an electronic circuit, a processor (shared, dedicated, or grouped), and / or memory (shared. Dedicated or grouped) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable hardware components that provide the described functionality. In some aspects, some functions associated with the circuitry may be implemented by one or more software or firmware modules. In some aspects, the circuitry may include logic that is at least partially operable in hardware.

[0027] The term "logic" may refer to, for example, computing logic embedded in the circuit system of a computing device and / or computing logic stored in the memory of a computing device. For example, the logic may be accessed by a processor of a computing device to execute the computing logic to perform computing functions and / or operations. In one example, the logic may be embedded in various types of memory and / or firmware, such as silicon blocks of various chips and / or processors. The logic may be included in various circuit systems and / or implemented as part of various circuit systems, such as processor circuit systems, control circuit systems, and / or the like. In one example, the logic may be embedded in volatile memory and / or non-volatile memory, including random access memory, read-only memory, programmable memory, magnetic memory, flash memory, persistent memory, etc. The logic may be executed by one or more processors using memory (e.g., registers, lags, buffers, and / or the like) coupled to one or more processors, for example, executing the logic as needed.

[0028] Reference now Figure 1 , which schematically illustrates a block diagram of a system 100 according to some exemplary aspects.

[0029] like Figure 1 As shown, in some demonstrative aspects, system 100 may include a computing device 102 .

[0030] In some demonstrative aspects, device 102 may be implemented using suitable hardware components and / or software components, such as processors, controllers, memory units, storage units, input units, output units, communication units, operating systems, applications, and the like.

[0031] In some demonstrative aspects, device 102 may comprise, for example, a computer, a mobile computing device, a non-mobile computing device, a laptop computer, a notebook computer, a tablet computer, a handheld computer, a personal computer (PC), or the like.

[0032] In some exemplary aspects, device 102 may include, for example, one or more of the following: processor 191, input unit 192, output unit 193, memory unit 194, and / or storage unit 195. Device 102 may optionally include other suitable hardware components and / or software components. In some exemplary aspects, some or all components of one or more of devices in device 102 may be enclosed in a common housing or packaging and may be interconnected or operably associated using one or more wired or wireless links. In other aspects, components of one or more of devices in device 102 may be distributed in multiple or separate devices.

[0033] In some exemplary aspects, processor 191 may include, for example, a central processing unit (CPU), a digital signal processor (DSP), one or more processor cores, a single core processor, a dual core processor, a multi-core processor, a microprocessor, a host processor, a controller, multiple processors or controllers, a chip, a microchip, one or more circuits, a circuit system, a logic unit, an integrated circuit (IC), an application specific IC (ASIC), or any other suitable general-purpose or specific processor or controller. Processor 191 may execute, for example, instructions of an operating system (OS) of device 102 and / or instructions of one or more suitable applications.

[0034] In some exemplary aspects, input unit 192 may include, for example, a keyboard, a keypad, a mouse, a touch screen, a touch pad, a trackball, a stylus, a microphone, or other suitable pointing device or input device. Output unit 193 may include, for example, a monitor, a screen, a touch screen, a flat panel display, a light emitting diode (LED) display unit, a liquid crystal display (LCD) display unit, a plasma display unit, one or more audio speakers or headphones, or other suitable output devices.

[0035] In some exemplary aspects, memory unit 194 includes, for example, random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous DRAM (SD-RAM), flash memory, volatile memory, non-volatile memory, cache memory, buffer, short-term memory unit, long-term memory unit, or other suitable memory unit. Storage unit 195 may include, for example, a hard disk drive, a solid-state drive (SSD), or other suitable removable or non-removable storage unit. Memory unit 194 and / or storage unit 195 may, for example, store data processed by device 102.

[0036] In some demonstrative aspects, device 102 may be configured to communicate with one or more other devices via at least one network 103 (eg, a wireless and / or wired network).

[0037] In some exemplary aspects, network 103 may include a wired network, a local area network (LAN), a wireless network, a wireless LAN (WLAN) network, a radio network, a cellular network, a WiFi network, an IR network, a Bluetooth (BT) network, and the like.

[0038] In some demonstrative aspects, device 102 may be configured to perform and / or execute one or more operations, modules, processes, procedures, and / or the like, e.g., as described herein.

[0039] In some demonstrative aspects, device 102 may include compiler 160, which may be configured to generate object code 115 based on source code 112, for example, as described below.

[0040] In some exemplary aspects, compiler 160 may be configured to translate source code 112 into target code 115, eg, as described below.

[0041] In some demonstrative aspects, compiler 160 may include or may be implemented as software, a software module, an application, a program, a subroutine, instructions, an instruction set, computing code, words, values, symbols, and / or the like.

[0042] In some exemplary aspects, source code 112 may include computer code written in a source language.

[0043] In some exemplary aspects, the source language may include a programming language. For example, the source language may include a high-level programming language, such as, for example, C language, C++ language and / or the like.

[0044] In some exemplary aspects, target code 115 may include computer code written in a target language.

[0045] In some exemplary aspects, the target language can include a low-level language such as, for example, assembly language, object code, machine code, or the like.

[0046] In some exemplary aspects, object code 115 may include one or more purpose files, which may, for example, create and / or form an executable program.

[0047] In some exemplary aspects, the executable program may be configured to be executed on a target computer. For example, the target computer may include specific computer hardware, a specific machine and / or a specific operating system.

[0048] In some exemplary aspects, the executable program may be configured to be executed on processor 180, eg, as described below.

[0049] In some demonstrative aspects, processor 180 may include a vector processor 180, eg, as described below. In other aspects, processor 180 may include any other type of processor.

[0050] Some exemplary aspects are described herein with respect to a compiler (e.g., compiler 160) that is configured to compile source code 112 into target code 115 that is configured to be executed by a vector processor 180, e.g., as described below. In other aspects, a compiler (e.g., compiler 160) is configured to compile source code 112 into target code 115 that is configured to be executed by any other type of processor 180.

[0051] In some demonstrative aspects, processor 180 may be implemented as part of device 102 .

[0052] In other aspects, processor 180 may be implemented as part of any other device separate from device 102 , for example.

[0053] In some demonstrative aspects, vector processor 180 (also referred to as an "array processor") may include a processor that may be configured to process an entire vector in one instruction, eg, as described below.

[0054] In other aspects, the executable program may be configured to be executed on any other additional or alternative type of processor.

[0055] In some exemplary aspects, vector processor 180 may be designed to support high-performance image and / or vector processing. For example, vector processor 180 may be configured to process 1 / 2 / 3 / 4D arrays and / or floating point arrays of fixed point data very quickly and / or efficiently.

[0056] In some exemplary aspects, vector processor 180 can be configured to process arbitrary data, such as structures with pointers to structures. For example, vector processor 180 can include a scalar processor to calculate non-vector data, such as assuming that the non-vector data is minimal.

[0057] In some demonstrative aspects, compiler 160 may be implemented as a native application to be executed by device 102. For example, memory unit 194 and / or storage unit 195 may store instructions obtained in compiler 160, and / or processor 191 may be configured to execute instructions obtained in compiler 160 and / or perform one or more calculations and / or processes of compiler 160, e.g., as described below.

[0058] In other aspects, compiler 160 may comprise a remote application to be executed by any suitable computing system (eg, server 170 ).

[0059] In some exemplary aspects, server 170 may include at least a remote server, a network-based server, a cloud server, and / or any other server.

[0060] In some exemplary aspects, server 170 may include a suitable memory and / or storage unit 174 having stored thereon instructions derived from compiler 160 and a suitable processor 171 for executing the instructions, e.g., as described below.

[0061] In some exemplary aspects, compiler 160 may include a combination of remote applications and local applications.

[0062] In one example, compiler 160 may be downloaded and / or received by a user of device 102 from another computing system (e.g., server 170) such that compiler 160 may be executed locally by the user of device 102. For example, instructions may be received and stored temporarily in a memory or any suitable short-term storage or buffer of device 102, e.g., prior to execution by processor 191 of device 102.

[0063] In another example, compiler 160 may include a client module to be executed locally by device 102 and a server module to be executed by server 170. For example, the client module may include and / or may be implemented as a local application, a web application, a website, a web client, e.g., a hypertext markup language (HTML) web application, etc.

[0064] For example, one or more first operations of compiler 160 may be performed locally, such as by device 102 , and / or one or more second operations of compiler 160 may be performed remotely, such as by server 170 .

[0065] In other aspects, compiler 160 may include or be implemented by any other suitable computing arrangement and / or scheme.

[0066] In some demonstrative aspects, system 100 may include an interface 110 (eg, a user interface) to interface between a user of device 102 and one or more elements of system 100 (eg, compiler 160).

[0067] In some demonstrative aspects, interface 110 may be implemented using any suitable hardware components and / or software components, such as a processor, a controller, a memory unit, a storage unit, an input unit, an output unit, a communication unit, an operating system, and / or an application.

[0068] In some aspects, interface 110 may be implemented as part of any suitable module, system, device, or component of system 100 .

[0069] In other aspects, interface 110 may be implemented as a separate element of system 100 .

[0070] In some demonstrative aspects, interface 110 may be implemented as part of device 102. For example, interface 110 may be associated with device 102 and / or included as part of the device.

[0071] In one example, interface 110 can be implemented as part of any suitable application, such as middleware and / or device 102. For example, interface 110 can be implemented as part of compiler 160 and / or part of the OS of device 102.

[0072] In some demonstrative aspects, interface 110 may be implemented as part of server 170. For example, interface 110 may be associated with server 170 and / or included as part of the server.

[0073] In one example, interface 110 may include or be part of: a web-based application, a website, a web page, a plug-in, an ActiveX control, a rich content component (eg, a Flash or Shockwave component), or the like.

[0074] In some exemplary aspects, interface 110 may be associated therewith and / or may include, for example, a gateway (GW) 113 and / or an application programming interface (API) 114, for example, to transmit information and / or communicate between elements of system 100 and / or to one or more other parties (e.g., internal or external parties), users, applications and / or systems.

[0075] In some aspects, interface 110 may include any suitable graphical user interface (GUI) 116 and / or any other suitable interface.

[0076] In some demonstrative aspects, interface 110 may be configured to receive source code 112 from, for example, a user of device 102 via GUI 116 and / or API 114 .

[0077] In some exemplary aspects, interface 110 may be configured to transfer source code 112 to, for example, compiler 160 , for example, to generate object code 115 , for example, as described below.

[0078] refer to Figure 2 , which schematically illustrates a compiler 200 according to some exemplary aspects. For example, the compiler 160 ( Figure 1 ) may implement one or more elements of compiler 200 and / or may perform one or more operations and / or functionalities of compiler 200.

[0079] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may be configured to generate target code 233, for example, by compiling source code 212 in a source language.

[0080] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may include a front end 210 configured to receive and analyze source code 212 in a source language.

[0081] In some exemplary aspects, front end 210 may be configured to generate intermediate code 213 , for example, based on source code 212 .

[0082] In some exemplary aspects, intermediate code 213 may comprise a lower-level representation of source code 212 .

[0083] In some exemplary aspects, front end 210 can be configured to perform, for example, lexical analysis, syntactic analysis, semantic analysis, and / or any other additional or alternative types of analysis of source code 212 .

[0084] In some exemplary aspects, front end 210 can be configured to identify errors and / or problems using the results of the analysis of source code 212. For example, front end 210 can be configured to generate error information, e.g., including error and / or warning messages, which can identify a location in source code 212, e.g., where an error or problem is detected.

[0085] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may include a middle end 220 configured to receive and process intermediate code 213 and generate adjusted (eg, optimized) intermediate code 223 .

[0086] In some demonstrative aspects, middle end 220 may be configured to perform one or more adjustments (eg, optimizations) to intermediate code 213 , eg, to generate adjusted intermediate code 223 .

[0087] In some exemplary aspects, middle end 220 can be configured to perform one or more optimizations on intermediate code 213 , eg, independent of the type of target computer used to execute target code 233 .

[0088] In some exemplary aspects, middle end 220 can be implemented to support the use of optimized intermediate code 223, eg, for different machine types.

[0089] In some exemplary aspects, middle end 220 may be configured to optimize the intermediate representation of intermediate code 223 , for example, to improve the performance and / or quality of the generated target code 233 .

[0090] In some exemplary aspects, one or more optimizations of intermediate code 213 may include, for example, inline expansion, dead code elimination, constant propagation, loop transformation, parallelization, and / or the like.

[0091] In some exemplary aspects, such as Figure 2 As shown, the compiler 200 may include a back end 230 configured to receive and process the adjusted intermediate code 213 , and generate a target code 233 based on the adjusted intermediate code 213 .

[0092] In some exemplary aspects, backend 230 may be configured to perform one or more operations and / or processes that may be specific to a target computer used to execute target code 233. For example, backend 230 may be configured to process optimized intermediate code 213 by applying analysis, transformation, and / or optimization operations to adjusted intermediate code 213, which operations may be configured, for example, based on a target computer used to execute target code 233.

[0093] In some exemplary aspects, the one or more analysis, transformation, and / or optimization operations applied to the adjusted intermediate code 213 may include, for example, resource and storage decisions, such as register allocation, instruction scheduling, and / or the like.

[0094] In some exemplary aspects, the object code 233 may include target-dependent assembly code that may be specific to a target computer used to execute the object code 233 and / or a target operating system of the target computer.

[0095] In some exemplary aspects, the object code 233 may include code for a processor (e.g., vector processor 180 ( Figure 1 ))'s target-dependent assembly code.

[0096] In some exemplary aspects, compiler 200 may include a vector microcode processor (VMP) open computing language (OpenCL) compiler, for example, as described below. In other aspects, compiler 200 may include any other type of vector processor compiler, or may be implemented as part of any other type of vector processor compiler.

[0097] In some exemplary aspects, the VMP OpenCL compiler may include a low-level virtual machine (LLVM)-based compiler that may be configured according to an LLVM-based compilation scheme, for example, to reduce OpenCL C code to VMP accelerator assembly code, for example, suitable for use by vector processor 180 ( Figure 1 )implement.

[0098] In some exemplary aspects, compiler 200 may include one or more techniques that may be required to compile code into a format suitable for a VMP architecture, for example, in addition to an open source LLVM compiler pass.

[0099] In some exemplary aspects, FE 210 may be configured to parse OpenCL C code and translate it, for example, via an abstract syntax tree (AST), into, for example, an LLVM intermediate representation (IR).

[0100] In some exemplary aspects, compiler 200 may include a dedicated API, e.g., to detect the correct pattern for compiler pattern matching, e.g., a pattern suitable for VMP. For example, VMP may be configured as a complex instruction set computer (CISC) machine that implements a very complex instruction set architecture (ISA) that may be difficult to target from standard C code. Accordingly, compiler pattern matching may not be able to easily detect the correct pattern, and for such cases, the compiler may require a dedicated API.

[0101] In some exemplary aspects, FE 210 may implement one or more vendor extension builtins that may target a VMP-specific ISA, for example, in addition to standard OpenCL builtins that may be optimized for VMP machines.

[0102] In some exemplary aspects, FE 210 may be configured to implement OpenCL constructs and / or work-item functionality.

[0103] In some exemplary aspects, ME 220 may be configured to process LLVM IR code, which may be generic and target-independent, e.g., although it may include one or more hooks for a specific target architecture.

[0104] In some demonstrative aspects, ME 220 may perform one or more custom passes, for example, to support a VMP architecture, for example, as described below.

[0105] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of control flow graph (CFG) linearization analysis, e.g., as described below.

[0106] In some exemplary aspects, the CFG linearization analysis can be configured to linearize the code, for example, by converting if statements to select modes, for example, where the VMP vector code does not support standard control flow.

[0107] In one example, ME 220 may receive a given code, for example, as follows:

[0108]

[0109] According to this example, ME 220 may be configured to apply CFG linearization analysis to a given code, for example, as follows:

[0110] tmpA=A+5;

[0111] tmpB = B*2;

[0112] mask=x>0;

[0113] A=Select mask,tmpA,A

[0114] B=Select not mask,tmpB,B

[0115] Example (1)

[0116] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of automatic vectorization analysis, e.g., as described below.

[0117] In some exemplary aspects, the auto-vectorization analysis may be configured to vectorize (eg, auto-vectorize) a given code, for example, to exploit the vector capabilities of the VMP.

[0118] In some exemplary aspects, ME 220 may be configured to perform automatic vectorization analysis, e.g., to vectorize code into scalar form. For example, some or all operations of automatic vectorization analysis may not be performed, such as when the code is already provided in vectorized form.

[0119] In some exemplary aspects, for example, in some use cases and / or scenarios, a compiler may not always be able to auto-vectorize code, for example, due to data dependencies between loop iterations.

[0120] In one example, ME 220 may receive a given code, for example, as follows:

[0121]

[0122] According to this example, ME 220 may be configured to perform CFG auto-vectorization analysis by applying a first transformation, for example, as follows:

[0123]

[0124] Example (2a)

[0125] For example, ME 220 may be configured to perform CFG auto-vectorization analysis by applying a second transformation, e.g., after a first transformation, e.g., as follows:

[0126]

[0127] Example (2b)

[0128] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of scratch pad memory cycle access analysis (SPMLAA), e.g., as described below.

[0129] In some exemplary aspects, the SPMLAA may define a processing block (PB), for example, that should later be outlined and compiled for VMP.

[0130] In some exemplary aspects, a processing block may include an accelerated loop that may be executed by a vector unit of a VMP.

[0131] In some exemplary aspects, a PB (eg, each PB) may include memory references. For example, some or all memory accesses may refer to a local memory bank.

[0132] In some exemplary aspects, the VMP may enable the AGU (e.g., as described below with reference to Figure 3 AGU 320) and scatter-gather unit (SG) described above are used to access the memory bank.

[0133] In some exemplary aspects, the AGU can be pre-configured, for example, before a loop is executed. For example, a loop trip count can be calculated, for example, before running a processing block.

[0134] In some exemplary aspects, image references can be created at this stage, eg, some or all image references, and strides and offsets can then be calculated, eg, per-dimension strides and offsets for each reference.

[0135] In some exemplary aspects, ME 220 may be configured to perform one or more operations of AGU planner analysis, e.g., as described below.

[0136] In some exemplary aspects, the AGU planner analysis can include an iterator specification that can cover image references from an entire processing block, eg, all image references.

[0137] In some exemplary aspects, an iterator may cover a single reference or a group of references.

[0138] In some exemplary aspects, one or more memory references may be combined via a shuffle instruction and / or reuse the same access, and / or preserve values ​​read from a previous iteration.

[0139] In some exemplary aspects, other memory references, such as those without a linear access pattern, may be processed using a scatter-gather (SG) unit, which may have a performance penalty, such as because it may need to maintain indexes and / or masks.

[0140] In some exemplary aspects, a plan may be configured as an arrangement of iterators in a processing block. For example, a processing block may, for example, theoretically have multiple plans.

[0141] In some exemplary aspects, the AGU planner analysis can be configured to construct all possible plans for all PBs and select a combination, eg, the best combination, from among all valid combinations.

[0142] In some exemplary aspects, the total number of iterators in a valid combination may be limited, eg, not to exceed the number of available AGUs on the VMP.

[0143] In some exemplary aspects, one or more parameters may be defined for an iterator (e.g., for each iterator), e.g., including stride, width, and / or cardinality, e.g., as part of an AGU planner analysis. For example, a minimum-maximum range for an iterator may be defined dimensionally, e.g., in each dimension, e.g., as part of an AGU planner analysis.

[0144] In some exemplary aspects, the AGU planner analysis can be configured to track and evaluate memory references to the image, eg, each memory reference, eg, to understand its access pattern.

[0145] In one example, according to Example 2a / 2b, image "a" as a base address can be accessed with 64 iterations using a step size of 32 bytes.

[0146] In some exemplary aspects, LLVM can include scalar evaluation analysis (SCEV) that can compute access patterns, for example, to understand each image reference.

[0147] In some exemplary aspects, ME 220 may exploit the masking capabilities of the AGU, eg, to avoid maintaining induction variables, which may have a performance penalty.

[0148] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of rewrite analysis, e.g., as described below.

[0149] In some exemplary aspects, the rewrite analysis may be configured to transform the code of a processing block, for example, when setting up iterators and / or modifying memory access instructions.

[0150] In some exemplary aspects, the setup of iterators (e.g., all iterators) can be implemented in the IR in a target-specific intrinsic function. For example, the setup of iterators can reside in the pre-header of the outermost loop.

[0151] In some exemplary aspects, the rewrite analysis may include a loop completion analysis, eg, as described below.

[0152] In some exemplary aspects, the code may be compiled with the goal that substantially all computations should be performed within the innermost loop.

[0153] For example, loop finishing analysis may promote instructions, for example, to move operations performed after the last iteration of the loop into the loop.

[0154] For example, loop finishing analysis may sink instructions, eg, to move operations performed before the first iteration of the loop into the loop.

[0155] For example, loop finishing analysis may hoist instructions and / or sink instructions, eg, such that substantially all instructions from an outer loop are moved to an innermost loop.

[0156] For example, loop completion analysis may be configured to provide technical solutions to support VMP iterators, for example, to work only on perfectly nested loops.

[0157] For example, loop completion analysis may lead to a situation where there are no instructions between "for" statements that make up a loop, e.g., to support VMP iterators, which cannot emulate such a situation.

[0158] In some exemplary aspects, loop completion analysis may be configured to collapse nested loops into a single collapsed loop.

[0159] In one example, ME 220 may receive a given code, for example, as follows:

[0160]

[0161]

[0162] According to this example, ME 220 may be configured to perform loop completion analysis to collapse nested loops in the code into a single collapsed loop, for example, as follows:

[0163]

[0164] Example (3)

[0165] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of vector loop delimitation analysis, eg, as described below.

[0166] In some exemplary aspects, the vector loop demarcation analysis can be configured to partition the code between the scalar subsystem and the vector subsystem, for example, as described below with reference to Figure 3 The vector processing block 310 ( Figure 3 ) and scalar processor 330( Figure 3 )between.

[0167] In some exemplary aspects, a VMP accelerator may include scalar and / or vector subsystems, e.g., as described below. For example, each of the subsystems may have a different computational unit / processor. Accordingly, scalar code may be compiled on a scalar compiler (e.g., an SSC compiler), and / or accelerated vector code may run on a VMP vector processor.

[0168] In some exemplary aspects, vector loop demarcation analysis can be configured to create separate functions for accelerating loop bodies of vector code. For example, these functions can be marked for VMP and / or can proceed to the VMP backend, for example, while the rest of the code can be compiled by the SSC compiler.

[0169] In some exemplary aspects, one or more portions of a vector loop (e.g., configuration of a vector unit and / or initialization of vector registers) may be performed by a scalar unit. However, these portions may be performed at a later stage, e.g., by backfilling the scalar code, e.g., because the scalar code may still be in LLVM IR before being processed by the SSC compiler.

[0170] In some exemplary aspects, BE 230 may be configured to translate LLVM IR into machine instructions. For example, BE 230 may not be target agnostic and may be familiar with target specific architectures and optimizations, for example, compared to ME 220 which may be agnostic to target specific architectures.

[0171] In some exemplary aspects, BE 230 may be configured to perform one or more analyses that may be specific to the target machine (eg, a VMP machine) to which the code is being downgraded, for example, even though BE 230 may use a general-purpose LLVM.

[0172] In some exemplary aspects, BE 230 may be configured to perform one or more operations of instruction degradation analysis, eg, as described below.

[0173] In some exemplary aspects, instruction degradation analysis may be configured to translate LLVM IR into target-specific instruction machine IR (MIR), for example, by translating LLVM IR into a directed acyclic graph (DAG).

[0174] In some exemplary aspects, the DAG may undergo a legalization process for instructions, such as based on data types and / or VMP instructions, which may be supported by the VMP HW.

[0175] In some exemplary aspects, instruction demotion analysis may be configured to, for example, perform a pattern matching process after a legalization process of instructions, for example, to demotion nodes (eg, each node) in a DAG to, for example, VMP-specific machine instructions.

[0176] In some exemplary aspects, instruction degradation analysis may be configured to generate a MIR, for example, after a pattern matching process.

[0177] In some exemplary aspects, instruction demotion analysis may be configured to degrade instructions according to a machine application binary interface (ABI) and / or calling convention.

[0178] In some exemplary aspects, BE 230 can be configured to perform one or more operations of a cell balance analysis, eg, as described below.

[0179] In some exemplary aspects, the unit balancing analysis may be configured to balance instructions among VMP computing units, for example, as described below with reference to Figure 3 The data processing unit 316 ( Figure 3 )between.

[0180] In some exemplary aspects, the cell balance analysis may be aware of some or all available arithmetic transformations, and / or may perform transformations according to an optimal algorithm.

[0181] In some exemplary aspects, BE 230 may be configured to perform one or more operations of a modulo scheduler (pipeliner) analysis, eg, as described below.

[0182] In some exemplary aspects, the pipeliner may be configured to schedule instructions according to one or more constraints (e.g., data dependencies, resource bottlenecks, and / or any other constraints), for example using a swing modulo scheduling (SMS) heuristic and / or any other additional and / or alternative heuristics.

[0183] In some exemplary aspects, the pipeliner can be configured to schedule a set of very long instruction word (VLIW) instructions (eg, of initiation intervals (II)) over which a program will iterate, such as during a steady state.

[0184] In some exemplary aspects, a performance metric may be measured, which may be based on the number of cycles a typical loop may execute, for example, as follows:

[0185] (input data size in bytes)*II / (bytes consumed / produced per iteration)

[0186] In some exemplary aspects, the pipeliner can attempt to minimize II as much as possible, for example, to improve performance.

[0187] In some exemplary aspects, the pipeliner can be configured to calculate a minimum II and schedule accordingly. For example, if the pipeliner fails to schedule, the pipeliner can attempt to increase the II and retry scheduling, for example, until a predefined II threshold is violated.

[0188] In some demonstrative aspects, BE 230 may be configured to perform one or more operations of register allocation analysis, eg, as described below.

[0189] In some exemplary aspects, register allocation analysis can be configured to attempt to assign registers in an efficient (eg, optimal) manner.

[0190] In some exemplary aspects, register allocation analysis may assign values ​​to bypass vector registers, general purpose vector registers, and / or scalar registers.

[0191] In some exemplary aspects, the values ​​may include private variables, constants, and / or values ​​that rotate across iterations.

[0192] In some exemplary aspects, register allocation analysis may implement an optimal heuristic that fits one or more VMP register file (regfile) constraints. For example, in some use cases, register allocation analysis may not use standard LLVM register allocation.

[0193] In some exemplary aspects, in some cases, register allocation analysis may fail, which may mean that the loop cannot be compiled. Accordingly, register allocation analysis may implement a retry mechanism that may return to the modulo scheduler and may attempt to reschedule the loop, e.g., with an increased launch interval. For example, in many cases, increasing the launch interval may reduce register starvation and / or may support compilation of vector loops.

[0194] In some demonstrative aspects, BE 230 may be configured to perform one or more operations of SSC configuration analysis, eg, as described below.

[0195] In some exemplary aspects, the SSC configuration analysis may be configured to set a configuration for executing a kernel, such as an AGU configuration.

[0196] In some exemplary aspects, SSC configuration analysis may be performed at a later stage, such as due to configurations being calculated after legalization, register allocation analysis, and / or modulo scheduling analysis.

[0197] In some exemplary aspects, the SSC configuration analysis can include a zero overhead loop (ZOL) mechanism in a vector loop. For example, the ZOL mechanism can configure loop trip counts based on access patterns of memory references in the loop, e.g., to avoid running instructions that check loop exit conditions for each iteration.

[0198] In some exemplary aspects, a VMP compilation flow may include one or more (e.g., a small number) of steps that may be called during the compilation flow in a test library (testlib) (e.g., a wrapper script for compilation, execution, and / or program testing). For example, these steps may be performed outside of the LLVM compiler.

[0199] In some exemplary aspects, a PCB Hardware Description Language (PHDL) simulator can be implemented to perform one or more roles of an assembler, an encoder, and / or a linker.

[0200] In some exemplary aspects, compiler 200 can be configured to provide technical solutions to support robustness, which can enable compilation of a wide range of loop selections with HW limitations. For example, compiler 200 can be configured to support technical solutions that may not generate verification errors.

[0201] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support programmability, which may provide users with the ability to express code in a variety of ways that may compile correctly to a VMP architecture.

[0202] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support an improved user experience, which may allow a user to debug and / or profile code. For example, the improved user experience may provide informative error messages, reporting tools, and / or profiling tools.

[0203] In some exemplary aspects, compiler 200 can be configured to provide technical solutions to support improved performance, for example, to optimize VMP assembly code and / or iterator access, which may result in faster execution. For example, improved performance can be achieved through high utilization compute units and using their complex CISC.

[0204] refer to Figure 3 , which schematically illustrates a vector processor 300 according to some exemplary aspects. For example, the vector processor 180 ( Figure 1 ) may implement one or more elements of the vector processor 300 and / or may perform one or more operations and / or functionalities of the vector processor 300.

[0205] In some exemplary aspects, vector processor 300 may comprise a vector microcode processor (VMP).

[0206] In some exemplary aspects, vector processor 300 may include a wide vector machine, eg, supporting a very long instruction word (VLIW) architecture and / or a single instruction / multiple data (SIMD) architecture.

[0207] In some exemplary aspects, vector processor 300 may be configured to provide a technical solution to support high performance for short integer types, which may be common in, for example, computer vision and / or deep learning algorithms.

[0208] In other aspects, the vector processor 300 may include any other type of vector processor, and / or may be configured to support any other additional or alternative functionality.

[0209] In some exemplary aspects, such as Figure 3 As shown, the vector processor 300 may include a vector processing block (vector processor) 310, a scalar processor 330, and a direct memory access (DMA) 340, for example, as described below.

[0210] In some exemplary aspects, such as Figure 3 As shown, the vector processing block 310 may be configured to process (eg, efficiently process) image data and / or vector data. For example, the vector processing block 310 may be configured to use a vector computing unit, for example, to accelerate computation.

[0211] In some exemplary aspects, scalar processor 330 may be configured to perform scalar calculations. For example, scalar processor 330 may be used as "glue logic" for a program that includes vector calculations. For example, some (e.g., even most) of the calculations of a program may be performed by vector processing block 310. However, several tasks (e.g., some basic tasks) (e.g., scalar calculations) may be performed by scalar processor 330.

[0212] In some demonstrative aspects, DMA 340 may be configured to interface with one or more memory elements in a chip including vector processor 300 .

[0213] In some demonstrative aspects, DMA 340 may be configured to read input from main memory, and / or write output to main memory.

[0214] In some exemplary aspects, scalar processor 330 and vector processing block 310 may use respective local memories to process data.

[0215] In some exemplary aspects, such as Figure 3 As shown, the vector processor 300 may include an extractor and decoder 350 , which may be configured to control the scalar processor 330 and / or the vector processing block 310 .

[0216] In some exemplary aspects, operations of scalar processor 330 and / or vector processing block 310 may be triggered by instructions stored in program memory 352 .

[0217] In some demonstrative aspects, DMA 340 may be configured to transfer data in parallel with the execution of program instructions in memory 352, for example.

[0218] In some exemplary aspects, DMA 340 may be controlled by software, such as via configuration registers, rather than instructions, for example, and accordingly may be considered a second “thread” of execution in vector processor 300 .

[0219] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include one or more data processing units, e.g., a group of data processing units, e.g., as described below.

[0220] In some exemplary aspects, a data processing unit may include hardware configured to perform calculations, such as an arithmetic logic unit (ALU).

[0221] In one example, the data processing unit may be configured to add numbers and / or store numbers in memory.

[0222] In some exemplary aspects, the data processing unit may be controlled by commands encoded in, for example, program memory 352 and / or configuration registers. For example, the configuration registers may be memory mapped and writeable by memory storage commands of scalar processor 330.

[0223] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include a state configuration including a set of registers and memory, eg, as described below.

[0224] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a set of vector memories 312 , which may be configured, for example, to store data to be processed by the vector processor block 310 .

[0225] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a set of vector registers 314 that may be configured for use, for example, in data processing performed by the vector processor block 310 .

[0226] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may be associated with a set of memory maps.

[0227] In some exemplary aspects, a memory map may include a set of addresses accessible by a data processing unit that may load data from / to registers and memory and / or store data.

[0228] In some exemplary aspects, such as Figure 3 As shown, the vector processing block 310 may include a plurality of address generation units (AGUs) 320 , which may include addresses accessible to them, for example, in one or more memories in the memory 312 .

[0229] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a plurality of data processing units 316, for example, as described below.

[0230] In some exemplary aspects, the data processing unit 316 may be configured to process commands, for example, including a number of digits at a time. In one example, the command may include 8 digits. In another example, the command may include 4 digits, 16 digits, or any other count of digits.

[0231] In some exemplary aspects, two or more data processing units 316 can be used simultaneously. In one example, data processing unit 316 can process and execute multiple different commands, for example, 3 different commands, for example, including 8 numbers, in a single cycle.

[0232] In some exemplary aspects, the data processing units 316 may be asymmetric. For example, the first and second data processing units 316 may support different commands. For example, addition may be performed by the first data processing unit 316, and / or multiplication may be performed by the second data processing unit 316. For example, both operations may be performed by one or more additional data processing units 316.

[0233] In some demonstrative aspects, data processing unit 316 may be configured to support arithmetic operations for many combinations of input and output data types.

[0234] In some exemplary aspects, data processing unit 316 may be configured to support one or more operations, which may be less common. For example, processing unit 316 may support operations to work with a lookup table (LUT) of vector processor 300 and / or any other operations.

[0235] In some exemplary aspects, data processing unit 316 may be configured to support efficient computation of nonlinear functions, histograms, and / or random data access, which may, for example, facilitate implementation of algorithms like image scaling, Hough transform, and / or any other algorithm.

[0236] In some exemplary aspects, vector memory 312 may include a bank of memory having a size of 16K, for example, or any other size, that may be accessed in the same cycle.

[0237] In one example, the maximum memory access size may be 64 bits. According to this example, the peak throughput may be 256 bits, for example, 64×4=256. For example, a high memory bandwidth may be achieved to utilize the computational power of the data processing unit 316 .

[0238] In one example, two data processing units 316 may support 16 8-bit multiply and accumulate operations (MACs) per cycle. According to this example, two data processing units 316 may not be useful, for example, if the input numbers are not extracted at that speed, and / or there is no input of exactly 256 bits, for example, 16x8x2=256.

[0239] In some exemplary aspects, AGU 320 may be configured to perform memory access operations, such as loading and storing data from / to vector memory 314 .

[0240] In some exemplary aspects, AGU 320 may be configured to calculate addresses of input and output data items, for example, to handle I / O in situations where high bandwidth is not sufficient to utilize data processing unit 316 .

[0241] In some exemplary aspects, AGU 320 may be configured to calculate addresses of input and / or output data items, for example, based on configuration registers written by scalar processor 330, prior to entering a vector command block (eg, a loop).

[0242] For example, the AGU 320 may be configured to write an image base pointer, width, height, and / or stride to configuration registers, for example, to iterate over an image.

[0243] In some exemplary aspects, the AGU 320 may be configured to handle addressing (e.g., all addressing), for example, to provide a technical solution in which the data processing unit 316 may not have the burden of incrementing a pointer or counter in a loop and / or the burden of checking a line end condition, for example, to zero a counter in a loop.

[0244] In some exemplary aspects, such as Figure 3 As shown, the AGU 320 may include four AGUs, and accordingly, four memories 312 may be accessed in the same cycle. In other aspects, any other count of AGUs 32 may be implemented.

[0245] In some exemplary aspects, AGUs 320 may not be "bound" to memory banks 312. For example, an AGU 320 (e.g., each AGU 320) may access a memory bank 312 (e.g., each memory bank 312), e.g., as long as two or more AGUs 320 do not attempt to access the same memory bank 312 in the same cycle.

[0246] In some demonstrative aspects, vector registers 314 may be configured to support communications between data processing unit 316 and AGU 320 .

[0247] In one example, the total number of vector registers 314 may be 28, which may be divided into several subsets, for example, based on their functions. For example, a first subset of vector registers 314 may be used for input / output of, for example, all data processing units 316 and / or AGU 320; and / or a second subset of vector registers 314 may not be used for output of some operations (e.g., most operations) and may be used for one or more other operations, for example, to store loop-invariant inputs.

[0248] In some exemplary aspects, a data processing unit 316 (e.g., each data processing unit 316) may have one or more registers to host the output of the last performed operation, e.g., which may be fed as input to other data processing units 316. For example, these registers may "bypass" vector registers 314 and may operate faster than writing these outputs to the first set of vector registers 314.

[0249] In some exemplary aspects, the extractor and decoder 350 may be configured to support low-overhead vector loops, e.g., very low-overhead vector loops (also referred to as "zero-overhead vector loops"), e.g., where a termination (exit) condition of the vector loop may not need to be checked during execution of the vector loop.

[0250] For example, the AGU 320 may signal a termination (exit) condition, such as when the AGU 320 completes iterations over a configured memory region.

[0251] For example, the fetcher and decoder 350 may exit the loop when, for example, the AGU 320 signals a termination condition.

[0252] For example, the scalar processor 330 may be utilized to configure loop parameters, such as the first and last instructions and / or exit conditions.

[0253] In one example, vector loops may be utilized, for example, together with high memory bandwidth and / or cheap addressing, for example, to solve control and data flow problems, for example, to provide a technical solution to allow data processing unit 316 to process data with substantially no additional overhead.

[0254] In some exemplary aspects, scalar processor 330 may be configured to provide one or more functionalities that may be complementary to the functionality of vector processing block 310. For example, a large portion (e.g., most) of the work in a vector program may be performed by data processing unit 316. For example, scalar processor 330 may be utilized, for example, to "glue" together various vector code blocks of a vector program.

[0255] In some exemplary aspects, the scalar processor 330 may be implemented separately from the vector processing block 310. In other aspects, the scalar processor 330 may be configured to share one or more components and / or functionality with the vector processing block 310.

[0256] In some exemplary aspects, scalar processor 330 may be configured to perform operations that may not be suitable for execution on vector processing block 310 .

[0257] For example, the scalar processor 330 may be utilized to execute a 32-bit C program. For example, the scalar processor 330 may be configured to support 1, 2, and / or 4-byte data types of the C code and / or some or all arithmetic operators of the C code.

[0258] For example, the scalar processor 330 may be configured to provide a technical solution to perform operations that cannot be performed on the vector processing block 310 without, for example, using a full CPU.

[0259] In some exemplary aspects, scalar processor 330 may include a scalar data memory 332 , for example, having a size of 16K or any other size, which may be configured to store data, such as variables used by a scalar portion of a program.

[0260] For example, scalar processor 330 may store local and / or global variables declared by portable C code, which may be compiled by a compiler (e.g., compiler 200 ( Figure 2 )) is allocated to scalar data memory.

[0261] In some exemplary aspects, such as Figure 3 As shown, the scalar processor 330 may include or may be associated with a set of vector registers 334 that may be used for data processing by the scalar processor 330 .

[0262] In some exemplary aspects, the scalar processor 330 can be associated with a scalar memory map that can enable the scalar processor 330 to access substantially all states of the vector processor 300. For example, the scalar processor 330 can configure a vector unit and / or a DMA channel via the scalar memory map.

[0263] In some exemplary aspects, the scalar processor 330 may not be allowed to access one or more block control registers that may be used by an external processor to run and debug a vector program.

[0264] In some exemplary aspects, DMA 340 can be configured to communicate, for example, via main memory, with one or more other components of a chip implementing vector processor 300. For example, DMA 340 can be configured to transfer blocks of data, for example, large, contiguous blocks of data, for example, to support scalar processor 330 and / or vector processing blocks that can manipulate data stored in local memory. For example, a vector program may be able to use DMA 340 to read data from main chip memory.

[0265] In some exemplary aspects, DMA 340 may be configured to communicate with other elements of the chip, for example, via a plurality of DMA channels (e.g., 8 DMA channels or any other count of DMA channels). For example, a DMA channel (e.g., each DMA channel) may be able to transfer a rectangular patch from a local memory to a main chip memory, or vice versa. In other aspects, a DMA channel may transfer any other type of data block between a local memory and a main chip memory.

[0266] In some exemplary aspects, a rectangular tile may be defined by a base pointer, a width, a height, and a stride.

[0267] For example, at peak throughput, 8 bytes may be transferred per cycle, however, there may be an overhead for each tile and / or for each row in a tile.

[0268] In some exemplary aspects, DMA 340 can be configured to transfer data in parallel with computations, such as via multiple DMA channels, for example, as long as the executed commands do not access local memory involved in the transfer.

[0269] In one example, since all channels can access the same memory bus, using several channels to implement a transfer may not save I / O cycles, for example, compared to when a single channel is used. However, multiple DMA channels can be utilized to schedule several transfers and execute them in parallel with the calculation. For example, this may be advantageous compared to a single channel, which may not allow a second transfer to be scheduled before the first transfer is completed.

[0270] In some exemplary aspects, DMA 340 can be associated with a memory map that can support DMA channels accessing vector memory and / or scalar data. For example, access to vector memory can be performed in parallel with computation. For example, access to scalar data may not generally allow for parallelism, e.g., because scalar processor 330 may be involved in almost any reasonable program and may access its local variables while performing a transfer, which may result in memory contention with active DMA channels.

[0271] In some exemplary aspects, DMA 340 can be configured to provide a technical solution to support parallelization of I / O and computation. For example, a program performing computations may not have to wait for I / O, for example, when these computations can be run quickly by vector processing block 310.

[0272] In some exemplary aspects, an external processor (eg, a CPU) may be configured to initiate execution of a program on vector processor 300. For example, vector processor 300 may remain idle, for example, as long as program execution is not initiated.

[0273] In some exemplary aspects, the external processor may be configured to debug the program, for example, to execute a single step at a time, to stop when the program reaches a breakpoint, and / or to examine the contents of registers and memory storing program variables.

[0274] In some exemplary aspects, external memory mapping may be implemented to enable an external processor to control the vector processor 300 and / or a debugger, for example, by writing to control registers of the vector processor 300 .

[0275] In some exemplary aspects, the external memory map may be implemented by a superset of the scalar memory map. For example, the implementation may make all registers and memories defined by the architecture of the vector processor 300 accessible to a debugger backend running on an external processor.

[0276] In some exemplary aspects, the vector processor 300 may issue an interrupt signal, such as when the vector processor 300 terminates a program.

[0277] In some exemplary aspects, the interrupt signal may be used, for example, to implement a driver to maintain a queue of programs scheduled for execution by vector processor 300 and / or may be used to start a new program, for example, by an external processor, when a previously executed program completes.

[0278] Return to reference Figure 1In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on one or more loops, which may be based on, for example, source code 112, for example, as described below.

[0279] In some exemplary aspects, compiler 160 may be configured to generate object code 115 based on one or more loops, which may be configured, for example, according to a loop execution scheme, e.g., as described below.

[0280] In some exemplary aspects, the loop execution scheme may be configured to provide a technical solution to support one or more vector processing architectures, such as a VLIW architecture and / or any other architecture, for example, as described below.

[0281] In some exemplary aspects, the loop execution scheme may be configured to provide technical solutions to support improvement of one or more types of loop nests (eg, imperfect loop nests), for example, as described below.

[0282] In some exemplary aspects, a loop nest may include at least an outer loop and an inner loop, eg, as described below.

[0283] In some exemplary aspects, a loop nest may include an outer loop (e.g., an outermost loop), an inner loop (e.g., an innermost loop), and one or more nested loops (also referred to as “middle nested loops” or “middle loops”), which may be nested between the outer loop and the inner loop, e.g., as described below.

[0284] In some exemplary aspects, a loop nest may include multiple loops nested in multiple nesting levels, eg, as described below.

[0285] In one example, the plurality of loops may include a first loop (eg, an outer loop), eg, in a first nesting level, and a second loop (eg, an inner loop), eg, in a second nesting level.

[0286] In one example, the plurality of loops can include one or more intermediate loops, eg, in one or more intermediate nesting levels, eg, between a first nesting level and a second nesting level.

[0287] In one example, the plurality of loops may include three loops in three nesting levels. For example, the three loops may include a first loop at a first nesting level, e.g., an outermost loop; a second loop at a second nesting level, e.g., an intermediate loop; and a third loop at a third nesting level, e.g., an innermost loop. For example, the second loop may be nested in the first loop, and the third loop may be nested in the second loop.

[0288] In some exemplary aspects, it may be desirable to provide technical solutions to efficiently transform imperfect loop nests into perfect loop nests, e.g., to improve the performance of an executable program, e.g., when executed by a processor (e.g., a vector processor or any other target processor), e.g., as described below.

[0289] In some exemplary aspects, a perfect loop nest may be configured to include a loop nest where all computational operations of the loop nest reside in an innermost loop of the perfect loop nest.

[0290] For example, the outer loop of a perfect loop nest may not include any computational instructions.

[0291] For example, all computational instructions of a perfect loop nest may be in the innermost loop of the perfect loop nest.

[0292] In one example, one or more processor architectures may require and / or may benefit from the use of perfect loop nests in a program.

[0293] In another example, the perfect loop nest may be applicable to one or more (eg, more) loop optimizations.

[0294] In another example, one or more scheduling schemes (eg, modulo scheduling) that may be key optimizations for a VLIW target may not be able to optimize code across loop levels / basic blocks, such as when perfect loop nesting is not used.

[0295] In another example, one or more processor architectures may only support perfect loop nests. For example, these architectures may rely on being able to round loop nests into perfect loops.

[0296] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on one or more loops, which may be configured, for example, according to a loop execution scheme, which may be configured to provide a technical solution to improve the performance of a program executed by a target processor (e.g., a vector processor), for example, by transforming an imperfect loop nest into a perfect loop nest, for example, as described below.

[0297] In some exemplary aspects, the loop execution scheme may be configured to provide a technical solution to improve performance of a program executed by a target processor (e.g., a vector processor), such as by efficiently transforming imperfect loop nests into folded loops, for example, as described below.

[0298] In some exemplary aspects, a folded loop of a perfect loop nest may be configured to include a single basic block loop that includes all nested loops of the perfect loop, eg, as described below.

[0299] In some exemplary aspects, execution of a single basic block loop can be preconfigured, eg, along different dimensions, which can correspond to primitive loops in a primitive loop nest, for example.

[0300] For example, execution of the folded loop may be pre-configured, for example, by controlling hardware ("HW controlled"), for example, as described below.

[0301] In one example, one or more processor architectures may only support folded loops. For example, these processor architectures may rely on the ability to fold and / or perfect loop nests into single basic blocks and / or perfect loops.

[0302] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on one or more loops, which may be configured, for example, according to a loop execution scheme, which may be configured to provide a technical solution to support the transformation of imperfect loop nests into perfect loop nests, for example, to improve the performance of the executable program, for example, as described below.

[0303] In some exemplary aspects, the loop execution scheme may be configured to provide a technical solution to efficiently transform an imperfect loop nest into a folded loop, such as by transforming a perfect loop nest into a folded loop, for example, as described below.

[0304] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on a compilation scheme that may be configured, for example, to provide a technical solution to support computing one or more predicates that may be used to transform an imperfect loop nest into a perfect loop nest and / or transform an imperfect loop nest into a folded loop, e.g., as described below.

[0305] In some exemplary aspects, a predicate may be configured to indicate, identify, confirm, predict, and / or assert the start of a loop and / or the end of a loop, eg, the first iteration or the last iteration of a loop.

[0306] In some exemplary aspects, predicates can be configured to identify the start of a loop and / or the end of a loop, eg, even without processing and / or maintaining induction variables, eg, as described below.

[0307] In one example, one or more predicates can be utilized to indicate the start and / or end of execution of one or more inner loops nested within an original loop nest, for example, as described below.

[0308] In one example, efficiently computing predicates may be important, for example, where loop completion and / or loop folding relies on predication, for example, as described below.

[0309] In another example, efficiently computing predicates may be important, for example, to support processor architectures that may not be able to efficiently compute induction variables, for example, as described below.

[0310] In some exemplary aspects, the loop execution scheme may be configured to provide a technical solution to support processor architectures that may not support predicated instructions for non-memory access operations.

[0311] In some exemplary aspects, the loop execution scheme may be configured to provide technical solutions to support computing predicates, for example, to support transforming loop nests into perfect loop nests and / or folded loops, for example, as described below.

[0312] In some exemplary aspects, the loop execution scheme can be configured to provide a technical solution to enable more efficient execution of a program, for example, while avoiding the need to compute predicates based on induction variables, for example, as described below.

[0313] In some exemplary aspects, the loop execution scheme may be configured to provide a technical solution to support one or more architectures that do not have predicated operations in hardware and / or have loop nesting controlled by hardware, e.g., as described below.

[0314] In some exemplary aspects, compiler 160 may be configured to identify one or more loop nests based on source code, eg, as described below.

[0315] In some demonstrative aspects, compiler 160 may be configured to identify one or more of the loop nests in source code 112 , for example, where the loop nests are included in source code 112 .

[0316] In some exemplary aspects, compiler 160 may be configured to identify one or more of loop nests in code (eg, mid-level code that may be compiled from source code 112 ).

[0317] In some exemplary aspects, compiler 160 may be configured to transform one or more identified loop nests into one or more perfect loop nests, eg, as described below.

[0318] In some exemplary aspects, compiler 160 may be configured to compile source code into target code 115, for example, such that target code 115 may be based on one or more perfect loop nests, for example, as described below.

[0319] In some exemplary aspects, compiler 160 may be configured to transform one or more identified loop nests into one or more perfect loop nests, eg, according to a loop polishing scheme, eg, as described below.

[0320] In some exemplary aspects, compiler 160 may be configured to move one or more outer instructions from an outer loop level of a loop nest into an inner loop (e.g., an innermost loop) of the loop nest, for example, while using one or more predicates to protect execution of the outer instructions moved into the inner loop, for example, as described below.

[0321] In some exemplary aspects, the predicate can be configured to check and / or represent a state of an induction variable that counts the number of iterations of an inner loop, eg, as described below.

[0322] In some exemplary aspects, the predicate (also referred to as a “loop start predicate”) may be configured to identify when the induction variable may be equal to the start of the corresponding inner loop, e.g., for instructions moved from before the inner loop into the inner loop (also referred to as “sunk instructions”), e.g., as described below.

[0323] In some exemplary aspects, the predicate (also referred to as a “loop-ending predicate”) may be configured to identify when an induction variable may be equal to the end of a corresponding inner loop, e.g., for instructions moved from after the inner loop into the inner loop (also referred to as “promoted instructions”), e.g., as described below.

[0324] In some exemplary aspects, compiler 160 may be configured to move all instructions from an outer loop stage (nesting stage) of a loop nest into an innermost loop of the loop nest, e.g., to transform the loop nest into a perfect loop nest, e.g., as described below.

[0325] In some exemplary aspects, compiler 160 may be configured to identify one or more outer loop instructions of an outer loop of a loop nest that is external to an inner loop of the loop nest, eg, as described below.

[0326] In some exemplary aspects, compiler 160 may be configured to move outer loop instructions into an inner loop of a loop nest, for example, based on position-based criteria related to the position of the outer loop instructions relative to the inner loop, eg, as described below.

[0327] In some exemplary aspects, the position-based criteria may be used to identify whether an outer loop instruction precedes an inner loop (a "pre-header instruction") or follows an inner loop (a "latch instruction"), e.g., as described below.

[0328] In some exemplary aspects, compiler 160 may be configured to transform outer loop instructions into conditional instructions in an inner loop, which may be within an inner loop of a loop nest, eg, as described below.

[0329] For example, conditional instructions may be configured based on the location-based criteria, e.g., as described below.

[0330] In some exemplary aspects, the conditional instruction may be configured, for example, based on a predicate to indicate, identify, confirm, predict and / or assert an iteration count of an inner loop, for example, as described below.

[0331] In some exemplary aspects, predicates may be utilized to indicate, identify, confirm, predict, and / or assert whether the inner loop is at the first iteration of the inner loop or at the last iteration of the inner loop, eg, as described below.

[0332] In some exemplary aspects, conditional instructions can be configured as memory access operations that can be based on a predicate, such as an iteration count of an inner loop, eg, as described below.

[0333] In some exemplary aspects, the outer loop instructions may include a load operation, and the memory access operation may be configured to perform the load operation, for example, based on a predicate on an iteration count of the inner loop, eg, as described below.

[0334] In some exemplary aspects, the outer loop instructions may include a store operation, and the memory access operation may be configured to perform the store operation, for example, based on a predicate on an iteration count of the inner loop, eg, as described below.

[0335] In some exemplary aspects, compiler 160 may be configured to sink instructions, for example, by moving instructions to be performed before a first iteration of the inner loop into the inner loop, eg, as described below.

[0336] In some exemplary aspects, the pre-header instructions may be sunk, for example, by moving the pre-header instructions into an inner loop and transforming the pre-header instructions into pre-header conditional instructions, eg, as described below.

[0337] In some exemplary aspects, a pre-header conditional instruction may include a condition to configure an outcome of the pre-header conditional instruction based on a predicate, eg, on an inner loop, eg, as described below.

[0338] In some exemplary aspects, compiler 160 may generate target code 115 based on compiled code that may be configured to configure a particular outcome of a pre-header conditional instruction, e.g., when a predicate identifies execution of an inner loop before a first iteration of the inner loop, e.g., as described below.

[0339] In some exemplary aspects, compiler 160 may be configured to promote instructions, such as by moving instructions to be performed after a last iteration of the inner loop into the inner loop, eg, as described below.

[0340] In some exemplary aspects, a latch instruction may be promoted, for example, by moving the latch instruction into an inner loop and transforming the latch instruction into a latch conditional instruction (also referred to as a "promoted conditional instruction"), for example, as described below.

[0341] In some exemplary aspects, a latch conditional instruction may include a condition to configure a result of the latch conditional instruction based on a predicate, eg, on an inner loop, eg, as described below.

[0342] In some exemplary aspects, compiler 160 may generate target code 115 based on compiled code that may be configured to configure latching of specific results of conditional instructions, such as when a predicate identifies that execution of an inner loop is after a last iteration of the inner loop, e.g., as described below.

[0343] In some exemplary aspects, compiler 160 may be configured to repeatedly and / or iteratively perform lifting and / or sinking operations, e.g., to iterate over all instructions in a loop nest, e.g., until substantially all instructions are moved from an outer loop into an innermost loop, e.g., as described below.

[0344] For example, the loop execution scheme may be configured to provide a technical solution to support VMP iterators to work only on perfect loop nests. For example, the loop execution scheme may result in a situation where there are no instructions between two subsequent "for" statements that constitute a loop, for example, as described below.

[0345] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into target code 115 , for example, by transforming one or more loop nests into folded loops, eg, as described below.

[0346] In some exemplary aspects, one or more loop nests may be transformed into folded loops, e.g., to provide a technical solution to improve performance of execution of a program by a target processor (e.g., a vector processor and / or any other processor), e.g., as described below.

[0347] In some exemplary aspects, compiler 160 may be configured to transform one or more identified loop nests into folded loops, eg, according to a loop folding scheme, eg, as described below.

[0348] In some exemplary aspects, compiler 160 may be configured to apply a loop folding scheme, for example, based on the results of a loop polishing scheme, as described below.

[0349] In some exemplary aspects, compiler 160 may be configured to apply a loop folding scheme, eg, even without performing a loop perfecting scheme, eg, when a loop perfecting scheme is unnecessary, eg, when an input loop nest includes a perfect loop nest.

[0350] In one example, compiler 160 may be configured to apply a loop folding scheme, for example, to provide object code 115 that is configured for execution by one or more processor architectures that support only single basic block loops.

[0351] In some exemplary aspects, compiler 160 may be configured to collapse multiple separate loops of a loop nest into a single loop, which may be configured, for example, to execute substantially all iterations of the original loop nest, eg, as described below.

[0352] In some exemplary aspects, execution of the folded loop may be pre-configured, for example, to configure advancement along a dimension of the original loop.

[0353] In some exemplary aspects, the loop folding scheme may be configured to provide a technical solution to support execution of target code 115 by one or more processor architectures, including processor architectures in which calculation of induction variables may be computationally expensive.

[0354] In some exemplary aspects, compiler 160 may be configured to identify, for example, based on source code 112, a perfect loop nest including a plurality of nested loops, for example, as described below.

[0355] In some exemplary aspects, multiple nested loops may correspond to a corresponding multiple dimensions, eg, as described below.

[0356] In some exemplary aspects, dimensioning of the nested loops may be performed during multiple iterations of the nested loops, eg, as described below.

[0357] In some exemplary aspects, compiler 160 may be configured to configure folded loops based on loop nests, such as by folding multiple loop nests into a single loop based on multiple dimensions, for example, as described below.

[0358] In some exemplary aspects, compiler 160 may compile source code 112 of a program to be executed by a target processor (eg, processor 180 ), as described below.

[0359] For example, compiler 160 may identify, e.g., based on source code 112, a loop nest that includes an outer loop (y-loop) along a dimension based on a variable height, a nested loop (x-loop) along a dimension based on a variable width, and an inner loop (z-loop) along a dimension based on a variable area, e.g., as follows:

[0360]

[0361] Example (4)

[0362] For example, as shown in Example 4, a nested loop may include a pre-header instruction, such as out1[y*width+x]=inp1[y*width+x]+7, which may reside before the header of the inner loop (z-loop).

[0363] For example, as shown in Example 4, the pre-header instruction (out1[y*width+x]=inp1[y*width+x]+7) may be executed before the execution of the inner loop begins.

[0364] For example, as shown in Example 4, the pre-header instruction (out1[y*width+x]=inp1[y*width+x]+7) may include an external load operation (e.g., inp1[y*width+x]) and an external storage operation (e.g., "out1[y*width+x]="), which may be executed, for example, each time before the execution of the inner loop starts.

[0365] In some exemplary aspects, compiler 160 may be configured to sink pre-header instructions into inner loops, for example, to generate a perfect loop nest, for example, as described below.

[0366] For example, compiler 160 may be configured to move the pre-header instruction (out1[y*width+x]=inp1[y*width+x]+7) into the inner loop and transform the pre-header instruction (out1[y*width+x]=inp1[y*width+x]+7) into a conditional pre-header instruction (also called a "sunken conditional instruction"), for example, as described below.

[0367] For example, as shown in Example 4, the outer loop may include a latch instruction, such as out2[y]=a, which may follow the nested loop and the inner loop.

[0368] For example, as shown in Example 4, the latch instruction may include an external storage operation, such as out2[y]=a.

[0369] For example, as shown in Example 4, the latch instruction (out2[y]=a) may be executed, for example, each time after the last iteration of the inner loop and the nested loop.

[0370] In some exemplary aspects, compiler 160 may be configured to hoist a latch instruction (out2[y]=a) into an inner loop, for example, to generate a perfect loop nest, for example, as described below.

[0371] For example, compiler 160 may be configured to move a latch instruction (out2[y]=a) into an inner loop and transform the latch instruction (out2[y]=a) into a conditional latch (also called a "promoted conditional instruction") instruction, which may be based on the latch instruction, for example, as described below.

[0372] For example, as shown in Example 4, the inner loop may include an inner load instruction, such as a=inp2[y*width*area+x*area+z], which may be within the inner loop.

[0373] In some exemplary aspects, compiler 160 may be configured to transform a perfect loop nest into a folded loop along a dimension, which may be based on, for example, a value height, a value width, and a value area, e.g., as follows:

[0374]

[0375] Example (5)

[0376] In some exemplary aspects, as shown in Example 5, the folded loop may include a single block that includes instructions based on all instructions of Example 4.

[0377] In some exemplary aspects, as shown in Example 5, the load and store operations in the pre-header instruction out1[y*width+x]=inp1[y*width+x]+7 can be transformed into a conditional pre-header instruction "if (first_iteration_of_z_loop) out1[out1_ind]=result".

[0378] In some exemplary aspects, as shown in Example 5, the conditional pre-header instruction may include a condition based on a predicate that may indicate, identify, confirm, predict, and / or assert the start of an inner loop.

[0379] In some exemplary aspects, as shown in Example 5, the conditional pre-header instruction can be executed, for example, only when a predicate (first_iteration_of_z_loop) is true, for example, only when the inner loop starts executing.

[0380] In some exemplary aspects, as shown in Example 5, the latch instruction out2[y]=a may be transformed into a conditional latch instruction “if (last_iteration_of_x_and_z_loops) out2[out2_ind]=a”.

[0381] In some exemplary aspects, as shown in Example 5, the conditional latch instruction may include a condition based on a predicate that may indicate, identify, confirm, predict, and / or assert a last iteration of an inner loop and a last iteration of a nested loop.

[0382] In some exemplary aspects, as shown in Example 5, the conditional latch instruction may be executed, for example, only when a predicate (last_iteration_of_x_and_z_loops) is true, for example, only when the inner loop and the nested loop are after the last iteration.

[0383] In some exemplary aspects, as shown in Example 5, an intrinsic load instruction (e.g., a=inp2[y*width*area+x*area+z]) may be transformed into a load instruction, e.g., char a=inp2[inp2_ind]. For example, as shown in Example 5, an intrinsic load instruction (a=inp2[y*width*area+x*area+z]) may not be transformed into a conditional instruction.

[0384] In some exemplary aspects, as shown in Example 5, the index of the load and store instructions can be calculated, for example, as an AGU parameter.

[0385] In some exemplary aspects, as shown in Example 5, an index of a conditional latch instruction and / or a conditional pre-header instruction can be calculated, for example, as an AGU parameter, for example, as described below.

[0386] In some exemplary aspects, compiler 160 may be configured to provide technical solutions to support computing predicates (eg, conditional instructions) that may be used to transform outer loop instructions into folded loop instructions, eg, as described below.

[0387] In some exemplary aspects, for example, in some use cases, implementations, and / or scenarios, transforming outer loop instructions using conditional store / load instructions may be inefficient.

[0388] In one example, conditional store / load instructions may require additional conditional instructions.

[0389] In another example, conditional store / load instructions may require maintaining one or more induction variables in a loop.

[0390] In some exemplary aspects, compiler 160 may be configured to configure a folded loop (e.g., the folded loop of Example 5), for example, according to a predicate-based memory access mechanism, which may be configured to support configuration of memory access operations, for example, based on one or more predicates, for example, as described below.

[0391] In some exemplary aspects, the predicate-based memory access mechanism may be configured to provide a technical solution to support configuration of memory access operations (e.g., load instructions and / or store instructions) based on loop start predicates and / or loop end predicates, for example, as described below.

[0392] In some exemplary aspects, the predicate-based memory access mechanism may be configured to provide a technical solution to support memory access operations based on loop start predicates and / or loop end predicates, e.g., even without computing induction variables for one or more loops in a loop nest, e.g., as described below.

[0393] In some exemplary aspects, the predicate-based memory access mechanism may be configured to provide a technical solution to support memory access operations based on loop start predicates and / or loop end predicates, for example, at processor architectures that may not support computation of induction variables.

[0394] In some exemplary aspects, the predicate-based memory access mechanism may be configured to provide a technical solution to support memory access operations based on loop start predicates and / or loop end predicates, for example, on processor architectures where computing induction variables may be computationally expensive, such as when the loop is fully pre-configured.

[0395] In some exemplary aspects, compiler 160 may be configured to implement the conditional latch instruction (promoted conditional instruction) and / or conditional pre-header instruction (sunk conditional instruction) of Example 5, for example, by setting one or more AGU parameters corresponding to the conditional latch instruction and / or conditional pre-header instruction, for example, as described below.

[0396] In some exemplary aspects, compiler 160 may be configured to configure the AGU to perform memory access operations, which may be performed, for example, at the beginning of an inner loop or at the end of an inner loop based on an outer loop instruction, for example, as described below.

[0397] In some exemplary aspects, compiler 160 may be configured to identify loop nests based on source code 112 to be compiled into target code 150 to be executed by target processor 180 , eg, as described below.

[0398] In some exemplary aspects, compiler 160 may be configured to generate object code 115 that is configured for execution, for example, by a target vector processor (eg, vector processor 180), for example, as described below.

[0399] In some exemplary aspects, compiler 160 may be configured to generate target code 115 configured for execution by, for example, a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor (eg, processor 180).

[0400] In other aspects, compiler 160 may be configured to generate object code 115 configured for execution by, for example, any other suitable type of processor.

[0401] In some exemplary aspects, compiler 160 may be configured to generate object code 115 based on source code 112 including Open Computing Language (OpenCL) code, for example.

[0402] In other aspects, compiler 160 may be configured to generate object code 115 based on source code 112 , including any other suitable type of code, for example.

[0403] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into target code 115 , for example, according to a low-level virtual machine (LLVM)-based compilation scheme.

[0404] In other aspects, compiler 160 may be configured to compile source code 112 into target code 115 according to any other suitable compilation scheme.

[0405] In some demonstrative aspects, compiler 160 may be configured to identify loop nests in source code 112 , for example, where loop nests are included in source code 112 .

[0406] In some exemplary aspects, compiler 160 may be configured to identify loop nests in code (eg, mid-end code or any other code) that may be compiled from source code 112 .

[0407] In some exemplary aspects, a loop nest may include multiple loops, eg, including at least an outer loop and an inner loop within the outer loop, eg, as described below.

[0408] In some demonstrative aspects, the plurality of loops may include at least a first loop (eg, an outer loop) and a second loop (eg, an inner loop) nested within the first loop, eg, as described below.

[0409] In some exemplary aspects, the first loop may include at least one first loop instruction that may be outside of the second loop ("outer loop instruction"), eg, as described below.

[0410] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, eg, to configure an AGU of target processor 180 , eg, based on a first loop instruction, eg, as described below.

[0411] In some exemplary aspects, the AGU configuration code may be configured to configure a first dimension of the AGU, for example, based on a first cycle, eg, as described below.

[0412] In some exemplary aspects, the AGU configuration code may be configured to configure a second dimension of the AGU, for example, based on a second cycle, eg, as described below.

[0413] In some exemplary aspects, the AGU configuration code may be configured to configure a second dimension of the AGU, eg, to configure memory access operations to be performed at the beginning of a second cycle or at the end of the second cycle, eg, as described below.

[0414] In some exemplary aspects, the memory access operations may be based on, for example, first loop instructions, eg, as described below.

[0415] In some exemplary aspects, a memory access operation may comprise a load operation or a store operation, eg, as described below.

[0416] In some exemplary aspects, compiler 160 may be configured to generate object code 115 , for example, based on compiling source code 112 , for example, as described below.

[0417] In some exemplary aspects, target code 115 may be based on, for example, AGU configuration code, e.g., as described below.

[0418] In some exemplary aspects, the plurality of loops can include a third loop, which can, for example, be nested within the first loop, eg, as described below.

[0419] In some exemplary aspects, the second loop may be nested within the third loop, eg, as described below.

[0420] In some exemplary aspects, the first loop instructions may be outside of the third loop, eg, as described below.

[0421] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, eg, to configure a third dimension of the AGU, eg, based on a third cycle, eg, as described below.

[0422] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, eg, to configure the third dimension, eg, to configure memory access operations to be performed at the beginning of the second loop or at the end of the second loop, eg, as described below.

[0423] In some exemplary aspects, the third loop may include third loop instructions that may, for example, be outside of the second loop, eg, as described below.

[0424] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, eg, to configure another AGU of the target processor, eg, based on a third loop instruction, eg, as described below.

[0425] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, eg, to configure a first dimension of another AGU, eg, based on a third cycle, eg, as described below.

[0426] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, eg, to configure a second dimension of another AGU, eg, based on a second cycle, eg, as described below.

[0427] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, e.g., to configure a second dimension of another AGU, e.g., to configure another memory access operation to be performed at the beginning of a second loop or at the end of the second loop, e.g., as described below.

[0428] In some exemplary aspects, another memory access operation may be based on, for example, a third loop instruction, eg, as described below.

[0429] In some exemplary aspects, compiler 160 may be configured to transform a loop nest into a transformed loop that includes a memory access operation, eg, as described below.

[0430] In some exemplary aspects, target code 115 may be based on, for example, transformed loops, e.g., as described below.

[0431] In some exemplary aspects, a transformed cycle may include a memory access operation and another memory access operation, eg, as described below.

[0432] In some exemplary aspects, the transformed loop may include a perfectly flat loop in which, for example, all computation operations of the loop nest are implemented in the transformed loop, eg, as described below.

[0433] In some exemplary aspects, the transformed loop may include a fully folded loop, eg, including only a single basic block loop, eg, based on multiple loops, eg, as described below.

[0434] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, eg, to set a base address parameter of the AGU, eg, based on a memory pointer of a first loop instruction, eg, as described below.

[0435] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, eg, to set a maximum (Max) parameter of a second dimension of the AGU, eg, based on an entry size corresponding to a first loop instruction, eg, as described below.

[0436] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Max parameter for a second dimension of the AGU, for example, based on an entry size corresponding to a first loop instruction, and to set a Max parameter for a third dimension of the AGU, for example, where a second loop is nested within a third loop, for example, as described below.

[0437] In some exemplary aspects, the at least one first loop instruction may include a pre-header instruction to be performed before a first iteration of the inner loop, eg, as described below.

[0438] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to configure a second dimension of the AGU, for example, to configure memory access operations to be performed only at the beginning of a second loop, for example, when the first loop instructions include a pre-header instruction, for example, as described below.

[0439] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, e.g., to set a base address parameter of the AGU to a memory pointer of the pre-header instruction, e.g., when the first loop instruction includes a pre-header instruction, e.g., as described below.

[0440] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, e.g., to set a minimum (Min) parameter of a second dimension of the AGU to zero, e.g., when the first loop instruction includes a pre-header instruction, e.g., as described below.

[0441] In some exemplary aspects, the pre-header instruction may include a load operation, eg, as described below.

[0442] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to set a stride parameter of a second dimension of the AGU to zero, e.g., based on a determination that the pre-header instruction includes a load operation, e.g., as described below.

[0443] In some exemplary aspects, the pre-header instruction may include a store operation, eg, as described below.

[0444] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code, e.g., based on a determination that the pre-header instruction includes a storage operation, to set a stride parameter of a second dimension of the AGU, e.g., based on an entry size corresponding to the pre-header instruction, e.g., as described below.

[0445] In some exemplary aspects, the at least one first loop instruction may include a latch instruction to be performed after a last iteration of the second loop, eg, as described below.

[0446] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, e.g., to configure a second dimension of the AGU, e.g., to configure memory access operations to be performed only at the end of a second loop, e.g., when the first loop instructions include a latch instruction, e.g., as described below.

[0447] In some exemplary aspects, a latch instruction may include a load operation, eg, as described below.

[0448] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, e.g., to set a base address parameter of a second dimension of the AGU to, e.g., a memory pointer of a latch instruction, e.g., when the latch instruction comprises a load operation, e.g., as described below.

[0449] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Min parameter of a second dimension of the AGU to zero, for example, when a latch instruction includes a load operation, for example, as described below.

[0450] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, e.g., to set a Max parameter of a second dimension of the AGU to, e.g., an entry size corresponding to a latch instruction, e.g., when the latch instruction includes a load operation, e.g., as described below.

[0451] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a stride parameter of a second dimension of the AGU to zero, for example, when a latch instruction includes a load operation, for example, as described below.

[0452] In some exemplary aspects, a latch instruction may include a store operation, eg, as described below.

[0453] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a base address parameter of the AGU based on a first parameter value, a second parameter value, and a third parameter value, for example, when a latch instruction includes a storage operation, for example, as described below.

[0454] In some exemplary aspects, the first parameter value may comprise an entry size corresponding to a latch instruction, eg, as described below.

[0455] In some demonstrative aspects, the second parameter value may comprise a total iteration count over one or more loops in the first loop and including the second loop, eg, as described below.

[0456] In some demonstrative aspects, the third parameter value may include a dimension count of the AGU corresponding to one or more cycles, eg, as described below.

[0457] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a base address parameter denoted as Base of the AGU, for example, as follows:

[0458] Base=OrigBase+EntrySize*([∑TripCount(L)]-#InnerDims),

[0459] Where OrigBase represents the memory pointer of the latch instruction, where EntrySize represents the entry size corresponding to the latch instruction,

[0460] Where [∑TripCount(L)] represents the total iteration count over one or more loops in the first loop and including the second loop, and where #InnerDims represents the dimension count of the AGU corresponding to the one or more loops, for example, as described below.

[0461] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, e.g., to set a stride parameter of a second dimension of the AGU, e.g., based on an entry size corresponding to a latch instruction, e.g., when the latch instruction includes a store operation, e.g., as described below.

[0462] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Min parameter of a second dimension of the AGU, for example, based on an entry size corresponding to the latch instruction and an iteration count in a second loop, for example, when the latch instruction includes a storage operation, for example, as described below.

[0463] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Max parameter of a second dimension of the AGU, for example, based on a Min parameter of the second dimension of the AGU and an entry size corresponding to the latch instruction, for example, when a latch instruction includes a storage operation, for example, as described below.

[0464] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a stride parameter of a second dimension of the AGU, for example, based on an additive inverse of an entry size corresponding to the latch instruction, when the latch instruction includes a storage operation, for example, as described below.

[0465] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Min parameter of a second dimension of the AGU, for example, based on an entry size corresponding to the latch instruction and an iteration count in a second loop, for example, when the latch instruction includes a storage operation, for example, as described below.

[0466] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Min parameter of a second dimension of the AGU, for example, based on a product of an additive inverse of an entry size corresponding to the latch instruction and a subtraction result of subtracting one from an iteration count in a second loop, for example, when the latch instruction includes a store operation, for example, as described below.

[0467] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Max parameter of a second dimension of the AGU, for example, based on a Min parameter of the second dimension of the AGU and an entry size corresponding to the latch instruction, for example, when a latch instruction includes a storage operation, for example, as described below.

[0468] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Max parameter of a second dimension of the AGU, for example, based on a sum of a Min parameter of the second dimension of the AGU and an entry size corresponding to the latch instruction, for example, when the latch instruction includes a storage operation, for example, as described below.

[0469] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, based on a determination that the plurality of loops includes a third loop nested within a first loop, the second loop is nested within the third loop, and a latch instruction including a storage operation is outside the third loop, for example, as described below.

[0470] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, e.g., to configure a third dimension of the AGU based on a third cycle, e.g., when a latch instruction including a store operation is outside the third cycle, e.g., as described below.

[0471] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a stride parameter of a third dimension of the AGU, for example, based on an entry size corresponding to a latch instruction when a latch instruction including a storage operation is outside a third loop, for example, as described below.

[0472] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Min parameter of a third dimension of the AGU, for example, based on an entry size corresponding to the latch instruction and an iteration count in the third loop, when a latch instruction including a storage operation is outside the third loop, for example, as described below.

[0473] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to set a Max parameter of a third dimension of the AGU, for example, based on a Min parameter of a third dimension of the AGU and an entry size corresponding to the latch instruction, for example, when a latch instruction including a storage operation is outside a third loop, for example, as described below.

[0474] In some demonstrative aspects, compiler 160 may be configured to perform one or more operations, for example, according to a loop compilation scheme, which may be configured to compile instructions of a loop, for example, as described below.

[0475] In some demonstrative aspects, compiler 160 may be configured to recognize load and / or store operations and, for example, for each load and / or store operation, analyze one or more attributes of the load / store operation, eg, as described below.

[0476] In some exemplary aspects, compiler 160 may be configured to analyze, for a load / store operation (e.g., for each load and / or store operation), its innermost enclosing loop, its base address parameter, its stride parameter (e.g., stride), its offset, its bounds (e.g., Min / Max parameters), and / or its pass-through values.

[0477] In some exemplary aspects, compiler 160 may be configured to divide identified load and / or store operations into one or more groups, for example, according to their attributes (eg, according to their strides, bounds, and / or passthrough values).

[0478] In some exemplary aspects, compiler 160 can be configured to assign an AGU to a set of identified load / store operations, for example, to each set of identified load / store operations.

[0479] In one example, each set of identified load / store operations can be implemented using a single AGU.

[0480] In some exemplary aspects, compiler 160 may be configured to configure AGU configuration code for an AGU that implements load and / or store operations that may be located outside of an innermost loop.

[0481] For example, the AGU configuration code may set one or more specific parameters for load and / or store operations in an outer loop, including, for example, a specific base address parameter, a specific stride parameter, and / or specific Min / Max parameters, for example, as described below.

[0482] In some exemplary aspects, compiler 160 may be configured to sink or hoist external load operations that may be outside of inner loops.

[0483] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to configure the AGU based on an external load operation.

[0484] In some exemplary aspects, the AGU configuration code may configure load operations that may be performed at the beginning or end of an inner loop based on the outer load operations.

[0485] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to configure the AGU for performing a load operation in a first iteration of an inner loop or in a last operation of an inner loop.

[0486] In some exemplary aspects, compiler 160 may be configured to set a base address parameter of the AGU to, for example, a memory pointer for an external load operation, eg, similar to a typical setting of a base address parameter.

[0487] In some exemplary aspects, compiler 160 may be configured, for example, to set a stride parameter of the AGU for a dimension (denoted as L) corresponding to the inner loop to zero, for example, as follows:

[0488] For each inner loop L: Step L =0

[0489] For example, you can apply Step to a load operation. L =0 setting.

[0490] In some exemplary aspects, compiler 160 may be configured to set the minimum parameter of the AGU for the dimension corresponding to the inner loop L to zero, and set the maximum parameter of the AGU for the dimension corresponding to the inner loop L based on the entry size of the outer load operation, for example, as follows:

[0491] For each inner loop L:Min L =0,Max L =EntrySize

[0492] In some exemplary aspects, compiler 160 may be configured to, for example, sink the load operation inp1[y*width+x] of the pre-header instruction of Example 4 into, for example, the inner loop of Example 4.

[0493] In some exemplary aspects, compiler 160 may be configured to identify the innermost enclosing loop of the outer load operation of Example 4.

[0494] For example, compiler 160 may identify the loop over x (X loop) as the innermost enclosing loop for the load operation inp1[y*width+x], for example, while the loop over z (Z loop) may be inner than the loop for the load operation inp1[y*width+x].

[0495] In some exemplary aspects, compiler 160 may be configured to identify an entry size of memory pointer inp1. For example, the entry size of memory pointer inp1 may be 1, for example, because memory pointer inp1 may be configured to include a character (Char).

[0496] In some exemplary aspects, compiler 160 may be configured to transform an external load operation inp1[y*width+x] into, for example, a load instruction, such as char val=inp1[inp1_ind], for example, by configuring AGU configuration code for a dimension corresponding to an inner loop (e.g., a dimension corresponding to a Z loop) of an AGU that implements a memory pointer inp1, for example, as follows:

[0497] Base=inp1

[0498] Step(Z-Loop)=0

[0499] Min(Z-Loop)=0

[0500] Max(Z-Loop)=sizeof(char)=1

[0501] In some exemplary aspects, compiler 160 may be configured to sink external store operations, which may be outside of an inner loop, eg, as described below.

[0502] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to configure the AGU based on sinking of external memory operations.

[0503] In some exemplary aspects, the AGU configuration code may configure a storage operation that may be performed at the beginning of the inner loop based on the sinking of the outer storage operation. For example, the storage operation may be performed at the first iteration of the inner loop.

[0504] In some exemplary aspects, compiler 160 may be configured to set a base address parameter of the AGU, such as a memory pointer for external memory operations, eg, similar to a typical setting of a base address parameter.

[0505] In some exemplary aspects, compiler 160 may be configured to set a stride parameter of the AGU for a dimension (denoted as L) corresponding to the inner loop, for example, for each inner loop, based on an entry size of an outer storage operation, for example, as follows:

[0506] For each inner loop L: Step L =EntrySize

[0507] In some exemplary aspects, compiler 160 may be configured to set the minimum parameter of the AGU for the dimension corresponding to the inner loop L to, for example, zero, and set the maximum parameter of the AGU for the dimension corresponding to the inner loop L based on, for example, the entry size of the outer storage operation, e.g., as follows:

[0508] For each inner loop L:Min L =0,Max L =EntrySize

[0509] In some exemplary aspects, compiler 160 may be configured to sink the outer store operation “out1[y*width+x]=..” of the pre-fetcher instruction of Example 4 into, for example, the inner loop of Example 4.

[0510] In some exemplary aspects, compiler 160 may be configured to identify the innermost enclosing loop of the external storage operation of Example 4.

[0511] For example, compiler 160 may identify X loop as the innermost enclosing loop for the storage operation “out1 [y*width+x]=..”, while Z loop may be inner than the loop for the storage operation “out1 [y*width+x]=..”.

[0512] In some exemplary aspects, compiler 160 may be configured to identify an entry size of memory pointer out1.

[0513] For example, the entry size of the memory pointer out1 may be 1, for example, because the memory pointer out1 may be configured to include Char.

[0514] In some exemplary aspects, compiler 160 may be configured to transform an external storage operation “out1[y*width+x]=..” into, for example, a storage instruction, such as “if (first_iteration_of_z_loop) out1[out1_ind]=result”, for example, by configuring an AGU configuration code for a dimension corresponding to an inner loop (e.g., a dimension corresponding to a Z loop) of an AGU that implements a memory pointer out1, for example, as follows:

[0515] Base=out1

[0516] Step(Z-Loop)=sizeof(char)=1

[0517] Min(Z-Loop)=0

[0518] Max(Z-Loop)=sizeof(char)=1

[0519] In some exemplary aspects, compiler 160 may be configured to hoist external store operations, which may be outside of an inner loop, eg, as described below.

[0520] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to configure the AGU based on promotion of external memory operations.

[0521] In some exemplary aspects, the AGU configuration code may configure, for example, a store operation that may be performed at the end of an inner loop based on the promotion of the outer store operation. For example, the store operation may be performed at the last iteration of the inner loop.

[0522] In some exemplary aspects, compiler 160 may be configured to set a base address parameter of the AGU based on, for example, an entry size corresponding to an external memory operation, e.g., as follows:

[0523] Set Base = OrigBase + EntrySize * ([∑TripCount(L)] - #InnerDims), where OrigBase represents the base address of the external storage operation,

[0524] Where TripCount(L) represents the number of iterations for loop L,

[0525] Where #InnerDims represents the number of dimensions (total count) of the AGU corresponding to loops that are inner than loops with external storage operations. For example, the summation over loop L may be over all loops L that are inner than loops with external storage operations being lifted.

[0526] In some exemplary aspects, compiler 160 may be configured to set a stride parameter of the AGU, for example, for a dimension corresponding to an inner loop, based on an entry size of an outer storage operation, for example, as follows:

[0527] Set inner loops Step=-EntrySize

[0528] In some exemplary aspects, compiler 160 may be configured to, for example, for each inner loop, set, for example, a minimum parameter of an AGU for a dimension corresponding to inner loop L, and set a maximum parameter of an AGU for a dimension corresponding to inner loop L, for example, based on an entry size of an outer storage operation, for example, as follows:

[0529] For inner loop L setting: Min L =-EntrySize*(TripCount(L)–1)

[0530] Max L =Min L +EntrySize

[0531] In some exemplary aspects, compiler 160 may be configured to promote the storage operation “out2[y]=a” of the latch instruction of Example 4 into, for example, the inner loop of Example 4.

[0532] In some exemplary aspects, compiler 160 may be configured to identify the innermost enclosing loop of the external storage operation of Example 4.

[0533] For example, compiler 160 may identify a loop on y (Y loop) as the innermost enclosing loop for the storage operation “out2[y]=a”, while X loop and Z loop may be inner than the loop for the storage operation “out2[y]=a”.

[0534] In some exemplary aspects, compiler 160 may be configured to identify an entry size of memory pointer out2.

[0535] For example, the entry size of the memory pointer out2 may be 1, for example, because the memory pointer out2 may be configured to include Char.

[0536] In some exemplary aspects, compiler 160 may be configured to identify an iteration count (trip count) of an inner X loop and an iteration count (trip count) of an inner Z loop, for example, as follows:

[0537] TripCount(X-Loop)=width, TripCount(Z-Loop)=area

[0538] In some exemplary aspects, compiler 160 may be configured to transform an external storage operation “out2[y]=a” into, for example, a conditional storage instruction, for example, “if (last_iteration_of_x_and_z_loops) out2[out2_ind]=a”, for example, by configuring AGU configuration code for the dimensions of the inner loop (e.g., the dimensions corresponding to the X loop and the dimensions corresponding to the Z loop) of the AGU that implements the memory pointer out2, for example, as follows:

[0539] Base = out2 + (area + width – 2)

[0540] Step(Z-Loop)=-sizeof(int)=-1

[0541] Min(Z-Loop)=-1*(area–1)

[0542] Max(Z-Loop)=-1*(area–1)+1

[0543] Step(X-Loop)=-sizeof(int)=-1

[0544] Min(X-Loop)=-1*(width–1)

[0545] Max(X-Loop)=-1*(width–1)+1

[0546] Set the Y loop's step size, Min, Max, and trip count for all loops, for example, as usual:

[0547] Step(Y-Loop)=1,Min(Y-Loop)=0,Max(Y-Loop)=height*1,TripCount(Y-

[0548] Loop)=height, TripCount(X-Loop)=width, TripCount(Z-Loop)=area.

[0549] In some exemplary aspects, compiler 160 may configure the AGU configuration code, for example, to configure multiple AGUs based on the code of Example 4, for example, as follows:

[0550] agu1=allocate_agu("load");

[0551] set_base(agu1,inp1);

[0552] set_y_minmax(agu1,0,height*width);

[0553] set_y_count_stride(agu1,height,width);

[0554] set_x_minmax(agu1,0,width);

[0555] set_x_count_stride(agu1,width,1);

[0556] set_z_minmax(agu1,0,1);

[0557] set_z_count_stride(agu1,area,0);

[0558] agu2=allocate_agu(“load”);

[0559] set_base(agu2,inp2);

[0560] set_y_minmax(agu2,0,height*width*area);

[0561] set_y_count_stride(agu2,height,width*area);

[0562] set_x_minmax(agu2,0,width*area);

[0563] set_x_count_stride(agu2,width,area);

[0564] set_z_minmax(agu2,0,area);

[0565] set_z_count_stride(agu2,area,1);

[0566] agu3=allocate_agu(“store”);

[0567] set_base(agu3,out1);

[0568] set_y_minmax(agu3,0,height*width);

[0569] set_y_count_stride(agu3,height,width);

[0570] set_x_minmax(agu3,0,width);

[0571] set_x_count_stride(agu3,width,1);

[0572] set_z_minmax(agu3,0,1);

[0573] set_z_count_stride(agu3,area,1);

[0574] agu4=allocate_agu("store");

[0575] set_base(agu4,out2+(area+width–2));

[0576] set_y_minmax(agu4,0,height);

[0577] set_y_count_stride(agu4,height,1);

[0578] set_x_minmax(agu4,1-width,2-width);

[0579] set_x_count_stride(agu4,width,-1);

[0580] set_z_minmax(agu4,1-area,2-area);

[0581] set_z_count_stride(agu4,area,-1);

[0582] Example (6a)

[0583] In some exemplary aspects, compiler 160 may configure loop code to execute the loop of Example 4, for example, based on the AGU configuration code of Example 6a, for example, as follows:

[0584] cycle:

[0585] val = agu1.load() + 7;

[0586] agu3.store(val);

[0587] a=agu2.load();

[0588] agu4.store(a);

[0589] br(Loop);

[0590] Example (6b)

[0591] In some exemplary aspects, as shown in Example 6a, compiler 160 may, for example, specify the first AGU (eg, agu1) to perform a load operation “val=agu1.load()” based on the external load operation “inp1[y*width+x]” of Example 4.

[0592] In some exemplary aspects, as shown in Example 6a, compiler 160 may, for example, specify a second AGU (e.g., agu2) to perform a load operation “a=agu2.load()” based on the internal load operation “a=inp2[y*width*area+x*area+z]” of Example 4.

[0593] In some exemplary aspects, as shown in Example 6a, compiler 160 may, for example, specify a third AGU (eg, agu3) to perform a storage operation “agu3.store(val);” based on the external storage operation “out1[y*width+x]=…;” of Example 4.

[0594] In some exemplary aspects, as shown in Example 6a, compiler 160 may, for example, specify the fourth AGU (eg, agu4) to perform a storage operation "agu4.store(a)" based on the external storage operation "out2[y]=a" of Example 4.

[0595] In some exemplary aspects, as shown in Example 6a, compiler 160 may configure the AGU configuration code to configure agu1.

[0596] In some exemplary aspects, as shown in Example 6a, the AGU configuration code to configure agu1 may be configured to set a base address parameter of the first AGU to the memory pointer agu1.

[0597] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu1 can be configured to set the Min parameter for dimension z of the first AGU to zero and the Max parameter for dimension z of the first AGU to 1.

[0598] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu1 may be configured to set the iteration count for dimension z of the first AGU to the value area and set the stride (step size) for dimension z of the first AGU to zero.

[0599] For example, these settings for dimension z of the first AGU may configure a load operation val=agu1.load() to be performed, for example, at the beginning of inner loop z.

[0600] In some exemplary aspects, as shown in Example 6a, compiler 160 may configure the AGU configuration code to configure agu3.

[0601] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu3 can be configured to set the base address parameter of the third AGU to the memory pointer out1.

[0602] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu3 can be configured to set the Min parameter for dimension z of the third AGU to zero and set the Max parameter for dimension z of the third AGU to 1.

[0603] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu1 may be configured to set the iteration count for dimension z of the third AGU to the value area and set the stride (step size) for dimension z of the third AGU to 1.

[0604] For example, these settings for dimension z of the first AGU may configure a load operation agu3.store(val) to be performed, for example, at the beginning of inner loop z.

[0605] In some exemplary aspects, as shown in Example 6a, compiler 160 may configure the AGU configuration code to configure agu4.

[0606] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu4 can be configured to set the base address parameters of the fourth AGU based on, for example, the memory pointer out2 and the iteration count of the inner loop (e.g., the value area and the value width), for example, as follows:

[0607] Set Base(agu4)=OrigBase+EntrySize*([∑TripCount(L)]-

[0608] #InnerDims)=out2+1*([TripCount(Z-loop)+TripCount(X-loop)]-

[0609] 2)=out2+area+width-2=out2+1*([area+width]-2)=out2+area+width-2.

[0610] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu4 can be configured to set the Min parameter for the dimension x of the fourth AGU, for example, as follows:

[0611] Min x =-EntrySize*(TripCount(X-loop)–1)=(-1)*(width-1)=1-width

[0612] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu4 can be configured to set the Max parameter for the dimension x of the fourth AGU, for example, as follows:

[0613] Max x =Min x +EntrySize=1-width+1=2-width

[0614] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu4 can be configured to set the step size (step size) parameter for the dimension x of the fourth AGU, for example, as follows:

[0615] Step x =-EntrySize=-1

[0616] In some exemplary aspects, as shown in Example 6a, the AGU configuration code to configure agu4 can be configured to set a count parameter for dimension x of the fourth AGU to width, for example, based on an iteration count of the X loop.

[0617] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu4 can be configured to set the Min parameter for the dimension z of the fourth AGU, for example, as follows:

[0618] Min z =-EntrySize*(TripCount(Z-loop)–1)=(-1)*(area-1)=1-area

[0619] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu4 can be configured to set the Max parameter for the dimension z of the fourth AGU, for example, as follows:

[0620] Max z =Min z +EntrySize=1-area+1=2-area

[0621] In some exemplary aspects, as shown in Example 6a, the AGU configuration code used to configure agu4 can be configured to set a stride (step length) parameter for dimension z of the fourth AGU, for example, as follows:

[0622] Step z =-EntrySize=-1

[0623] In some exemplary aspects, as shown in Example 6a, the AGU configuration code to configure agu4 can be configured to set a g-count parameter for dimension z of the fourth AGU to area, for example, based on an iteration count of the Z loop.

[0624] In one example, compiler 160 may process source code 112, which may be based on Example 4, for example, with a setting of width=3 and a setting of area=5, for example, as follows:

[0625]

[0626] Example (7)

[0627] In some exemplary aspects, as shown in Example 7, the outermost loop (Y loop) may include an outer store instruction, such as out[y]=c, which may follow the nested loop (X loop) and the inner loop (Z loop).

[0628] In some exemplary aspects, the external store instruction may be executed after 3 iterations of the X loop, e.g., where each iteration of the X loop includes 5 iterations of the Z loop. For example, the external store instruction may be executed after a total of 15 iterations (e.g., 3*5=15).

[0629] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, for example, to configure the AGU based on the external store instruction of Example 7.

[0630] In some exemplary aspects, the AGU configuration code may, for example, set parameters for the x-dimension and z-dimension of the AGU based on an external storage instruction out[y]=c, for example, including a base address parameter, a total iteration count, a step size parameter, a Min parameter, and a Max parameter, for example, as described above.

[0631] refer to Figure 4 , which schematically illustrates an implementation scheme 400 for performing a latch-store operation in a loop nest according to some exemplary aspects.

[0632] In some exemplary aspects, implementation scheme 400 may demonstrate settings of an AGU to configure execution of an external storage operation out[y]=c of Example 7.

[0633] In some exemplary aspects, such as Figure 4 As shown, execution scheme 400 may include three execution steps, for example, corresponding to three iterations of the X loop of Example 7.

[0634] In some exemplary aspects, such as Figure 4 As shown, the execution scheme 400 may include a first execution step 410 corresponding to a first iteration of the X loop (eg, x=0).

[0635] In some exemplary aspects, such as Figure 4 As shown, the execution scheme 400 may include a second execution step 420 corresponding to a second iteration of the X loop (eg, x=1).

[0636] In some exemplary aspects, such as Figure 4 As shown, the execution scheme 400 may include a third execution step 430 corresponding to a third iteration of the X loop (eg, x=2).

[0637] In some exemplary aspects, such as Figure 4 As shown, the Z loop may be performed, for example, five times in each execution step of the execution scheme 400 .

[0638] In some exemplary aspects, such as Figure 4 As shown, the AGU configuration code may be configured based on Example 7, for example, to set the Min parameter to zero and the Max parameter to one for each of the x-dimension and the z-dimension of the AGU.

[0639] In some exemplary aspects, such as Figure 4 As shown, the AGU configuration code may be configured based on instance 7, for example, to set the base address parameter of the AGU to memory pointer 6.

[0640] In some exemplary aspects, such as Figure 4 As shown, during the first execution step 410, the first iteration of the Z loop may start at memory pointer 7, and the last iteration of the Z loop may be at memory pointer 2. Accordingly, a storage operation defined by a Min parameter of zero and a Max parameter of one may not be performed.

[0641] In some exemplary aspects, such as Figure 4 As shown, during the second execution step 420, the first iteration of the Z loop may start at memory pointer 6, and the last iteration of the Z loop may be at memory pointer 1. Accordingly, storage operations defined by a Min parameter of zero and a Max parameter of one may not be performed.

[0642] In some exemplary aspects, such as Figure 4As shown, during the third execution step 430, the first iteration of the Z loop may start at memory pointer 5, and the last iteration of the Z loop may be at memory pointer 0. Accordingly, a storage operation, which may be defined by a Min parameter of zero and a Max parameter of one, may be performed, for example, only in the last iteration of the Z loop in the last iteration of the X loop.

[0643] refer to Figure 5 , which schematically illustrates an execution scheme 500 for performing a pre-header load or store operation in a loop nest according to some exemplary aspects.

[0644] In some exemplary aspects, execution scheme 500 may demonstrate execution of a pre-header load or store instruction, for example, according to a loop execution scheme.

[0645] In some exemplary aspects, the pre-header load or store instruction may be within an outer loop, which may be outside of an inner loop.

[0646] In some exemplary aspects, such as Figure 5 As shown, the AGU configuration code corresponding to the AGU used to perform a pre-header load or store operation may, for example, set the Min parameter to zero and the Max parameter to one for the dimension of the AGU corresponding to the pre-header load or store instruction.

[0647] In some exemplary aspects, such as Figure 5 As shown, the AGU configuration code corresponding to the AGU used to perform a pre-header load or store operation may set the base address parameter of the AGU to, for example, a memory pointer of an external load or store pre-header instruction.

[0648] In some exemplary aspects, such as Figure 5 As shown, the AGU configuration code may be configured to provide a technical solution to ensure that a pre-header load or store operation will be performed only once at the beginning of an inner loop.

[0649] For example, the setting of the Min and Max parameters that may limit execution of a load or store operation to the first iteration of an inner loop may ensure that a pre-header load or store operation will be executed only once at the beginning of the inner loop, eg, as described above.

[0650] refer to Figure 6 , which schematically illustrates a method of compiling code for a processor. For example, Figure 6 One or more operations of the method may be performed by: a system, for example, system 100 ( Figure 1 ); devices, for example, device 102 ( Figure 1 ); a server, for example, server 170 ( Figure 1 ); and / or a compiler, for example, compiler 160 ( Figure 1 ) and / or compiler 200( Figure 2 ).

[0651] In some exemplary aspects, as indicated at block 602, the method may include identifying a loop nest based on source code to be compiled into target code to be executed by a target processor. For example, the loop nest may include a plurality of loops including at least a first loop and a second loop nested within the first loop. For example, the first loop may include at least one first loop instruction outside of the second loop. For example, the compiler 160( Figure 1 ) may be configured, for example, based on source code 112 ( Figure 1 ) to identify loop nests, for example, as described above.

[0652] In some exemplary aspects, as indicated at block 604, the method may include generating AGU configuration code to configure an AGU of a target processor based on the first loop instruction. For example, the AGU configuration code may configure a first dimension of the AGU based on the first loop, and configure a second dimension of the AGU based on the second loop. For example, the AGU configuration code may configure the second dimension of the AGU to configure a memory access operation to be performed at the beginning of the second loop or at the end of the second loop. For example, the memory access operation may be based on the first loop instruction. For example, the compiler 160( Figure 1 ) may be configured to generate AGU configuration code to configure the target processor 180 based on the first cycle instruction ( Figure 1 )'s AGU, for example, as described above.

[0653] In some exemplary aspects, as indicated at block 606, the method may include generating target code based on compiling the source code. For example, the target code may be based on the AGU configuration code. For example, the compiler 160 ( Figure 1 ) may be configured, for example, based on the source code 112 ( Figure 1 ) to generate the target code 115( Figure 1 ), the target code is based on AGU configuration code, for example, as described above.

[0654] refer to Figure 7 , which schematically illustrates an article of manufacture 700 according to some exemplary aspects. Article 700 may include one or more tangible computer-readable ("machine-readable") non-transitory storage media 702, which may include computer-executable instructions implemented, for example, by logic 704, which are operable to enable at least one computer processor to perform operations on a device 102 ( Figure 1 )、Server 170( Figure 1 ) and / or compiler 160( Figure 1) to implement one or more operations to enable device 102 ( Figure 1 )、Server 170( Figure 1 ) and / or compiler 160( Figure 1 ) performs, triggers and / or implements one or more operations and / or functionalities, and / or performs, triggers and / or implements reference Figures 1 to 6 One or more operations and / or functionalities described, and / or one or more operations described herein. The phrases "non-transitory machine-readable medium" and "computer-readable non-transitory storage medium" may be directed to include all computer-readable media, with the sole exception of transitory propagating signals.

[0655] In some exemplary aspects, the product 700 and / or the machine-readable storage medium 702 may include one or more types of computer-readable storage media capable of storing data, including volatile memory, non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. For example, the machine-readable storage medium 702 may include RAM, DRAM, double data rate DRAM (DDR-DRAM), SDRAM, static RAM (SRAM), ROM, programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory (e.g., NOR or NAND flash memory), content addressable memory (CAM), polymer memory, phase change memory, ferroelectric memory, silicon oxide nitride oxide silicon (SONOS) memory, disk, hard drive, etc. The computer-readable storage medium may include any suitable medium involved in downloading or transferring a computer program from a remote computer to a requesting computer via a communication link (e.g., a modem, radio, or network connection), the computer program being carried by a data signal embedded in a carrier wave or other propagation medium.

[0656] In some exemplary aspects, logic 704 may include instructions, data, and / or code that, if executed by a machine, may cause the machine to perform methods, processes, and / or operations as described herein. The machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, etc., and may be implemented using any suitable combination of hardware, software, firmware, etc.

[0657] In some exemplary aspects, logic 704 may include or may be implemented as software, a software module, an application, a program, a subroutine, an instruction, an instruction set, a computing code, a word, a value, a symbol, etc. Instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Instructions may be implemented according to a predefined computer language, manner, or syntax for instructing a processor to perform a specific function. Instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language, machine code, etc.

[0658] Examples

[0659] The following examples relate to further aspects.

[0660] Example 1 includes a product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions operable to, when executed by at least one processor, enable the at least one processor to cause a compiler to: identify a loop nest based on source code to be compiled into target code to be executed by a target processor, the loop nest comprising a plurality of loops, the plurality of loops comprising at least a first loop and a second loop nested within the first loop, wherein the first loop comprises at least one first loop instruction outside of the second loop; generate an address generation unit (AGU) configuration code to configure an AGU of the target processor based on the first loop instruction, wherein the AGU configuration code is used to configure a first dimension of the AGU based on the first loop, and to configure a second dimension of the AGU based on the second loop, wherein the AGU configuration code is used to configure the second dimension of the AGU to configure a memory access operation to be performed at the beginning of the second loop or at the end of the second loop, wherein the memory access operation is based on the first loop instruction; and generate target code based on the compilation of the source code, wherein the target code is based on the AGU configuration code.

[0661] Example 2 includes the subject matter of Example 1, and optionally wherein the plurality of loops includes a third loop nested within the first loop, the second loop nested within the third loop, the first loop instructions being outside the third loop, wherein the AGU configuration code is used to configure a third dimension of the AGU based on the third loop, wherein the AGU configuration code is used to configure the third dimension to configure a memory access operation to be performed at the beginning of the second loop or at the end of the second loop.

[0662] Example 3 includes the subject matter of Example 2, and optionally wherein the third loop includes a third loop instruction outside of the second loop, wherein the AGU configuration code is used to configure another AGU of the target processor based on the third loop instruction, wherein the AGU configuration code is used to configure a first dimension of the another AGU based on the third loop, and configure a second dimension of the another AGU based on the second loop, wherein the AGU configuration code is used to configure the second dimension of the another AGU to configure another memory access operation to be performed at the beginning of the second loop or at the end of the second loop, wherein the another memory access operation is based on the third loop instruction.

[0663] Example 4 includes the subject matter of example 3, and optionally wherein the instructions when executed cause a compiler to transform a loop nest into a transformed loop comprising a memory access operation and another memory access operation, wherein the target code is based on the transformed loop.

[0664] Example 5 includes the subject matter of any of Examples 2 to 4, and optionally wherein the AGU configuration code is to set a maximum (Max) parameter of a second dimension of the AGU and a Max parameter of a third dimension of the AGU based on an entry size corresponding to the first loop instruction.

[0665] Example 6 includes the subject matter of any of Examples 1 to 5, and optionally wherein the AGU configuration code is used to set a base address parameter of the AGU based on a memory pointer of a first loop instruction, and to set a maximum (Max) parameter of a second dimension of the AGU based on an entry size corresponding to the first loop instruction.

[0666] Example 7 includes the subject matter of any of Examples 1 to 6, and optionally wherein at least one first loop instruction includes a pre-header instruction to be performed before a first iteration of a second loop, wherein the AGU configuration code is used to configure a second dimension of the AGU to configure memory access operations to be performed only at the beginning of the second loop.

[0667] Example 8 includes the subject matter of Example 7, and optionally wherein the AGU configuration code is to set a minimum (Min) parameter of a second dimension of the AGU to zero.

[0668] Example 9 includes the subject matter of example 7 or 8, and optionally wherein the instructions when executed cause the compiler to configure the AGU configuration code to set a stride parameter of a second dimension of the AGU to zero based on a determination that the pre-header instruction comprises a load operation.

[0669] Example 10 includes subject matter according to any of Examples 7 to 9, and optionally wherein the instructions, when executed, cause the compiler to configure the AGU configuration code to set a stride parameter of a second dimension of the AGU based on an entry size corresponding to the pre-header instruction based on a determination that the pre-header instruction includes a storage operation.

[0670] Example 11 includes the subject matter of any of Examples 7 to 10, and optionally wherein the AGU configuration code is to set a base address parameter of the AGU to a memory pointer of the pre-header instruction.

[0671] Example 12 includes the subject matter of any of Examples 1 to 11, and optionally wherein at least one first loop instruction includes a latch instruction to be performed after a last iteration of the second loop, wherein the AGU configuration code is used to configure a second dimension of the AGU to configure a memory access operation to be performed only at the end of the second loop.

[0672] Example 13 includes the subject matter of Example 12, and optionally wherein the latch instruction comprises a load operation.

[0673] Example 14 includes the subject matter of Example 13, and optionally wherein the AGU configuration code is used to: set a base address parameter of the AGU to a memory pointer of a latch instruction, set a minimum (Min) parameter of a second dimension of the AGU to zero, set a maximum (Max) parameter of the second dimension of the AGU to an entry size corresponding to the latch instruction, and set a stride parameter of the second dimension of the AGU to zero.

[0674] Example 15 includes the subject matter of Example 12, and optionally wherein the latch instruction comprises a store operation.

[0675] Example 16 includes the subject matter of Example 15, and optionally wherein the AGU configuration code is used to set a base address parameter of the AGU based on a first parameter value, a second parameter value, and a third parameter value, wherein the first parameter value comprises an entry size corresponding to a latch instruction, the second parameter value comprises a total iteration count over one or more loops in the first loop and including the second loop, and the third parameter value comprises a dimension count of the AGU corresponding to the one or more loops.

[0676] Example 17 includes the subject matter of example 16, and optionally wherein the AGU configuration code is to set a base address parameter denoted as Base of the AGU as follows:

[0677] Base=OrigBase+EntrySize*([∑TripCount(L)]-#InnerDims),

[0678] Where OrigBase represents the memory pointer of the latch instruction, EntrySize represents the entry size, [ΣTripCount(L)] represents the total iteration count over one or more loops, and #InnerDims represents the dimension count of the AGU corresponding to the one or more loops.

[0679] Example 18 includes the subject matter of any of Examples 15 to 17, and optionally wherein the AGU configuration code is used to: set a step parameter of a second dimension of the AGU based on an entry size corresponding to the latch instruction; set a minimum (Min) parameter of the second dimension of the AGU based on the entry size and an iteration count in a second loop; and set a maximum (Max) parameter of the second dimension of the AGU based on the Min parameter of the second dimension of the AGU and the entry size.

[0680] Example 19 includes the subject matter of Example 18, and optionally wherein the AGU configuration code is to set a stride parameter of a second dimension of the AGU based on an additive inverse of an entry size.

[0681] Example 20 includes the subject matter of Example 18 or 19, and optionally wherein the AGU configuration code is to set a Min parameter of a second dimension of the AGU based on a product of an additive inverse of the entry size and a subtraction result of subtracting one from an iteration count in a second loop.

[0682] Example 21 includes the subject matter of any of Examples 18 to 20, and optionally wherein the AGU configuration code is to set a Max parameter of a second dimension of the AGU based on a sum of a Min parameter of the second dimension of the AGU and an entry size.

[0683] Example 22 includes the subject matter of any one of Examples 16 to 21, and optionally wherein the plurality of loops includes a third loop nested within the first loop, the second loop nested within the third loop, a latch instruction outside the third loop, and wherein the AGU configuration code is used to configure a third dimension of the AGU based on the third loop.

[0684] Example 23 includes the subject matter of Example 22, and optionally wherein the AGU configuration code is used to: set a step parameter of a third dimension of the AGU based on an entry size, set a minimum (Min) parameter of the third dimension of the AGU based on the entry size and an iteration count in a third loop, and set a maximum (Max) parameter of the third dimension of the AGU based on the Min parameter of the third dimension of the AGU and the entry size.

[0685] Example 24 includes the subject matter of any of Examples 1 to 23, and optionally wherein the instructions when executed cause a compiler to transform a loop nest into a transformed loop including a memory access operation, wherein the target code is based on the transformed loop.

[0686] Example 25 includes the subject matter of example 24, and optionally wherein the transformed loop comprises a perfectly flat loop in which all computation operations of the loop nest are implemented in the transformed loop.

[0687] Example 26 includes the subject matter of example 24 or 25, and optionally wherein the transformed loop comprises a fully folded loop, the fully folded loop comprising only a single basic block loop based on the multiple loops.

[0688] Example 27 includes the subject matter of any of Examples 1 to 26, and optionally wherein the memory access operation comprises a load operation or a store operation.

[0689] Example 28 includes the subject matter of any one of Examples 1-27, and optionally wherein the source code includes Open Computing Language (OpenCL) code.

[0690] Example 29 includes the subject matter of any one of Examples 1 to 28, and optionally wherein the computer-executable instructions, when executed, cause a compiler to compile source code into target code according to a low-level virtual machine (LLVM-based) compilation scheme.

[0691] Example 30 includes the subject matter of any of Examples 1 to 29, and optionally wherein the target code is configured for execution by a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor.

[0692] Example 31 includes the subject matter of any of Examples 1 to 30, and optionally wherein the object code is configured for execution by a target vector processor.

[0693] Example 32 includes a compiler configured to perform any of the operations described in any of Examples 1 to 31.

[0694] Example 33 includes a computing device configured to perform any of the operations described in any of Examples 1 to 31.

[0695] Example 34 includes a computing system comprising: at least one memory for storing instructions; and at least one processor for retrieving instructions from the memory and executing the instructions to cause the computing system to perform any of the operations described in any of Examples 1 to 31.

[0696] Example 35 includes a computing system comprising: a compiler configured to generate a target code according to any of the operations described in any of Examples 1 to 31; and a processor configured to execute the target code.

[0697] Example 36 includes a device comprising means for performing any of the operations described in any of Examples 1 to 31.

[0698] Example 37 includes an apparatus comprising: a memory interface; and a processing circuit system configured to: perform any of the operations described in any of Examples 1 to 31.

[0699] Example 38 includes a method comprising any of the operations described in any of Examples 1 to 31.

[0700] The functions, operations, components, and / or features described herein with reference to one or more aspects may be combined with or utilized in combination with one or more other functions, operations, components, and / or features described herein with reference to one or more other aspects, or vice versa.

[0701] While certain features have been illustrated and described herein, numerous modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the disclosure.

Claims

1. A product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions operable to, when executed by at least one processor, enable the at least one processor to cause a compiler to: identifying a loop nest based on source code to be compiled into target code to be executed by a target processor, the loop nest comprising a plurality of loops, the plurality of loops comprising at least a first loop and a second loop nested within the first loop, wherein the first loop comprises at least one first loop instruction outside of the second loop; generating an address generation unit (AGU) configuration code to configure an AGU of the target processor based on the first loop instruction, wherein the AGU configuration code is used to configure a first dimension of the AGU based on the first loop and to configure a second dimension of the AGU based on the second loop, wherein the AGU configuration code is used to configure the second dimension of the AGU to configure a memory access operation to be performed at the beginning of the second loop or at the end of the second loop, wherein the memory access operation is based on the first loop instruction; as well as The object code is generated based on compiling the source code, wherein the object code is based on the AGU configuration code.

2. The product of claim 1 , wherein the plurality of loops comprises a third loop nested in the first loop, the second loop nested in the third loop, the first loop instruction being outside the third loop, wherein the AGU configuration code is used to configure a third dimension of the AGU based on the third loop, wherein the AGU configuration code is used to configure the third dimension to configure the memory access operation to be performed at the beginning of the second loop or at the end of the second loop.

3. The product of claim 2, wherein the third loop comprises a third loop instruction outside of the second loop, wherein the AGU configuration code is used to configure another AGU of the target processor based on the third loop instruction, wherein the AGU configuration code is used to configure a first dimension of the another AGU based on the third loop and to configure a second dimension of the another AGU based on the second loop, wherein the AGU configuration code is used to configure the second dimension of the another AGU to configure another memory access operation to be performed at the beginning of the second loop or at the end of the second loop, wherein the another memory access operation is based on the third loop instruction.

4. The product of claim 3, wherein the instructions, when executed, cause the compiler to transform the loop nest into a transformed loop comprising the memory access operation and the another memory access operation, wherein the target code is based on the transformed loop.

5. The product of claim 2, wherein the AGU configuration code is to set a maximum (Max) parameter of the second dimension of the AGU and a Max parameter of the third dimension of the AGU based on an entry size corresponding to the first loop instruction.

6. The product of claim 1 , wherein the AGU configuration code is to set a base address parameter of the AGU based on a memory pointer of the first loop instruction, and to set a maximum (Max) parameter of the second dimension of the AGU based on an entry size corresponding to the first loop instruction.

7. The article of manufacture of claim 1 , wherein the at least one first loop instruction comprises a pre-header instruction to be performed prior to a first iteration of the second loop, wherein the AGU configuration code is to configure the second dimension of the AGU to configure the memory access operation to be performed only at the beginning of the second loop.

8. The product of claim 7, wherein the AGU configuration code is to set a minimum (Min) parameter of the second dimension of the AGU to zero.

9. The article of manufacture of claim 7, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code to set a stride parameter of the second dimension of the AGU to zero based on a determination that the pre-header instruction comprises a load operation.

10. The article of manufacture of claim 7, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code based on a determination that the pre-header instruction comprises a store operation to set a stride parameter of the second dimension of the AGU based on an entry size corresponding to the pre-header instruction.

11. The article of claim 7, wherein the AGU configuration code is to set a base address parameter of the AGU to a memory pointer of the pre-header instruction.

12. The product of any one of claims 1 to 11, wherein the at least one first loop instruction comprises a latch instruction to be performed after a last iteration of the second loop, wherein the AGU configuration code is to configure the second dimension of the AGU to configure the memory access operation to be performed only at the end of the second loop.

13. The article of manufacture of claim 12, wherein the latch instruction comprises a load operation.

14. The product of claim 13, wherein the AGU configuration code is used to: set a base address parameter of the AGU to a memory pointer of the latch instruction, set a minimum (Min) parameter of the second dimension of the AGU to zero, set a maximum (Max) parameter of the second dimension of the AGU to an entry size corresponding to the latch instruction, and set a stride parameter of the second dimension of the AGU to zero.

15. The article of manufacture of claim 12, wherein the latch instruction comprises a store operation.

16. The product of claim 15, wherein the AGU configuration code is used to set a base address parameter of the AGU based on a first parameter value, a second parameter value, and a third parameter value, wherein the first parameter value comprises an entry size corresponding to the latch instruction, the second parameter value comprises a total iteration count in the first loop and over one or more loops including the second loop, and the third parameter value comprises a dimension count of the AGU corresponding to the one or more loops.

17. The article of manufacture of claim 16, wherein the AGU configuration code is to set the base address parameter denoted as Base of the AGU as follows: Base=OrigBase+EntrySize*([∑TripCount(L)]-#InnerDims), Wherein OrigBase represents a memory pointer of the latch instruction, EntrySize represents the entry size, [∑TripCount(L)] represents the total iteration count over the one or more loops, and #InnerDims represents the dimension count of the AGU corresponding to the one or more loops.

18. The product of claim 15, wherein the AGU configuration code is used to: set a stride parameter of the second dimension of the AGU based on an entry size corresponding to the latch instruction; set a minimum (Min) parameter of the second dimension of the AGU based on the entry size and an iteration count in the second loop; and set a maximum (Max) parameter of the second dimension of the AGU based on the Min parameter of the second dimension of the AGU and the entry size.

19. The article of claim 18, wherein the AGU configuration code is to set the stride parameter of the second dimension of the AGU based on an additive inverse of the entry size.

20. The article of claim 18, wherein the AGU configuration code is to set the Min parameter of the second dimension of the AGU based on a product of an additive inverse of the entry size and a subtraction result of subtracting one from the iteration count in the second loop.

21. The article of claim 18, wherein the AGU configuration code is to set the Max parameter of the second dimension of the AGU based on a sum of the Min parameter of the second dimension of the AGU and the entry size.

22. The product of any one of claims 1 to 11, wherein the instructions, when executed, cause the compiler to transform the loop nest into a transformed loop that includes the memory access operation, wherein the target code is based on the transformed loop.

23. The article of manufacture of claim 22, wherein the transformed loop comprises a perfectly flat loop in which all computational operations of the loop nest are implemented in the transformed loop.

24. The article of manufacture of claim 22, wherein the transformed loop comprises a fully folded loop comprising only a single basic block loop based on the plurality of loops.

25. The product of any one of claims 1 to 11, wherein the memory access operation comprises a load operation or a store operation.

26. The product of any one of claims 1 to 11, wherein the source code comprises Open Computing Language (OpenCL) code.

27. The product of any one of claims 1 to 11, wherein the computer executable instructions, when executed, cause the compiler to compile the source code into the target code according to a Low Level Virtual Machine (LLVM)-based compilation scheme.

28. The product of any one of claims 1 to 11, wherein the object code is configured for execution by a Very Long Instruction Word (VLIW) Single Instruction / Multiple Data (SIMD) target processor.

29. The product of any one of claims 1 to 11, wherein the object code is configured for execution by a target vector processor.

30. A computing system comprising: at least one memory for storing instructions; as well as at least one processor to retrieve the instructions from the memory and to execute the instructions to cause the computing system to: identifying a loop nest based on source code to be compiled into target code to be executed by a target processor, the loop nest comprising a plurality of loops, the plurality of loops comprising at least a first loop and a second loop nested within the first loop, wherein the first loop comprises at least one first loop instruction outside of the second loop; generating an address generation unit (AGU) configuration code to configure an AGU of the target processor based on the first loop instruction, wherein the AGU configuration code is used to configure a first dimension of the AGU based on the first loop and to configure a second dimension of the AGU based on the second loop, wherein the AGU configuration code is used to configure the second dimension of the AGU to configure a memory access operation to be performed at the beginning of the second loop or at the end of the second loop, wherein the memory access operation is based on the first loop instruction; as well as The object code is generated based on compiling the source code, wherein the object code is based on the AGU configuration code.

31. The computing system of claim 30, comprising the target processor to execute the target code.

32. A method comprising: identifying a loop nest based on source code to be compiled into target code to be executed by a target processor, the loop nest comprising a plurality of loops, the plurality of loops comprising at least a first loop and a second loop nested within the first loop, wherein the first loop comprises at least one first loop instruction outside of the second loop; generating an address generation unit (AGU) configuration code to configure an AGU of the target processor based on the first loop instruction, wherein the AGU configuration code is used to configure a first dimension of the AGU based on the first loop and to configure a second dimension of the AGU based on the second loop, wherein the AGU configuration code is used to configure the second dimension of the AGU to configure a memory access operation to be performed at the beginning of the second loop or at the end of the second loop, wherein the memory access operation is based on the first loop instruction; as well as The object code is generated based on compiling the source code, wherein the object code is based on the AGU configuration code.

33. The method of claim 32, wherein the plurality of loops comprises a third loop nested in the first loop, the second loop nested in the third loop, the first loop instruction being outside the third loop, wherein the AGU configuration code is used to configure a third dimension of the AGU based on the third loop, wherein the AGU configuration code is used to configure the third dimension to configure the memory access operation to be performed at the beginning of the second loop or at the end of the second loop.