Apparatus, system and method for compiling code for processor
By introducing the AGU configuration scheme in the compiler, using the additional AGU dimensions and time shift schemes, the problem of insufficient number of AGUs in the vector processor is solved, and the effect of efficiently handling multiple memory access operations is achieved.
Patent Information
- Application Number
- CN202380071513.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively support the functional requirements of efficient processing, especially when handling multiple memory access operations in a vector processor environment, there is a problem of insufficient number of AGUs.
By introducing an AGU configuration scheme in the compiler, using additional AGU dimensions and time shift schemes, the same AGU is configured to perform multiple memory access operations, reducing the number of AGUs.
It realizes efficient processing of multiple memory access operations under the condition of finite number of AGUs, and improves the processing efficiency of vector processors.
Smart Images

Figure CN119998787A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 415,310, filed on October 12, 2022, entitled “APPARATUS, SYSTEM, AND METHOD OF VECTOR PROCESSING,” the entire disclosure of which is incorporated herein by reference. Background Art
[0003] The compiler may be configured to compile source code into object code configured for execution by the processor.
[0004] There is a need to provide technical solutions to support efficient processing functionality. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity of presentation. In addition, reference numerals may be repeated in the drawings to indicate corresponding or similar elements. The drawings are listed below.
[0006] Figure 1 is a schematic block diagram illustration of a system according to some exemplary aspects.
[0007] Figure 2 is a schematic illustration of a compiler according to some exemplary aspects.
[0008] Figure 3 is a schematic illustration of a vector processor according to some exemplary aspects.
[0009] Figure 4 is a schematic illustration of a memory access scheme to load data from a memory relative to a memory pointer according to some exemplary aspects.
[0010] Figure 5 is a schematic illustration of a memory access scheme to load data from a memory relative to a memory pointer according to some exemplary aspects.
[0011] Figure 6 is a schematic flow chart illustration of a method of compiling code for a processor according to some exemplary aspects.
[0012] Figure 7 is a schematic illustration of a product according to some exemplary aspects. DETAILED DESCRIPTION
[0013] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of some aspects. However, it will be appreciated by those of ordinary skill in the art that some aspects may be practiced without these specific details. In other cases, well-known methods, procedures, components, units and / or circuits are not described in detail to avoid obscuring the discussion.
[0014] Some portions of the following detailed description are presented in terms of algorithms and symbolic representations of operations on data bits or binary digital signals within a computer memory. These algorithmic descriptions and representations may be techniques used by those skilled in the data processing arts to convey the substance of their work to others skilled in the art.
[0015] An algorithm is here and generally considered to be a self-consistent sequence of acts or operations leading to a desired result. These include physical manipulations of physical quantities. Typically, but not necessarily, these quantities capture forms of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for common sense reasons, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. It should be understood, however, that all of these terms and similar terms should be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities.
[0016] Discussions herein utilizing terms such as, for example, "process," "compute," "calculate," "determine," "create," "analyze," "verify," and the like may refer to the operations and / or processes of a computer, computing platform, computing system, or other electronic computing device that manipulates data represented as physical (e.g., electronic) quantities within the computer's registers and / or memories and / or transforms that data into other data similarly represented as physical quantities within the computer's registers and / or memories or other information storage media that may store instructions for performing operations and / or processes.
[0017] As used herein, the terms "plurality" and "a plurality" include, for example, "a plurality" or "two or more." For example, "a plurality of items" includes two or more items.
[0018] References to "one aspect," "an aspect," "exemplary aspect," "various aspects," etc. indicate that the aspects so described may include particular features, structures, or characteristics, but not every aspect necessarily includes the particular features, structures, or characteristics. Furthermore, repeated use of the phrase "in one aspect" does not necessarily refer to the same aspect, although it may.
[0019] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe common objects merely indicates that different instances of the same object are being referenced and is not intended to imply that the objects so described must be in a given sequence in time, space, ranking, or in any other manner.
[0020] For example, some aspects may capture the form of entirely hardware aspects, entirely software aspects, or aspects including both hardware and software elements.Some aspects may be implemented in software, including but not limited to firmware, resident software, microcode, etc.
[0021] Furthermore, some aspects may be captured in the form of a computer program product that can be accessed from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For example, a computer-usable or computer-readable medium can be or can include any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with an instruction execution system, device, or apparatus.
[0022] In some exemplary aspects, the medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium.
[0023] In some exemplary aspects, a data processing system suitable for storing and / or executing program code may include at least one processor coupled directly or indirectly to a memory element, for example, via a system bus. The memory element may include, for example, local memory employed during actual execution of the program code, a mass storage device, and a cache memory that may provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from a mass storage device during execution.
[0024] In some exemplary aspects, input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system directly or through an intermediate I / O controller. In some exemplary aspects, a network adapter can be coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices, such as through an intermediate private or public network. In some exemplary aspects, modems, cable modems, and Ethernet cards are exemplary examples of network adapter types. Other suitable components can be used.
[0025] Some aspects may be used in connection with various devices and systems, such as computing devices, computers, mobile computers, non-mobile computers, server computers, and the like.
[0026] As used herein, the term "circuitry" may refer to, be a part of, or include an application specific integrated circuit (ASIC), an integrated circuit, an electronic circuit, a processor (shared, dedicated, or grouped), and / or memory (shared. Dedicated or grouped) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable hardware components that provide the described functionality. In some aspects, some functions associated with the circuitry may be implemented by one or more software or firmware modules. In some aspects, the circuitry may include logic that is at least partially operable in hardware.
[0027] The term "logic" may refer to, for example, computing logic embedded in the circuit system of a computing device and / or computing logic stored in the memory of a computing device. For example, the logic may be accessed by a processor of a computing device to execute the computing logic to perform computing functions and / or operations. In one example, the logic may be embedded in various types of memory and / or firmware, such as silicon blocks of various chips and / or processors. The logic may be included in various circuit systems and / or implemented as part of various circuit systems, such as processor circuit systems, control circuit systems, and / or the like. In one example, the logic may be embedded in volatile memory and / or non-volatile memory, including random access memory, read-only memory, programmable memory, magnetic memory, flash memory, persistent memory, etc. The logic may be executed by one or more processors using memory (e.g., registers, lags, buffers, and / or the like) coupled to one or more processors, for example, executing the logic as needed.
[0028] Reference now Figure 1 , which schematically illustrates a block diagram of a system 100 according to some exemplary aspects.
[0029] like Figure 1 As shown, in some demonstrative aspects, system 100 may include a computing device 102 .
[0030] In some demonstrative aspects, device 102 may be implemented using suitable hardware components and / or software components, such as processors, controllers, memory units, storage units, input units, output units, communication units, operating systems, applications, and the like.
[0031] In some demonstrative aspects, device 102 may comprise, for example, a computer, a mobile computing device, a non-mobile computing device, a laptop computer, a notebook computer, a tablet computer, a handheld computer, a personal computer (PC), or the like.
[0032] In some exemplary aspects, device 102 may include, for example, one or more of the following: processor 191, input unit 192, output unit 193, memory unit 194, and / or storage unit 195. Device 102 may optionally include other suitable hardware components and / or software components. In some exemplary aspects, some or all components of one or more of devices in device 102 may be enclosed in a common housing or packaging and may be interconnected or operably associated using one or more wired or wireless links. In other aspects, components of one or more of devices in device 102 may be distributed in multiple or separate devices.
[0033] In some exemplary aspects, processor 191 may include, for example, a central processing unit (CPU), a digital signal processor (DSP), one or more processor cores, a single core processor, a dual core processor, a multi-core processor, a microprocessor, a host processor, a controller, multiple processors or controllers, a chip, a microchip, one or more circuits, a circuit system, a logic unit, an integrated circuit (IC), an application specific IC (ASIC), or any other suitable general-purpose or specific processor or controller. Processor 191 may execute, for example, instructions of an operating system (OS) of device 102 and / or instructions of one or more suitable applications.
[0034] In some exemplary aspects, input unit 192 may include, for example, a keyboard, a keypad, a mouse, a touch screen, a touch pad, a trackball, a stylus, a microphone, or other suitable pointing device or input device. Output unit 193 may include, for example, a monitor, a screen, a touch screen, a flat panel display, a light emitting diode (LED) display unit, a liquid crystal display (LCD) display unit, a plasma display unit, one or more audio speakers or headphones, or other suitable output devices.
[0035] In some exemplary aspects, memory unit 194 includes, for example, random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous DRAM (SD-RAM), flash memory, volatile memory, non-volatile memory, cache memory, buffer, short-term memory unit, long-term memory unit, or other suitable memory unit. Storage unit 195 may include, for example, a hard disk drive, a solid-state drive (SSD), or other suitable removable or non-removable storage unit. Memory unit 194 and / or storage unit 195 may, for example, store data processed by device 102.
[0036] In some demonstrative aspects, device 102 may be configured to communicate with one or more other devices via at least one network 103 (eg, a wireless and / or wired network).
[0037] In some exemplary aspects, network 103 may include a wired network, a local area network (LAN), a wireless network, a wireless LAN (WLAN) network, a radio network, a cellular network, a WiFi network, an IR network, a Bluetooth (BT) network, and the like.
[0038] In some demonstrative aspects, device 102 may be configured to perform and / or execute one or more operations, modules, processes, procedures, and / or the like, e.g., as described herein.
[0039] In some demonstrative aspects, device 102 may include compiler 160, which may be configured to generate object code 115 based on source code 112, for example, as described below.
[0040] In some exemplary aspects, compiler 160 may be configured to translate source code 112 into target code 115, eg, as described below.
[0041] In some demonstrative aspects, compiler 160 may include or may be implemented as software, a software module, an application, a program, a subroutine, instructions, an instruction set, computing code, words, values, symbols, and / or the like.
[0042] In some exemplary aspects, source code 112 may include computer code written in a source language.
[0043] In some exemplary aspects, the source language may include a programming language. For example, the source language may include a high-level programming language, such as, for example, C language, C++ language and / or the like.
[0044] In some exemplary aspects, target code 115 may include computer code written in a target language.
[0045] In some exemplary aspects, the target language can include a low-level language such as, for example, assembly language, object code, machine code, or the like.
[0046] In some exemplary aspects, object code 115 may include one or more purpose files, which may, for example, create and / or form an executable program.
[0047] In some exemplary aspects, the executable program may be configured to be executed on a target computer. For example, the target computer may include specific computer hardware, a specific machine and / or a specific operating system.
[0048] In some exemplary aspects, the executable program may be configured to be executed on processor 180, eg, as described below.
[0049] In some demonstrative aspects, processor 180 may include a vector processor 180, eg, as described below. In other aspects, processor 180 may include any other type of processor.
[0050] Some exemplary aspects are described herein with respect to a compiler (e.g., compiler 160) that is configured to compile source code 112 into target code 115 that is configured to be executed by a vector processor 180, e.g., as described below. In other aspects, a compiler (e.g., compiler 160) is configured to compile source code 112 into target code 115 that is configured to be executed by any other type of processor 180.
[0051] In some demonstrative aspects, processor 180 may be implemented as part of device 102 .
[0052] In other aspects, processor 180 may be implemented as part of any other device separate from device 102 , for example.
[0053] In some demonstrative aspects, vector processor 180 (also referred to as an "array processor") may include a processor that may be configured to process an entire vector in one instruction, eg, as described below.
[0054] In other aspects, the executable program may be configured to be executed on any other additional or alternative type of processor.
[0055] In some exemplary aspects, vector processor 180 may be designed to support high-performance image and / or vector processing. For example, vector processor 180 may be configured to process 1 / 2 / 3 / 4D arrays and / or floating point arrays of fixed point data very quickly and / or efficiently.
[0056] In some exemplary aspects, vector processor 180 can be configured to process arbitrary data, such as structures with pointers to structures. For example, vector processor 180 can include a scalar processor to calculate non-vector data, such as assuming that the non-vector data is minimal.
[0057] In some demonstrative aspects, compiler 160 may be implemented as a native application to be executed by device 102. For example, memory unit 194 and / or storage unit 195 may store instructions obtained in compiler 160, and / or processor 191 may be configured to execute instructions obtained in compiler 160 and / or perform one or more calculations and / or processes of compiler 160, e.g., as described below.
[0058] In other aspects, compiler 160 may comprise a remote application to be executed by any suitable computing system (eg, server 170 ).
[0059] In some exemplary aspects, server 170 may include at least a remote server, a network-based server, a cloud server, and / or any other server.
[0060] In some exemplary aspects, server 170 may include a suitable memory and / or storage unit 174 having stored thereon instructions derived from compiler 160 and a suitable processor 171 for executing the instructions, e.g., as described below.
[0061] In some exemplary aspects, compiler 160 may include a combination of remote applications and local applications.
[0062] In one example, compiler 160 may be downloaded and / or received by a user of device 102 from another computing system (e.g., server 170) such that compiler 160 may be executed locally by the user of device 102. For example, instructions may be received and stored temporarily in a memory or any suitable short-term storage or buffer of device 102, e.g., prior to execution by processor 191 of device 102.
[0063] In another example, compiler 160 may include a client module to be executed locally by device 102 and a server module to be executed by server 170. For example, the client module may include and / or may be implemented as a local application, a web application, a website, a web client, e.g., a hypertext markup language (HTML) web application, etc.
[0064] For example, one or more first operations of compiler 160 may be performed locally, such as by device 102 , and / or one or more second operations of compiler 160 may be performed remotely, such as by server 170 .
[0065] In other aspects, compiler 160 may include or be implemented by any other suitable computing arrangement and / or scheme.
[0066] In some demonstrative aspects, system 100 may include an interface 110 (eg, a user interface) to interface between a user of device 102 and one or more elements of system 100 (eg, compiler 160).
[0067] In some demonstrative aspects, interface 110 may be implemented using any suitable hardware components and / or software components, such as a processor, a controller, a memory unit, a storage unit, an input unit, an output unit, a communication unit, an operating system, and / or an application.
[0068] In some aspects, interface 110 may be implemented as part of any suitable module, system, device, or component of system 100 .
[0069] In other aspects, interface 110 may be implemented as a separate element of system 100 .
[0070] In some demonstrative aspects, interface 110 may be implemented as part of device 102. For example, interface 110 may be associated with device 102 and / or included as part of the device.
[0071] In one example, interface 110 can be implemented as part of any suitable application, such as middleware and / or device 102. For example, interface 110 can be implemented as part of compiler 160 and / or part of the OS of device 102.
[0072] In some demonstrative aspects, interface 110 may be implemented as part of server 170. For example, interface 110 may be associated with server 170 and / or included as part of the server.
[0073] In one example, interface 110 may include or be part of: a web-based application, a website, a web page, a plug-in, an ActiveX control, a rich content component (eg, a Flash or Shockwave component), or the like.
[0074] In some exemplary aspects, interface 110 may be associated therewith and / or may include, for example, a gateway (GW) 113 and / or an application programming interface (API) 114, for example, to transmit information and / or communicate between elements of system 100 and / or to one or more other parties (e.g., internal or external parties), users, applications and / or systems.
[0075] In some aspects, interface 110 may include any suitable graphical user interface (GUI) 116 and / or any other suitable interface.
[0076] In some demonstrative aspects, interface 110 may be configured to receive source code 112 from, for example, a user of device 102 via GUI 116 and / or API 114 .
[0077] In some exemplary aspects, interface 110 may be configured to transfer source code 112 to, for example, compiler 160 , for example, to generate object code 115 , for example, as described below.
[0078] refer to Figure 2 , which schematically illustrates a compiler 200 according to some exemplary aspects. For example, the compiler 160 ( Figure 1 ) may implement one or more elements of compiler 200 and / or may perform one or more operations and / or functionalities of compiler 200.
[0079] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may be configured to generate target code 233, for example, by compiling source code 212 in a source language.
[0080] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may include a front end 210 configured to receive and analyze source code 212 in a source language.
[0081] In some exemplary aspects, front end 210 may be configured to generate intermediate code 213 , for example, based on source code 212 .
[0082] In some exemplary aspects, intermediate code 213 may comprise a lower-level representation of source code 212 .
[0083] In some exemplary aspects, front end 210 can be configured to perform, for example, lexical analysis, syntactic analysis, semantic analysis, and / or any other additional or alternative types of analysis of source code 212 .
[0084] In some exemplary aspects, front end 210 can be configured to identify errors and / or problems using the results of the analysis of source code 212. For example, front end 210 can be configured to generate error information, e.g., including error and / or warning messages, which can identify a location in source code 212, e.g., where an error or problem is detected.
[0085] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may include a middle end 220 configured to receive and process intermediate code 213 and generate adjusted (eg, optimized) intermediate code 223 .
[0086] In some demonstrative aspects, middle end 220 may be configured to perform one or more adjustments (eg, optimizations) to intermediate code 213 , eg, to generate adjusted intermediate code 223 .
[0087] In some exemplary aspects, middle end 220 can be configured to perform one or more optimizations on intermediate code 213 , eg, independent of the type of target computer used to execute target code 233 .
[0088] In some exemplary aspects, middle end 220 can be implemented to support the use of optimized intermediate code 223, eg, for different machine types.
[0089] In some exemplary aspects, middle end 220 may be configured to optimize the intermediate representation of intermediate code 223 , for example, to improve the performance and / or quality of the generated target code 233 .
[0090] In some exemplary aspects, one or more optimizations of intermediate code 213 may include, for example, inline expansion, dead code elimination, constant propagation, loop transformation, parallelization, and / or the like.
[0091] In some exemplary aspects, such as Figure 2 As shown, the compiler 200 may include a back end 230 configured to receive and process the adjusted intermediate code 213 , and generate a target code 233 based on the adjusted intermediate code 213 .
[0092] In some exemplary aspects, backend 230 may be configured to perform one or more operations and / or processes that may be specific to a target computer used to execute target code 233. For example, backend 230 may be configured to process optimized intermediate code 213 by applying analysis, transformation, and / or optimization operations to adjusted intermediate code 213, which operations may be configured, for example, based on a target computer used to execute target code 233.
[0093] In some exemplary aspects, the one or more analysis, transformation, and / or optimization operations applied to the adjusted intermediate code 213 may include, for example, resource and storage decisions, such as register allocation, instruction scheduling, and / or the like.
[0094] In some exemplary aspects, the object code 233 may include target-dependent assembly code that may be specific to a target computer used to execute the object code 233 and / or a target operating system of the target computer.
[0095] In some exemplary aspects, the object code 233 may include code for a processor (e.g., vector processor 180 ( Figure 1 ))'s target-dependent assembly code.
[0096] In some exemplary aspects, compiler 200 may include a vector microcode processor (VMP) open computing language (OpenCL) compiler, for example, as described below. In other aspects, compiler 200 may include any other type of vector processor compiler, or may be implemented as part of any other type of vector processor compiler.
[0097] In some exemplary aspects, the VMP OpenCL compiler may include a low-level virtual machine (LLVM)-based compiler that may be configured according to an LLVM-based compilation scheme, for example, to reduce OpenCL C code to VMP accelerator assembly code, for example, suitable for use by vector processor 180 ( Figure 1 )implement.
[0098] In some exemplary aspects, compiler 200 may include one or more techniques that may be required to compile code into a format suitable for a VMP architecture, for example, in addition to an open source LLVM compiler pass.
[0099] In some exemplary aspects, FE 210 may be configured to parse OpenCL C code and translate it, for example, via an abstract syntax tree (AST), into, for example, an LLVM intermediate representation (IR).
[0100] In some exemplary aspects, compiler 200 may include a dedicated API, e.g., to detect the correct pattern for compiler pattern matching, e.g., a pattern suitable for VMP. For example, VMP may be configured as a complex instruction set computer (CISC) machine that implements a very complex instruction set architecture (ISA) that may be difficult to target from standard C code. Accordingly, compiler pattern matching may not be able to easily detect the correct pattern, and for such cases, the compiler may require a dedicated API.
[0101] In some exemplary aspects, FE 210 may implement one or more vendor extension builtins that may target a VMP-specific ISA, for example, in addition to standard OpenCL builtins that may be optimized for VMP machines.
[0102] In some exemplary aspects, FE 210 may be configured to implement OpenCL constructs and / or work-item functionality.
[0103] In some exemplary aspects, ME 220 may be configured to process LLVM IR code, which may be generic and target-independent, e.g., although it may include one or more hooks for a specific target architecture.
[0104] In some demonstrative aspects, ME 220 may perform one or more custom passes, for example, to support a VMP architecture, for example, as described below.
[0105] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of control flow graph (CFG) linearization analysis, e.g., as described below.
[0106] In some exemplary aspects, the CFG linearization analysis can be configured to linearize the code, for example, by converting if statements to select modes, for example, where the VMP vector code does not support standard control flow.
[0107] In one example, ME 220 may receive a given code, for example, as follows:
[0108]
[0109]
[0110] According to this example, ME 220 may be configured to apply CFG linearization analysis to a given code, for example, as follows:
[0111] tmpA=A+5;
[0112] tmpB = B*2;
[0113] mask=x>0;
[0114] A=Select mask,tmpA,A
[0115] B=Select not mask,tmpB,B
[0116] Example (1)
[0117] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of automatic vectorization analysis, e.g., as described below.
[0118] In some exemplary aspects, the auto-vectorization analysis may be configured to vectorize (eg, auto-vectorize) a given code, for example, to exploit the vector capabilities of the VMP.
[0119] In some exemplary aspects, ME 220 may be configured to perform automatic vectorization analysis, e.g., to vectorize code into scalar form. For example, some or all operations of automatic vectorization analysis may not be performed, such as when the code is already provided in vectorized form.
[0120] In some exemplary aspects, for example, in some use cases and / or scenarios, a compiler may not always be able to auto-vectorize code, for example, due to data dependencies between loop iterations.
[0121] In one example, ME 220 may receive a given code, for example, as follows:
[0122]
[0123] According to this example, ME 220 may be configured to perform CFG auto-vectorization analysis by applying a first transformation, for example, as follows:
[0124]
[0125] Example (2a)
[0126] For example, ME 220 may be configured to perform CFG auto-vectorization analysis by applying a second transformation, e.g., after a first transformation, e.g., as follows:
[0127]
[0128] Example (2b)
[0129] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of scratch pad memory cycle access analysis (SPMLAA), e.g., as described below.
[0130] In some exemplary aspects, the SPMLAA may define a processing block (PB), for example, that should later be outlined and compiled for VMP.
[0131] In some exemplary aspects, a processing block may include an accelerated loop that may be executed by a vector unit of a VMP.
[0132] In some exemplary aspects, a PB (eg, each PB) may include memory references. For example, some or all memory accesses may refer to a local memory bank.
[0133] In some exemplary aspects, the VMP may enable the AGU (e.g., as described below with reference to Figure 3 The AGU 320 described above and the scatter-gather unit (SG) are used to access the memory bank.
[0134] In some exemplary aspects, the AGU can be pre-configured, for example, before a loop is executed. For example, a loop trip count can be calculated, for example, before running a processing block.
[0135] In some exemplary aspects, image references can be created at this stage, eg, some or all image references, and strides and offsets can then be calculated, eg, per-dimension strides and offsets for each reference.
[0136] In some exemplary aspects, ME 220 may be configured to perform one or more operations of AGU planner analysis, e.g., as described below.
[0137] In some exemplary aspects, the AGU planner analysis can include an iterator specification that can cover image references from an entire processing block, eg, all image references.
[0138] In some exemplary aspects, an iterator may cover a single reference or a group of references.
[0139] In some exemplary aspects, one or more memory references may be combined via a shuffle instruction and / or reuse the same access, and / or preserve values read from a previous iteration.
[0140] In some exemplary aspects, other memory references, such as those without a linear access pattern, may be processed using a scatter-gather (SG) unit, which may have a performance penalty, such as because it may need to maintain indexes and / or masks.
[0141] In some exemplary aspects, a plan may be configured as an arrangement of iterators in a processing block. For example, a processing block may, for example, theoretically have multiple plans.
[0142] In some exemplary aspects, the AGU planner analysis can be configured to construct all possible plans for all PBs and select a combination, eg, the best combination, from among all valid combinations.
[0143] In some exemplary aspects, the total number of iterators in a valid combination may be limited, eg, not to exceed the number of available AGUs on the VMP.
[0144] In some exemplary aspects, one or more parameters may be defined for an iterator (e.g., for each iterator), e.g., including stride, width, and / or cardinality, e.g., as part of an AGU planner analysis. For example, a minimum-maximum range for an iterator may be defined dimensionally, e.g., in each dimension, e.g., as part of an AGU planner analysis.
[0145] In some exemplary aspects, the AGU planner analysis can be configured to track and evaluate memory references to the image, eg, each memory reference, eg, to understand its access pattern.
[0146] In one example, according to Example 2a / 2b, image "a" as a base address can be accessed with 64 iterations using a step size of 32 bytes.
[0147] In some exemplary aspects, LLVM can include scalar evaluation analysis (SCEV) that can compute access patterns, for example, to understand each image reference.
[0148] In some exemplary aspects, ME 220 may exploit the masking capabilities of the AGU, eg, to avoid maintaining induction variables, which may have a performance penalty.
[0149] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of rewrite analysis, e.g., as described below.
[0150] In some exemplary aspects, the rewrite analysis may be configured to transform the code of a processing block, for example, when setting up iterators and / or modifying memory access instructions.
[0151] In some exemplary aspects, the setup of iterators (e.g., all iterators) can be implemented in the IR in a target-specific intrinsic function. For example, the setup of iterators can reside in the pre-header of the outermost loop.
[0152] In some exemplary aspects, the rewrite analysis may include a loop completion analysis, eg, as described below.
[0153] In some exemplary aspects, the code may be compiled with the goal that substantially all computations should be performed within the innermost loop.
[0154] For example, loop finishing analysis may promote instructions, for example, to move operations performed after the last iteration of the loop into the loop.
[0155] For example, loop finishing analysis may sink instructions, eg, to move operations performed before the first iteration of the loop into the loop.
[0156] For example, loop finishing analysis may hoist instructions and / or sink instructions, eg, such that substantially all instructions from an outer loop are moved to an innermost loop.
[0157] For example, loop completion analysis may be configured to provide technical solutions to support VMP iterators, for example, to work only on perfectly nested loops.
[0158] For example, loop completion analysis may lead to a situation where there are no instructions between "for" statements that make up a loop, e.g., to support VMP iterators, which cannot emulate such a situation.
[0159] In some exemplary aspects, loop completion analysis may be configured to collapse nested loops into a single collapsed loop.
[0160] In one example, ME 220 may receive a given code, for example, as follows:
[0161]
[0162] According to this example, ME 220 may be configured to perform loop completion analysis to collapse nested loops in the code into a single collapsed loop, for example, as follows:
[0163]
[0164] Example (3)
[0165] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of vector loop delimitation analysis, eg, as described below.
[0166] In some exemplary aspects, the vector loop demarcation analysis can be configured to partition the code between the scalar subsystem and the vector subsystem, for example, as described below with reference to Figure 3 The vector processing block 310 ( Figure 3 ) and scalar processor 330( Figure 3 )between.
[0167] In some exemplary aspects, a VMP accelerator may include scalar and / or vector subsystems, e.g., as described below. For example, each of the subsystems may have a different computational unit / processor. Accordingly, scalar code may be compiled on a scalar compiler (e.g., an SSC compiler), and / or accelerated vector code may run on a VMP vector processor.
[0168] In some exemplary aspects, vector loop demarcation analysis can be configured to create separate functions for accelerating loop bodies of vector code. For example, these functions can be marked for VMP and / or can proceed to the VMP backend, for example, while the rest of the code can be compiled by the SSC compiler.
[0169] In some exemplary aspects, one or more portions of a vector loop (e.g., configuration of a vector unit and / or initialization of vector registers) may be performed by a scalar unit. However, these portions may be performed at a later stage, e.g., by backfilling the scalar code, e.g., because the scalar code may still be in LLVM IR before being processed by the SSC compiler.
[0170] In some exemplary aspects, BE 230 may be configured to translate LLVM IR into machine instructions. For example, BE 230 may not be target agnostic and may be familiar with target specific architectures and optimizations, for example, compared to ME 220 which may be agnostic to target specific architectures.
[0171] In some exemplary aspects, BE 230 may be configured to perform one or more analyses that may be specific to the target machine (eg, a VMP machine) to which the code is being downgraded, for example, even though BE 230 may use a general-purpose LLVM.
[0172] In some exemplary aspects, BE 230 may be configured to perform one or more operations of instruction degradation analysis, eg, as described below.
[0173] In some exemplary aspects, instruction degradation analysis may be configured to translate LLVM IR into target-specific instruction machine IR (MIR), for example, by translating LLVM IR into a directed acyclic graph (DAG).
[0174] In some exemplary aspects, the DAG may undergo a legalization process for instructions, such as based on data types and / or VMP instructions, which may be supported by the VMP HW.
[0175] In some exemplary aspects, instruction demotion analysis may be configured to, for example, perform a pattern matching process after a legalization process of instructions, for example, to demotion nodes (eg, each node) in a DAG to, for example, VMP-specific machine instructions.
[0176] In some exemplary aspects, instruction degradation analysis may be configured to generate a MIR, for example, after a pattern matching process.
[0177] In some exemplary aspects, instruction demotion analysis may be configured to degrade instructions according to a machine application binary interface (ABI) and / or calling convention.
[0178] In some exemplary aspects, BE 230 can be configured to perform one or more operations of a cell balance analysis, eg, as described below.
[0179] In some exemplary aspects, the unit balancing analysis may be configured to balance instructions among VMP computing units, for example, as described below with reference to Figure 3 The data processing unit 316 ( Figure 3 )between.
[0180] In some exemplary aspects, the cell balance analysis may be aware of some or all available arithmetic transformations, and / or may perform transformations according to an optimal algorithm.
[0181] In some exemplary aspects, BE 230 may be configured to perform one or more operations of a modulo scheduler (pipeliner) analysis, eg, as described below.
[0182] In some exemplary aspects, the pipeliner may be configured to schedule instructions according to one or more constraints (e.g., data dependencies, resource bottlenecks, and / or any other constraints), for example using a swing modulo scheduling (SMS) heuristic and / or any other additional and / or alternative heuristics.
[0183] In some exemplary aspects, the pipeliner can be configured to schedule a set of very long instruction word (VLIW) instructions (eg, of initiation intervals (II)) over which a program will iterate, such as during a steady state.
[0184] In some exemplary aspects, a performance metric may be measured, which may be based on the number of cycles a typical loop may execute, for example, as follows:
[0185] (input data size in bytes)*II / (bytes consumed / produced per iteration)
[0186] In some exemplary aspects, the pipeliner can attempt to minimize II as much as possible, for example, to improve performance.
[0187] In some exemplary aspects, the pipeliner can be configured to calculate a minimum II and schedule accordingly. For example, if the pipeliner fails to schedule, the pipeliner can attempt to increase the II and retry scheduling, for example, until a predefined II threshold is violated.
[0188] In some demonstrative aspects, BE 230 may be configured to perform one or more operations of register allocation analysis, eg, as described below.
[0189] In some exemplary aspects, register allocation analysis can be configured to attempt to assign registers in an efficient (eg, optimal) manner.
[0190] In some exemplary aspects, register allocation analysis may assign values to bypass vector registers, general purpose vector registers, and / or scalar registers.
[0191] In some exemplary aspects, the values may include private variables, constants, and / or values that rotate across iterations.
[0192] In some exemplary aspects, register allocation analysis may implement an optimal heuristic that fits one or more VMP register file (regfile) constraints. For example, in some use cases, register allocation analysis may not use standard LLVM register allocation.
[0193] In some exemplary aspects, in some cases, register allocation analysis may fail, which may mean that the loop cannot be compiled. Accordingly, register allocation analysis may implement a retry mechanism that may return to the modulo scheduler and may attempt to reschedule the loop, e.g., with an increased launch interval. For example, in many cases, increasing the launch interval may reduce register starvation and / or may support compilation of vector loops.
[0194] In some demonstrative aspects, BE 230 may be configured to perform one or more operations of SSC configuration analysis, eg, as described below.
[0195] In some exemplary aspects, the SSC configuration analysis may be configured to set a configuration for executing a kernel, such as an AGU configuration.
[0196] In some exemplary aspects, SSC configuration analysis may be performed at a later stage, such as due to configurations being calculated after legalization, register allocation analysis, and / or modulo scheduling analysis.
[0197] In some exemplary aspects, the SSC configuration analysis can include a zero overhead loop (ZOL) mechanism in a vector loop. For example, the ZOL mechanism can configure loop trip counts based on access patterns of memory references in the loop, e.g., to avoid running instructions that check loop exit conditions for each iteration.
[0198] In some exemplary aspects, a VMP compilation flow may include one or more (e.g., a small number) of steps that may be called during the compilation flow in a test library (testlib) (e.g., a wrapper script for compilation, execution, and / or program testing). For example, these steps may be performed outside of the LLVM compiler.
[0199] In some exemplary aspects, a PCB Hardware Description Language (PHDL) simulator can be implemented to perform one or more roles of an assembler, an encoder, and / or a linker.
[0200] In some exemplary aspects, compiler 200 can be configured to provide technical solutions to support robustness, which can enable compilation of a wide range of loop selections with HW limitations. For example, compiler 200 can be configured to support technical solutions that may not generate verification errors.
[0201] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support programmability, which may provide users with the ability to express code in a variety of ways that may compile correctly to a VMP architecture.
[0202] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support an improved user experience, which may allow a user to debug and / or profile code. For example, the improved user experience may provide informative error messages, reporting tools, and / or profiling tools.
[0203] In some exemplary aspects, compiler 200 can be configured to provide technical solutions to support improved performance, for example, to optimize VMP assembly code and / or iterator access, which may result in faster execution. For example, improved performance can be achieved through high utilization compute units and using their complex CISC.
[0204] refer to Figure 3 , which schematically illustrates a vector processor 300 according to some exemplary aspects. For example, the vector processor 180 ( Figure 1) may implement one or more components of the vector processor 300 and / or may perform one or more operations and / or functionalities of the vector processor 300.
[0205] In some exemplary aspects, vector processor 300 may comprise a vector microcode processor (VMP).
[0206] In some exemplary aspects, vector processor 300 may include a wide vector machine, eg, supporting a very long instruction word (VLIW) architecture and / or a single instruction / multiple data (SIMD) architecture.
[0207] In some exemplary aspects, vector processor 300 may be configured to provide a technical solution to support high performance for short integer types, which may be common in, for example, computer vision and / or deep learning algorithms.
[0208] In other aspects, the vector processor 300 may include any other type of vector processor, and / or may be configured to support any other additional or alternative functionality.
[0209] In some exemplary aspects, such as Figure 3 As shown, the vector processor 300 may include a vector processing block (vector processor) 310, a scalar processor 330, and a direct memory access (DMA) 340, for example, as described below.
[0210] In some exemplary aspects, such as Figure 3 As shown, the vector processing block 310 may be configured to process (eg, efficiently process) image data and / or vector data. For example, the vector processing block 310 may be configured to use a vector computing unit, for example, to accelerate computation.
[0211] In some exemplary aspects, scalar processor 330 may be configured to perform scalar calculations. For example, scalar processor 330 may be used as "glue logic" for a program that includes vector calculations. For example, some (e.g., even most) of the calculations of a program may be performed by vector processing block 310. However, several tasks (e.g., some basic tasks) (e.g., scalar calculations) may be performed by scalar processor 330.
[0212] In some demonstrative aspects, DMA 340 may be configured to interface with one or more memory elements in a chip including vector processor 300 .
[0213] In some demonstrative aspects, DMA 340 may be configured to read input from main memory, and / or write output to main memory.
[0214] In some exemplary aspects, scalar processor 330 and vector processing block 310 may use respective local memories to process data.
[0215] In some exemplary aspects, such as Figure 3 As shown, the vector processor 300 may include an extractor and decoder 350 , which may be configured to control the scalar processor 330 and / or the vector processing block 310 .
[0216] In some exemplary aspects, operations of scalar processor 330 and / or vector processing block 310 may be triggered by instructions stored in program memory 352 .
[0217] In some demonstrative aspects, DMA 340 may be configured to transfer data in parallel with the execution of program instructions in memory 352, for example.
[0218] In some exemplary aspects, DMA 340 may be controlled by software, such as via configuration registers, rather than instructions, for example, and accordingly may be considered a second “thread” of execution in vector processor 300 .
[0219] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include one or more data processing units, e.g., a group of data processing units, e.g., as described below.
[0220] In some exemplary aspects, a data processing unit may include hardware configured to perform calculations, such as an arithmetic logic unit (ALU).
[0221] In one example, the data processing unit may be configured to add numbers and / or store numbers in memory.
[0222] In some exemplary aspects, the data processing unit may be controlled by commands encoded in, for example, program memory 352 and / or configuration registers. For example, the configuration registers may be memory mapped and writeable by memory storage commands of scalar processor 330.
[0223] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include a state configuration including a set of registers and memory, eg, as described below.
[0224] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a set of vector memories 312 , which may be configured, for example, to store data to be processed by the vector processor block 310 .
[0225] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a set of vector registers 314 that may be configured for use, for example, in data processing performed by the vector processor block 310 .
[0226] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may be associated with a set of memory maps.
[0227] In some exemplary aspects, a memory map may include a set of addresses accessible by a data processing unit that may load data from / to registers and memory and / or store data.
[0228] In some exemplary aspects, such as Figure 3 As shown, the vector processing block 310 may include a plurality of address generation units (AGUs) 320 , which may include addresses accessible to them, for example, in one or more memories in the memory 312 .
[0229] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a plurality of data processing units 316, for example, as described below.
[0230] In some exemplary aspects, the data processing unit 316 may be configured to process commands, for example, including a number of digits at a time. In one example, the command may include 8 digits. In another example, the command may include 4 digits, 16 digits, or any other count of digits.
[0231] In some exemplary aspects, two or more data processing units 316 can be used simultaneously. In one example, data processing unit 316 can process and execute multiple different commands, for example, 3 different commands, for example, including 8 numbers, in a single cycle.
[0232] In some exemplary aspects, the data processing units 316 may be asymmetric. For example, the first and second data processing units 316 may support different commands. For example, addition may be performed by the first data processing unit 316, and / or multiplication may be performed by the second data processing unit 316. For example, both operations may be performed by one or more additional data processing units 316.
[0233] In some demonstrative aspects, data processing unit 316 may be configured to support arithmetic operations for many combinations of input and output data types.
[0234] In some exemplary aspects, data processing unit 316 may be configured to support one or more operations, which may be less common. For example, processing unit 316 may support operations to work with a lookup table (LUT) of vector processor 300 and / or any other operations.
[0235] In some exemplary aspects, data processing unit 316 may be configured to support efficient computation of nonlinear functions, histograms, and / or random data access, which may, for example, facilitate implementation of algorithms like image scaling, Hough transform, and / or any other algorithm.
[0236] In some exemplary aspects, vector memory 312 may include a bank of memory having a size of 16K, for example, or any other size, that may be accessed in the same cycle.
[0237] In one example, the maximum memory access size may be 64 bits. According to this example, the peak throughput may be 256 bits, for example, 64×4=256. For example, a high memory bandwidth may be achieved to utilize the computational power of the data processing unit 316 .
[0238] In one example, two data processing units 316 may support 16 8-bit multiply and accumulate operations (MACs) per cycle. According to this example, two data processing units 316 may not be useful, for example, if the input numbers are not extracted at that speed, and / or there is no input of exactly 256 bits, for example, 16x8x2=256.
[0239] In some exemplary aspects, AGU 320 may be configured to perform memory access operations, such as loading and storing data from / to vector memory 314 .
[0240] In some exemplary aspects, AGU 320 may be configured to calculate addresses of input and output data items, for example, to handle I / O in situations where high bandwidth is not sufficient to utilize data processing unit 316 .
[0241] In some exemplary aspects, AGU 320 may be configured to calculate addresses of input and / or output data items, for example, based on configuration registers written by scalar processor 330, prior to entering a vector command block (eg, a loop).
[0242] For example, the AGU 320 may be configured to write an image base pointer, width, height, and / or stride to configuration registers, for example, to iterate over an image.
[0243] In some exemplary aspects, the AGU 320 may be configured to handle addressing (e.g., all addressing), for example, to provide a technical solution in which the data processing unit 316 may not have the burden of incrementing a pointer or counter in a loop and / or the burden of checking a line end condition, for example, to zero a counter in a loop.
[0244] In some exemplary aspects, such as Figure 3As shown, the AGU 320 may include four AGUs, and accordingly, four memories 312 may be accessed in the same cycle. In other aspects, any other count of AGUs 32 may be implemented.
[0245] In some exemplary aspects, AGUs 320 may not be "bound" to memory banks 312. For example, an AGU 320 (e.g., each AGU 320) may access a memory bank 312 (e.g., each memory bank 312), e.g., as long as two or more AGUs 320 do not attempt to access the same memory bank 312 in the same cycle.
[0246] In some demonstrative aspects, vector registers 314 may be configured to support communications between data processing unit 316 and AGU 320 .
[0247] In one example, the total number of vector registers 314 may be 28, which may be divided into several subsets, for example, based on their functions. For example, a first subset of vector registers 314 may be used for input / output of, for example, all data processing units 316 and / or AGU 320; and / or a second subset of vector registers 314 may not be used for output of some operations (e.g., most operations) and may be used for one or more other operations, for example, to store loop-invariant inputs.
[0248] In some exemplary aspects, a data processing unit 316 (e.g., each data processing unit 316) may have one or more registers to host the output of the last performed operation, e.g., which may be fed as input to other data processing units 316. For example, these registers may "bypass" vector registers 314 and may operate faster than writing these outputs to the first set of vector registers 314.
[0249] In some exemplary aspects, the extractor and decoder 350 may be configured to support low-overhead vector loops, e.g., very low-overhead vector loops (also referred to as "zero-overhead vector loops"), e.g., where a termination (exit) condition of the vector loop may not need to be checked during execution of the vector loop.
[0250] For example, the AGU 320 may signal a termination (exit) condition, such as when the AGU 320 completes iterations over a configured memory region.
[0251] For example, the fetcher and decoder 350 may exit the loop when, for example, the AGU 320 signals a termination condition.
[0252] For example, the scalar processor 330 may be utilized to configure loop parameters, such as the first and last instructions and / or exit conditions.
[0253] In one example, vector loops may be utilized, for example, together with high memory bandwidth and / or cheap addressing, for example, to solve control and data flow problems, for example, to provide a technical solution to allow data processing unit 316 to process data with substantially no additional overhead.
[0254] In some exemplary aspects, scalar processor 330 may be configured to provide one or more functionalities that may be complementary to the functionality of vector processing block 310. For example, a large portion (e.g., most) of the work in a vector program may be performed by data processing unit 316. For example, scalar processor 330 may be utilized, for example, to "glue" together various vector code blocks of a vector program.
[0255] In some exemplary aspects, the scalar processor 330 may be implemented separately from the vector processing block 310. In other aspects, the scalar processor 330 may be configured to share one or more components and / or functionality with the vector processing block 310.
[0256] In some exemplary aspects, scalar processor 330 may be configured to perform operations that may not be suitable for execution on vector processing block 310 .
[0257] For example, the scalar processor 330 may be utilized to execute a 32-bit C program. For example, the scalar processor 330 may be configured to support 1, 2, and / or 4-byte data types of the C code and / or some or all arithmetic operators of the C code.
[0258] For example, the scalar processor 330 may be configured to provide a technical solution to perform operations that cannot be performed on the vector processing block 310 without, for example, using a full CPU.
[0259] In some exemplary aspects, scalar processor 330 may include a scalar data memory 332 , for example, having a size of 16K or any other size, which may be configured to store data, such as variables used by a scalar portion of a program.
[0260] For example, scalar processor 330 may store local and / or global variables declared by portable C code, which may be compiled by a compiler (e.g., compiler 200 ( Figure 2 )) is allocated to scalar data memory.
[0261] In some exemplary aspects, such as Figure 3 As shown, the scalar processor 330 may include or may be associated with a set of vector registers 334 that may be used for data processing by the scalar processor 330 .
[0262] In some exemplary aspects, the scalar processor 330 can be associated with a scalar memory map that can enable the scalar processor 330 to access substantially all states of the vector processor 300. For example, the scalar processor 330 can configure a vector unit and / or a DMA channel via the scalar memory map.
[0263] In some exemplary aspects, the scalar processor 330 may not be allowed to access one or more block control registers that may be used by an external processor to run and debug a vector program.
[0264] In some exemplary aspects, DMA 340 can be configured to communicate, for example, via main memory, with one or more other components of a chip implementing vector processor 300. For example, DMA 340 can be configured to transfer blocks of data, for example, large, contiguous blocks of data, for example, to support scalar processor 330 and / or vector processing blocks that can manipulate data stored in local memory. For example, a vector program may be able to use DMA 340 to read data from main chip memory.
[0265] In some exemplary aspects, DMA 340 may be configured to communicate with other elements of the chip, for example, via a plurality of DMA channels (e.g., 8 DMA channels or any other count of DMA channels). For example, a DMA channel (e.g., each DMA channel) may be able to transfer a rectangular patch from a local memory to a main chip memory, or vice versa. In other aspects, a DMA channel may transfer any other type of data block between a local memory and a main chip memory.
[0266] In some exemplary aspects, a rectangular tile may be defined by a base pointer, a width, a height, and a stride.
[0267] For example, at peak throughput, 8 bytes may be transferred per cycle, however, there may be an overhead for each tile and / or for each row in a tile.
[0268] In some exemplary aspects, DMA 340 can be configured to transfer data in parallel with computations, such as via multiple DMA channels, for example, as long as the executed commands do not access local memory involved in the transfer.
[0269] In one example, since all channels can access the same memory bus, using several channels to implement a transfer may not save I / O cycles, for example, compared to when a single channel is used. However, multiple DMA channels can be utilized to schedule several transfers and execute them in parallel with the calculation. For example, this may be advantageous compared to a single channel, which may not allow a second transfer to be scheduled before the first transfer is completed.
[0270] In some exemplary aspects, DMA 340 can be associated with a memory map that can support DMA channels accessing vector memory and / or scalar data. For example, access to vector memory can be performed in parallel with computation. For example, access to scalar data may not generally allow for parallelism, e.g., because scalar processor 330 may be involved in almost any reasonable program and may access its local variables while performing a transfer, which may result in memory contention with active DMA channels.
[0271] In some exemplary aspects, DMA 340 can be configured to provide a technical solution to support parallelization of I / O and computation. For example, a program performing computations may not have to wait for I / O, for example, when these computations can be run quickly by vector processing block 310.
[0272] In some exemplary aspects, an external processor (eg, a CPU) may be configured to initiate execution of a program on vector processor 300. For example, vector processor 300 may remain idle, for example, as long as program execution is not initiated.
[0273] In some exemplary aspects, the external processor may be configured to debug the program, for example, to execute a single step at a time, to stop when the program reaches a breakpoint, and / or to examine the contents of registers and memory storing program variables.
[0274] In some exemplary aspects, external memory mapping may be implemented to enable an external processor to control the vector processor 300 and / or a debugger, for example, by writing to control registers of the vector processor 300 .
[0275] In some exemplary aspects, the external memory map may be implemented by a superset of the scalar memory map. For example, the implementation may make all registers and memories defined by the architecture of the vector processor 300 accessible to a debugger backend running on an external processor.
[0276] In some exemplary aspects, the vector processor 300 may issue an interrupt signal, such as when the vector processor 300 terminates a program.
[0277] In some exemplary aspects, the interrupt signal may be used, for example, to implement a driver to maintain a queue of programs scheduled for execution by vector processor 300 and / or may be used to start a new program, for example, by an external processor, when a previously executed program completes.
[0278] Return to reference Figure 1 In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on one or more loops, which may be based on, for example, source code 112, for example, as described below.
[0279] In some exemplary aspects, one or more loops may be nested within loop nests, eg, as described below.
[0280] In some exemplary aspects, a loop nest may include at least an outer loop and an inner loop, eg, as described below.
[0281] In some exemplary aspects, a loop nest may include an outer loop (e.g., an outermost loop), an inner loop (e.g., an innermost loop), and one or more nested loops (also referred to as “middle nested loops” or “middle loops”), which may be nested between the outer loop and the inner loop, e.g., as described below.
[0282] In some exemplary aspects, a loop nest may include multiple loops nested in multiple nesting levels, eg, as described below.
[0283] In one example, the plurality of loops may include a first loop (eg, an outer loop), eg, in a first nesting level, and a second loop (eg, an inner loop), eg, in a second nesting level.
[0284] In one example, the plurality of loops can include one or more intermediate loops, eg, in one or more intermediate nesting levels, eg, between a first nesting level and a second nesting level.
[0285] In one example, the plurality of loops may include three loops in three nesting levels. For example, the three loops may include a first loop at a first nesting level, e.g., an outermost loop; a second loop at a second nesting level, e.g., an intermediate loop; and a third loop at a third nesting level, e.g., an innermost loop. For example, the second loop may be nested in the first loop, and the third loop may be nested in the second loop.
[0286] In some exemplary aspects, one or more loops that may be based on, for example, source code 112 may include, for example, multiple memory access operations.
[0287] In some exemplary aspects, the memory access operation may include a load operation and / or a store operation. In other aspects, the memory access operation may include any other suitable type of memory access operation.
[0288] In some exemplary aspects, compiler 160 may configure one or more AGUs to perform memory access operations, eg, as described below.
[0289] In some exemplary aspects, for example, in some use cases, implementations, and / or scenarios, a memory access operation may require a certain number of AGUs, for example, while the number of available AGUs of a target processor (eg, a vector processor) may be limited.
[0290] In one example, a loop may include code, such as code of an OpenCL program and / or any other code, that may include many load operations and / or store operations, while a vector processor may have only a relatively low number of available AGUs.
[0291] In some exemplary aspects, there may be a need to provide a technical solution to handle multiple memory access operations, such as in situations when a count of memory access operations exceeds a count of available AGUs, such as described below.
[0292] In some exemplary aspects, eg, in some use cases, implementations, and / or scenarios, multiple memory access operations may be performed relative to the same memory pointer and different offsets.
[0293] For example, an implementation that uses a different AGU for each memory access operation with a different offset may result in a relatively high number of AGUs required to process multiple memory access operations, for example. For example, the count of available AGUs may be limited, for example, less than the required number of AGUs.
[0294] In some exemplary aspects, for example, compiler 160 may be configured to compile source code 112 of an executed program to be executed on a target processor including a limited number of AGUs, such as vector processor 180. For example, a vector processor may include four AGUs or any other count of AGUs.
[0295] For example, compiler 160 may identify, for example, based on source code 112, a loop nest including five memory access operations, e.g., as follows:
[0296]
[0297] Example (4)
[0298] For example, as shown in Example 4, the loop nest may include an outer loop (y loop) along a dimension (Y dimension) based on a variable height and an inner loop (x loop) along a dimension (X dimension) based on a variable width.
[0299] For example, as shown in Example 4, a loop nest may include a store operation out[index] and a load operation inp2[index], which may require two AGUs.
[0300] For example, as shown in Example 4, the loop nest may include three load operations, eg, with different offsets, to the same memory pointer denoted as inp1.
[0301] For example, as shown in Example 4, the loop nest may include a first load operation (e.g., inp1[index]) on the memory pointer inp1 with a first offset (e.g., index); a second load operation (e.g., inp1[index+width]) on the memory pointer inp1 with a second offset (e.g., index+width); and a third load operation (e.g., inp1[index+2*width]) on the memory pointer inp1 with a third offset (e.g., index+2*width).
[0302] refer to Figure 4 , which schematically illustrates a memory access scheme 400 to load data from a memory relative to a memory pointer, according to some exemplary aspects.
[0303] In one example, the memory access scheme 400 can demonstrate memory access relative to the same memory pointer (eg, memory pointer inp1 ), such as according to the loop nest of Example 4.
[0304] like Figure 4 As shown, the three load operations of Example 4 (eg, inp1[index], inp1[index+width], and inp1[index+2*width]) can load data with different offsets, eg, offsets of index, index+width, and index+2*width, respectively.
[0305] like Figure 4 As shown, the three load operations of Example 4 can load data with different offsets relative to the same memory pointer inp1, for example, in the same Y dimension, for example, with the same difference in the Y dimension, for example, offsets of 0, 1 and 2 (*width).
[0306] For example, an implementation that requires using a different AGU for each memory access with a different offset may require a total of 5 AGUs to process the memory access operation of instance 4. For example, a first AGU may be used to process the store operation out[index], a second AGU may be used to process the load operation inp2[index], and three additional AGUs may be used to process three corresponding load operations relative to the memory pointer inp1, e.g., operations inp1[index], inp1[index+width], and inp1[index+2*width]. According to this example, if the count of available AGUs of the target processor is less than 5, for example, the target processor may not be able to process the memory access operation of instance 4.
[0307] Return to reference Figure 1In some exemplary aspects, compiler 160 may be configured to generate target code 115 that may, for example, be configured to utilize an AGU of a target processor (e.g., vector processor 180), for example, according to an AGU configuration scheme, for example, as described below.
[0308] In some exemplary aspects, the AGU configuration scheme may be configured to provide a technical solution to support configuring the same AGU, e.g., to perform multiple memory access operations to the same memory pointer, e.g., with different offsets, e.g., as described below.
[0309] In some exemplary aspects, the AGU configuration scheme may be configured to provide a technical solution to support implementation of multiple data accesses, e.g., using a reduced number of AGUs, e.g., in zero-overhead loops and / or any other type of loops, e.g., as described below.
[0310] In some exemplary aspects, an AGU configuration scheme may be configured to use at least one additional dimension of the AGU, for example, to process multiple memory access operations, eg, as described below.
[0311] In some exemplary aspects, the AGU configuration scheme can be configured to use a time-shifting scheme, for example, to handle multiple memory access operations, for example, as described below.
[0312] In some exemplary aspects, the time-shifting scheme can be configured to provide a technical solution to preserve load history, for example, to handle multiple memory access operations, for example, as described below.
[0313] In some exemplary aspects, the AGU configuration scheme may be configured to utilize additional AGU dimensions and / or time-shifting schemes, for example, to provide a technical solution to support jumping across different memory access operations (e.g., load and / or store operations), for example, as described below.
[0314] In some exemplary aspects, the AGU configuration scheme can be configured to swap between two loop dimensions, for example, if determined necessary, for example, when using a time-shifting scheme, for example, as described below.
[0315] In some exemplary aspects, the AGU configuration scheme can be configured to provide a technical solution to support configurations (eg, even any configurations) of memory access operations with different offsets to the same memory pointer, for example, as described below.
[0316] In some exemplary aspects, the AGU configuration scheme may be configured to provide a technical solution to support configurations of load and / or store operations (e.g., even any configuration), such as where the offset of memory access operations to the same memory pointer may vary in up to 2 dimensions, e.g., as described below.
[0317] In other aspects, the AGU configuration scheme may be configured to support configurations for memory access operations where the offset varies in more than 2 dimensions, for example, where this is supported by the AGU.
[0318] In some exemplary aspects, compiler 160 may be configured to identify multiple memory access operations in a loop, for example, based on source code 112 to be compiled into target code 115 to be executed by target processor 180 , for example, as described below.
[0319] In some exemplary aspects, compiler 160 may be configured to generate object code 115 that is configured for execution, for example, by a target vector processor (eg, vector processor 180), for example, as described below.
[0320] In some exemplary aspects, compiler 160 may be configured to generate target code 115 configured for execution by, for example, a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor (eg, processor 180).
[0321] In other aspects, compiler 160 may be configured to generate object code 115 configured for execution by, for example, any other suitable type of processor.
[0322] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on source code 112 including Open Computing Language (OpenCL) code, for example.
[0323] In other aspects, compiler 160 may be configured to generate object code 115 based on source code 112 , including any other suitable type of code, for example.
[0324] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into target code 115 , for example, according to a low-level virtual machine (LLVM)-based compilation scheme.
[0325] In other aspects, compiler 160 may be configured to compile source code 112 into target code 115 according to any other suitable compilation scheme.
[0326] In some exemplary aspects, compiler 160 may be configured to identify a plurality of memory access operations in a loop in source code 112 , for example, where a loop is included in source code 112 .
[0327] In some exemplary aspects, compiler 160 may be configured to identify multiple memory access operations in a loop in code, such as mid-end code or any other code that may be compiled from source code 112 .
[0328] In some exemplary aspects, the plurality of memory access operations may include a plurality of load operations, eg, as described below. In other aspects, the plurality of memory access operations may include any other type of memory access operations.
[0329] In some exemplary aspects, the plurality of memory access operations may include at least a first memory access operation and a second memory access operation to the same memory pointer, eg, as described below.
[0330] In some exemplary aspects, the first memory access operation can have a first offset, eg, as described below.
[0331] In some exemplary aspects, the second memory access operation can have a second offset that is different than the first offset, eg, as described below.
[0332] In some exemplary aspects, compiler 160 may be configured to configure AGU configuration code to configure multiple memory access operations performed by the same AGU, eg, as described below.
[0333] In some exemplary aspects, compiler 160 may be configured to generate object code 115 , for example, based on compiling source code 112 , for example, as described below.
[0334] In some exemplary aspects, compiler 160 may be configured to generate target code 115 that will be based on, for example, AGU configuration code, e.g., as described below.
[0335] In some exemplary aspects, target code 115 may include some or all of the AGU configuration code.
[0336] In other aspects, target code 115 may be generated, for example, based on further processing and / or compilation of the AGU configuration code.
[0337] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to configure a plurality of memory access operations, for example, based on a dimension of a loop, eg, as described below.
[0338] In some demonstrative aspects, the first offset may comprise a first integer multiple of a dimension of the cycle, eg, as described below.
[0339] In some demonstrative aspects, the second offset may comprise a second integer multiple of the dimension of the cycle, eg, as described below.
[0340] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to configure a plurality of memory access operations, for example, based on a difference between a first offset and a second offset, eg, as described below.
[0341] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to configure multiple memory access operations based on the difference between the first offset and the second offset, e.g., based on a determination that the difference between the first offset and the second offset is independent of the dimension of the loop, e.g., as described below.
[0342] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to configure a first dimension of the AGU, for example, based on a dimension of a loop, eg, as described below.
[0343] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to configure the first dimension of the AGU based on the dimension of the loop, for example, based on determining that a difference between the first offset and the second offset is based on an integer multiple of the dimension of the loop and the difference between the first offset and the second offset is based on a dimension-independent value that may be independent of the dimension of the loop, e.g., as described below.
[0344] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to configure the second dimension of the AGU, eg, based on an irrelevant value, eg, as described below.
[0345] In some exemplary aspects, the multiple memory access operations may include a third memory access operation to the same memory pointer, eg, as described below.
[0346] In some exemplary aspects, the third memory access operation can have a third offset that is different from the first offset and the second offset, eg, as described below.
[0347] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to configure multiple memory access operations by the same AGU, for example, based on a third offset, such as described below.
[0348] In some exemplary aspects, the loop can include a first loop, eg, as described below.
[0349] In some exemplary aspects, a first loop may be nested within a second loop, eg, as described below.
[0350] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code, for example, based on the second cycle, eg, as described below.
[0351] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to, for example, include a first dimension code to configure a first dimension of the AGU, for example, based on a first cycle, eg, as described below.
[0352] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to, for example, include a second dimension code to configure a second dimension of the AGU, for example, based on a plurality of memory access operations, eg, as described below.
[0353] In some demonstrative aspects, compiler 160 may be configured to configure the AGU configuration code to, for example, include a third dimension code to configure a third dimension of the AGU, for example, based on the second cycle, for example, as described below.
[0354] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to, for example, include a loop swap instruction, which may be configured to indicate a loop swap between a dimension of the AGU corresponding to the first loop and a dimension of the AGU corresponding to the second loop, e.g., as described below.
[0355] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include a loop swap instruction, eg, based on determining a difference between the first offset and the second offset based on a dimension of the second loop, eg, as described below.
[0356] In some exemplary aspects, the AGU configuration code may include code to configure the same AGU, e.g., according to an extra-dimensional configuration scheme, e.g., as described below.
[0357] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU, for example, according to an extra dimension configuration scheme, e.g., as described below.
[0358] In some exemplary aspects, the additional dimension configuration scheme can be configured to configure additional dimensions of the AGU, for example, based on a plurality of memory access operations, eg, as described below.
[0359] In some exemplary aspects, the AGU configuration code may include, for example, first dimension code to configure a first dimension of the AGU, for example based on a cycle, eg, as described below.
[0360] In some exemplary aspects, the AGU configuration code may include second dimension code to configure an additional dimension of the AGU (eg, a second dimension of the AGU), eg, based on a plurality of memory access operations, eg, as described below.
[0361] For example, the first dimension code may configure multiple first dimensions of the AGU, e.g., including an x dimension and / or a y dimension, e.g., based on a cycle; and / or the second dimension code may configure a second (e.g., additional) dimension of the AGU, e.g., including a z dimension, e.g., based on multiple memory access operations, e.g., as described below.
[0362] In some exemplary aspects, compiler 160 may be configured to configure the first dimension code to set a count parameter of a first dimension of the AGU, for example, based on a count parameter of a loop, for example, as described below.
[0363] In some exemplary aspects, compiler 160 may be configured to configure the first dimension code to set a stride parameter of the first dimension of the AGU, for example, based on a stride parameter of a loop, for example, as described below.
[0364] In some exemplary aspects, compiler 160 may be configured to configure the second dimension code to set a count parameter of a second dimension of the AGU, for example, based on a total count of memory access operations in a plurality of memory access operations, for example, as described below.
[0365] In some exemplary aspects, compiler 160 may be configured to configure the second dimension code to set a stride parameter of the second dimension of the AGU, for example, based on a difference between the first offset and the second offset, for example, as described below.
[0366] In some exemplary aspects, compiler 160 may be configured to generate loop code, for example, based on the loop, eg, as described below.
[0367] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on loop code, for example, as described below.
[0368] In some exemplary aspects, loop code may include multiple memory access instructions to be executed by, for example, the same AGU, eg, as described below.
[0369] In some exemplary aspects, the plurality of memory access instructions may respectively correspond to a plurality of memory access operations, eg, as described below.
[0370] In some exemplary aspects, the multiple memory access operations may include a third memory access operation to the same memory pointer, eg, as described below.
[0371] In some exemplary aspects, the third memory access operation can have a third offset that is different than the first offset and the second offset, eg, as described below.
[0372] In some exemplary aspects, compiler 160 may be configured to generate loop code including first, second, and third memory access instructions to be executed by the same AGU, eg, based on the first, second, and third memory access instructions, respectively, eg, as described below.
[0373] In some exemplary aspects, target code 115 may be based on, for example, a loop code including first, second, and third memory access instructions to be executed by the same AGU, eg, as described below.
[0374] In some exemplary aspects, compiler 160 may be configured to recognize that, for example, a first difference between a first offset and a second offset is equal to a second difference between the second offset and a third offset, eg, as described below.
[0375] In some demonstrative aspects, compiler 160 may be configured to configure the second dimension of the AGU, eg, based on the second difference value, eg, based on a determination that the first difference value is equal to the second difference value, eg, as described below.
[0376] In some exemplary aspects, the loop can include a first loop, eg, an inner loop, eg, as described below.
[0377] In some exemplary aspects, a first loop can be nested within a second loop, eg, an outer loop, eg, as described below.
[0378] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include a third dimension code to configure a third dimension of the AGU, eg, based on the second cycle, eg, as described below.
[0379] In some exemplary aspects, compiler 160 may be configured to selectively configure the AGU configuration code based on a plurality of memory access operations, eg, based on one or more conditions, eg, as described below.
[0380] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code based on multiple memory access operations, for example, based on a determination that the dimension count of the first loop is less than the dimension count supported by the same AGU, eg, as described below.
[0381] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code based on multiple memory access operations, for example, based on a determination that the multiple memory access operations are unbounded and have no fall-through, eg, as described below.
[0382] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code based on multiple memory access operations, for example, based on a determination that differences between consecutive offsets of the multiple memory access operations are the same, eg, as described below.
[0383] In some exemplary aspects, the compiler 160 may be configured to configure the AGU configuration code based on multiple memory access operations, for example, based on a determination that: a dimension count of a first loop is less than a dimension count supported by the same AGU, the multiple memory access operations have no bounds and no fall-through, and / or the differences between consecutive offsets of the multiple memory access operations are the same, for example, as described below.
[0384] In other aspects, compiler 160 may be configured to configure the AGU configuration code based on a plurality of memory access operations, for example, based on any other additional and / or alternative conditions and / or criteria.
[0385] In some exemplary aspects, the AGU configuration code may include code to configure the same AGU, e.g., according to a time-shifting scheme, e.g., as described below.
[0386] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU, for example, according to a time-shifting scheme, e.g., as described below.
[0387] In some exemplary aspects, the time-shifting scheme can be configured to implement multiple memory access operations as multiple time-shifted memory access operations of the same AGU, eg, during multiple corresponding loop iterations, eg, as described below.
[0388] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include loop swap instructions for a dimension of the AGU corresponding to the loop, eg, as described below.
[0389] In some exemplary aspects, a loop swap instruction may be configured to indicate a loop swap between, for example, a dimension of an AGU corresponding to a loop and another dimension of an AGU corresponding to an outer loop, eg, as described below.
[0390] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include a loop swap instruction for the dimension of the AGU corresponding to the loop, e.g., based on a determination of a difference between the first offset and the second offset, e.g., based on a dimension of another loop (e.g., an outer loop including a loop including a plurality of memory access operations), e.g., as described below.
[0391] In some exemplary aspects, compiler 160 may be configured to configure a loop swap instruction to indicate a loop swap between a dimension of an AGU corresponding to a loop and another dimension of an AGU corresponding to another loop (e.g., an outer loop), e.g., as described below.
[0392] In some exemplary aspects, compiler 160 may be configured to selectively utilize loop swap instructions based on one or more conditions, eg, as described below.
[0393] In some exemplary aspects, compiler 160 may be configured to selectively utilize loop swap instructions, for example, based on a determination that a plurality of memory access operations are not in an inner loop, eg, as described below.
[0394] In some exemplary aspects, compiler 160 may be configured to selectively utilize loop swap instructions, for example, based on whether a difference between offsets for multiple memory access operations is based on a determination of a dimension of an inner loop, eg, as described below.
[0395] In some exemplary aspects, compiler 160 may be configured to utilize loop swap instructions based on a determination of a dimension of another loop (eg, an outer loop), for example, based on a difference between offsets for a plurality of memory access operations.
[0396] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to set the same count parameter for all AGUs to perform memory access operations, for example, during a cycle, for example, according to a time-shifting scheme, for example, as described below.
[0397] In some exemplary aspects, the same count parameter can be based on, for example, a count parameter of a cycle, a maximum offset of a plurality of memory access operations, and / or a minimum offset of a plurality of memory access operations, eg, as described below.
[0398] In other aspects, the same counting parameters may be configured based on any other additional or alternative parameters and / or criteria.
[0399] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code, for example, according to a time-shifting scheme, for example, to set the same count parameter, denoted as Count, for example, as follows:
[0400] Count=AlignUpTo((MaxXOffset–MinXOffset)+
[0401] Count(Loop),VF) / VF
[0402] Where MaxXOffset represents the maximum offset of multiple memory access operations.
[0403] Where MinXOffst represents the minimum offset of multiple memory access operations.
[0404] Where VF represents the vectorization factor,
[0405] Where Count(Loop) represents the counting parameter of the loop, and
[0406] The operation AlignUpTo(a,b) can align a upward to an integer multiple of b.
[0407] In some exemplary aspects, compiler 160 may be configured to generate loop code based on the loop, and pre-header code to be executed before a header of the loop code, eg, as described below.
[0408] In some exemplary aspects, compiler 160 may be configured to generate pre-header code, for example, when generating AGU configuration code according to a time-shifting scheme, eg, as described below.
[0409] For example, pre-header code may be generated for all AGU-based implementations of any of the loops, eg, to set AGU parameters, eg, when generating AGU configuration code according to a time-shifting scheme, eg, as described below.
[0410] In some exemplary aspects, target code 115 may be based on loop code and / or pre-header code, eg, as described below.
[0411] In some exemplary aspects, the pre-header code may include multiple AGU load instructions to be executed by the same AGU, eg, as described below.
[0412] In some exemplary aspects, the count of AGU load instructions in the plurality of AGU load instructions may be based on, for example, a count of memory access operations in the plurality of memory access operations, eg, as described below.
[0413] In some exemplary aspects, a count of AGU load instructions in the plurality of AGU load instructions may, for example, be equal to a count of memory access operations in the plurality of memory access operations, eg, as described below.
[0414] In other aspects, any other suitable count of AGU load instructions may be utilized.
[0415] In some exemplary aspects, compiler 160 may configure loop code to include multiple store instructions that may be configured to store results of corresponding multiple AGU load iterations performed by the same AGU, eg, as described below.
[0416] In some exemplary aspects, the count of the plurality of store instructions may be based, for example, on a count of memory access operations in the plurality of memory access operations, eg, as described below.
[0417] In some exemplary aspects, the count of the plurality of store instructions may, for example, be equal to the count of memory access operations in the plurality of memory access operations, eg, as described below.
[0418] In other aspects, any other suitable count of store instructions may be utilized.
[0419] In some exemplary aspects, compiler 160 may be configured to generate multiple storage instructions that may be configured to cause target processor 180 to store multiple load results of multiple AGU load instructions in pre-header code, for example, at the first iteration in the order of multiple AGU load iterations, e.g., as described below.
[0420] In some exemplary aspects, compiler 160 may be configured to selectively configure the AGU configuration code to include code to configure the same AGU according to a time-shifting scheme, eg, based on one or more conditions, eg, as described below.
[0421] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU according to a time-shifting scheme, eg, based on a determination that the plurality of memory access operations include only load operations, eg, as described below.
[0422] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU according to a time-shifting scheme, e.g., based on a determination that a difference between successive offsets of a plurality of memory access operations is independent of the dimensions of the loop, e.g., as described below.
[0423] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU according to a time-shifting scheme, e.g., based on a determination that multiple memory access operations are unbounded and / or do not fall through, e.g., as described below.
[0424] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU according to a time-shifting scheme, eg, based on a determination that multiple memory access operations are not masked, eg, as described below.
[0425] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU according to a time-shifting scheme, eg, based on a determination that the loop does not include any side effects, eg, as described below.
[0426] In some exemplary aspects, the compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU according to a time-shifting scheme, for example based on a determination that: the multiple memory access operations include only load operations, the difference between consecutive offsets of the multiple memory access operations is independent of the dimensions of the loop, the multiple memory access operations have no bounds and no fall-through, the multiple memory access operations are not masked, and / or the loop does not include any side effects, for example, as described below.
[0427] In other aspects, compiler 160 may be configured to configure the AGU configuration code to include code to configure the same AGU according to a time-shifting scheme, for example, based on any other additional and / or alternative conditions and / or criteria.
[0428] In some exemplary aspects, compiler 160 may be configured to compile a loop nest, such as the loop of Example 4, based on source code 112, such as according to an AGU configuration scheme, for example, as described below.
[0429] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code for a loop (e.g., the loop of Example 4), for example, to provide a technical solution to support memory access operations of Example 4, for example, using a reduced number of AGUs, for example, as described below.
[0430] In some exemplary aspects, compiler 160 may be configured to compile source code 112 , for example, according to an extra-dimensional configuration scheme, for example, as described below.
[0431] In some demonstrative aspects, compiler 160 may be configured to determine whether to compile source code 112 according to an extra dimension configuration scheme, for example, based on one or more predefined conditions, eg, as described below.
[0432] In some demonstrative aspects, the one or more predefined conditions may include a first condition, which may be based on, for example, a dimension count of a loop nest.
[0433] In some exemplary aspects, the compiler 160 may be configured to determine whether a loop nest based on the source code 112 can be compiled according to an additional dimension configuration scheme, for example based on conditions related to a dimension count of the loop nest and a supported dimension count that can be supported by an AGU (e.g., each AGU) of a target processor (e.g., a vector processor) used to execute the target code 115, for example, as described below.
[0434] In some exemplary aspects, compiler 160 may be configured to determine that a loop nest based on source code 112 may be compiled according to an extra dimension configuration scheme, e.g., based on a determination that a dimension count of the loop nest is less than a dimension count supported by a target processor 180 used to execute target code 115, e.g., as described below.
[0435] In one example, an AGU of a processor (e.g., a vector processor) may support four dimensions. According to this example, compiler 160 may be configured to determine whether a loop nest based on source code 112 may be compiled according to an additional dimension configuration scheme, for example, based on a determination of whether the dimension count of the loop nest is less than 4 (e.g., #LoopDimensions<=3).
[0436] In some exemplary aspects, the one or more predefined conditions may include a second condition, which may be based on, for example, a difference between offsets of memory access operations to the same memory pointer, eg, as described below.
[0437] In some exemplary aspects, compiler 160 may be configured to determine whether a loop nest based on source code 112 can be compiled according to an extra dimension configuration scheme, e.g., based on a determination of whether all differences between offsets of each two consecutive memory access operations to the same pointer are the same, e.g., as described below.
[0438] In one example, two consecutive memory access operations may include memory access operations that are consecutive in terms of offset, such as when the memory access operations are ordered by their offsets.
[0439] In some exemplary aspects, compiler 160 may be configured to determine that a loop nest based on source code 112 may be compiled according to an extra dimension configuration scheme, for example based on a determination that a difference between offsets of consecutive memory access operations (eg, every two consecutive memory access operations) is the same.
[0440] In some exemplary aspects, the one or more predefined conditions may include a third condition, which may be based on, for example, bounds on memory access operations to the same memory pointer.
[0441] In some exemplary aspects, compiler 160 may be configured to determine, for example, based on a determination that any boundary conditions (bounds) (e.g., with any fallthrough) are the same for all memory access operations to the same memory pointer, that a loop nest based on source code 112 may be compiled according to an extra dimension configuration scheme, e.g., as described below.
[0442] In some exemplary aspects, some or all of the conditions may be implemented to determine whether an extra-dimensional configuration scheme is applicable to a loop nest, and / or any other additional conditions, rules, and / or criteria may be implemented to determine whether an extra-dimensional configuration scheme is applicable to a loop nest.
[0443] In some exemplary aspects, compiler 160 may, for example, based on the loop nesting of Example 4, identify three load operations with different offsets to the same memory pointer inp1, for example, load operations inp1[index], inp1[index+width], and inp1[index+2*width].
[0444] In some exemplary aspects, compiler 160 may be configured to verify that three load operations to the same memory pointer have the same step size in all loops of the loop nest of Example 4, for example. For example, compiler 160 may verify that all three load operations inp1[index], inp1[index+width], and inp1[index+2*width] have the same step size in inner loop X, e.g., Step(X-Loop)=VectorSize, and have the same step size in outer loop Y, e.g., Step(Y-Loop)=width.
[0445] In some exemplary aspects, compiler 160 may be configured to recognize that the offsets of the three load operations inp1[index], inp1[index+width], and inp1[index+2*width] are different in the Y dimension. For example, compiler 160 may be configured to recognize that the load operation inp1[index] has an offset of 0(*width) in the Y dimension, the load operation inp1[index+width] has an offset of 1(*width) in the Y dimension, and the load operation inp1[index+2*width] has an offset of 2(*width) in the Y dimension.
[0446] In some exemplary aspects, compiler 160 may be configured to select to compile the loop nest of Example 4 using a non-commutative loop scheme, e.g., where outer loop Y remains the outermost loop and inner loop X remains the innermost loop, e.g., as described below.
[0447] In some exemplary aspects, compiler 160 may be configured to select to utilize a non-interchange loop scheme, for example, based on a selection to compile the loop nests of Example 4 according to an extra dimension configuration scheme, for example, as described below.
[0448] In some exemplary aspects, compiler 160 may be configured to verify whether predefined conditions for implementing an extra dimension allocation scheme are satisfied with respect to the loop nest of Example 4, eg, as described below.
[0449] In some exemplary aspects, compiler 160 may be configured to verify that the number of dimensions of the loop nest of Example 4 is less than four, eg, the number of AGU dimensions supported by the AGU.
[0450] In some exemplary aspects, compiler 160 may be configured to verify that the differences between the offsets of the three load operations are the same.
[0451] In some exemplary aspects, compiler 160 may be configured to verify that multiple memory access operations have no bounds and no fall-through, eg, that any bounds of three load operations are the same.
[0452] In some exemplary aspects, compiler 160 may determine that the predefined condition is satisfied with respect to the loop nest of Example 4, for example, by determining that the number of dimensions of a loop (e.g., two loops) satisfies a condition on the number of dimensions (e.g., #loop-dimensions=2<=3).
[0453] In some exemplary aspects, compiler 160 may determine that the predefined condition is satisfied with respect to the loop nest of Example 4, for example, by determining that the differences between the offsets are all equal to 1(*width) (eg, the differences between the offsets are (1-0=2-1=1)).
[0454] In some exemplary aspects, compiler 160 may determine that the predefined condition is satisfied with respect to the loop nest of Example 4, for example, by determining that there are no bounds or fall-throughs in a plurality of memory access operations of the loop nest.
[0455] In some exemplary aspects, compiler 160 may be configured to compile source code 112 , eg, according to an extra dimension configuration scheme, eg, based on a determination that all prerequisites for the extra dimension configuration scheme are satisfied.
[0456] In some exemplary aspects, compiler 160 may be configured to set AGU configuration code for the AGU, for example, according to an extra dimension configuration scheme, e.g., as described below.
[0457] In some exemplary aspects, compiler 160 may be configured to set count parameters for additional dimensions, for example, based on a total count of memory accesses.
[0458] For example, compiler 160 may recognize that the loop nest of Example 4 includes three load operations inp1[index], inp1[index+width], and inp1[index+2*width]. Accordingly, compiler 160 may set the count parameter of the additional dimension to three, for example, Count(Extra-dim)=#AccessesInTheGroup=3.
[0459] In some exemplary aspects, compiler 160 may be configured to set a stride parameter for an additional dimension of the AGU, for example, based on a difference between the offsets.
[0460] For example, compiler 160 may identify that the difference between consecutive offsets of three load operations is equal to 1(*width), e.g., as described above. Accordingly, compiler 160 may set the stride parameter of the additional dimension of the AGU to 1*width, e.g., Step(Extra-dim)=DifferenceBetween2ClosestOffsets*Stride=1*width.
[0461] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code and loop code for the loop nest of Example 4, for example, according to an extra dimension configuration scheme, for example, as follows:
[0462] AGU configuration code:
[0463] agu1=allocate_agu("load");
[0464] set_base(agu1,inp1);
[0465] set_y_minmax(agu1,0,(height–2)*width);
[0466] set_y_count_stride(agu1,height–2,width);
[0467] set_x_minmax(agu1,0,width);
[0468] VectorizationFactor = 32;
[0469] Xcount=AlignUpTo(width,VectorizationFactor) / VectorizationFactor;
[0470] VectorSize=VectorizationFactor*1;
[0471] set_x_count_stride(agu1, Xcount, VectorSize);
[0472] set_z_count_stride(agu1, 3, width);
[0473] agu2 = allocate_agu("load");
[0474] set_base(agu2, inp2);
[0475] set_y_minmax(agu2, 0, (height – 2) * width);
[0476] set_y_count_stride(agu2, height – 2, width);
[0477] set_x_minmax(agu2, 0, width);
[0478] set_x_count_stride(agu2, Xcount, VectorSize);
[0479] agu_out = allocate_agu("store");
[0480] set_base(agu_out, out);
[0481] set_y_minmax(agu_out, 0, (height – 2) * width);
[0482] set_y_count_stride(agu_out, height – 2, width);
[0483] set_x_minmax(agu_out, 0, width);
[0484] set_x_count_stride(agu_out, Xcount, VectorSize);
[0485] Loop:
[0486] %val0 = agu1.load(); agu1.bump();
[0487] %val1 = agu1.load(); agu1.bump();
[0488] %val2=agu1.load();agu1.bump();
[0489] %val_inp2=agu2.load();agu2.bump();
[0490] %temp=mul%val0,%val1;
[0491] %temp1=add%temp,val2;
[0492] %result=sub%temp1,%val_inp2;
[0493] agu_out.store(%result); agu_out.bump();
[0494] br(Loop);
[0495] Example (5)
[0496] In some exemplary aspects, as shown in Example 5, the AGU configuration code for agul may include AGU configuration code for an additional dimension (eg, dimension z).
[0497] In some exemplary aspects, as shown in Example 5, for example, in addition to dimension x that can be configured based on an X cycle and dimension y that can be configured based on a Y cycle, an additional dimension z can be configured.
[0498] In some exemplary aspects, AGU configuration code for agu1 may include, for example, AGU configuration code to set a count parameter of dimension z to three (eg, "3"), eg, set_z_count_stride(agu1,3,width).
[0499] In some exemplary aspects, AGU configuration code for agu1 may include, for example, AGU configuration code to set a stride (step length) parameter of dimension z to 1*width (eg, width), eg, set_z_count_stride(agu1,3,width).
[0500] In some exemplary aspects, as shown in Example 5, the loop code may include three load instructions for the same pointer (eg, agu1), which may correspond to three load operations, for example, as follows:
[0501] %val0=agu1.load();agu1.bump();
[0502] %val1=agu1.load();agu1.bump();
[0503] %val2=agu1.load();agu1.bump();
[0504] For example, the instruction %val0=agu1.load(); agu1.bump() may be configured to store in register %val0 the value loaded by agu1 in the first iteration over dimension z, eg, corresponding to the load operation inp1[index].
[0505] For example, the instruction %val1 = agu1.load(); agu1.bump() may be configured to store in register %val1 the value loaded by agu1 in the second iteration over dimension z, eg, corresponding to a load operation inp1[index+width].
[0506] For example, the instruction %val2=agu1.load(); agu1.bump() may be configured to store in register %val2 the value loaded by agu1 in the third iteration over dimension z, eg, corresponding to the load operation inp1[index+2*width].
[0507] In some exemplary aspects, compiler 160 may be configured to compile source code 112 , for example, according to a time-shifting scheme, eg, as described below.
[0508] In some demonstrative aspects, compiler 160 may be configured to determine whether to compile source code 112 according to a time-shifting scheme, for example, based on one or more predefined conditions, eg, as described below.
[0509] In some exemplary aspects, the one or more predefined conditions may include a first condition that may require that the time-shifting scheme may be applied, for example, only with respect to load operations (eg, not for store operations).
[0510] In some exemplary aspects, the one or more predefined conditions may include a second condition, which may entail including that a dimension of a load operation to be compiled according to the time-shifting scheme (“time-shifted dimension”) should be or may become a dimension of an innermost loop of a loop nest.
[0511] In some exemplary aspects, for example, the time-shifting scheme may include loop swap instructions, such as if possible, such as if the time-shifted dimension is not already a dimension of an innermost loop of a loop nest.
[0512] In some exemplary aspects, the one or more predefined conditions may include a third condition, which may be based on, for example, a difference between offsets of memory access operations to the same memory pointer, eg, as described below.
[0513] In some exemplary aspects, compiler 160 may be configured to determine whether a loop nest based on source code 112 can be compiled according to a time-shifting scheme, for example, based on a determination of whether a difference between offsets of each two (e.g., consecutive and / or non-consecutive) memory access operations to the same pointer is constant (e.g., known at compile time and / or does not change during execution).
[0514] In some exemplary aspects, compiler 160 may be configured to determine that a loop nest based on source code 112 may be compiled according to a time-shifting scheme, for example based on a determination that a difference between offsets of memory accesses is constant (eg, known at compile time).
[0515] In some exemplary aspects, the one or more predefined conditions may include a fourth condition, which may be based on, for example, bounds on memory access operations to the same pointer.
[0516] In some exemplary aspects, compiler 160 may be configured to determine, for example, based on a determination that any boundary conditions (bounds) (e.g., having any fall-through) are adaptable for all memory access operations to the same pointer, that a loop nest based on source code 112 may be compiled according to a time-shifting scheme, e.g., as described below.
[0517] In some demonstrative aspects, the one or more predefined conditions may include a fifth condition related to a side effect of the loop nest.
[0518] For example, compiler 160 may be configured to determine that a time-shifting scheme is applicable to a loop nest, for example based on a determination that the loop nest does not include one or more certain side effects (eg, only if the loop nest does not include one or more certain side effects).
[0519] In some exemplary aspects, the one or more predefined conditions may include a sixth condition related to masking of a load operation in a loop nest.
[0520] For example, compiler 160 may be configured to determine that a time-shifting scheme is applicable to a loop nest, for example, based on a determination that a load operation in the loop nest is not masked (eg, only when a load operation in the loop nest is not masked).
[0521] In some exemplary aspects, some or all of the conditions may be implemented to determine whether a time-shifting scheme is applicable to a loop nest, and / or any other additional conditions, rules, and / or criteria may be implemented to determine whether a time-shifting scheme is applicable to a loop nest.
[0522] In some exemplary aspects, compiler 160 may, for example, based on the loop nesting of Example 4, identify three load operations with different offsets to the same memory pointer inp1, for example, load operations inp1[index], inp1[index+width], and inp1[index+2*width].
[0523] In some exemplary aspects, compiler 160 may be configured to verify that three load operations to the same memory pointer have the same step size in all loops of the loop nest of Example 4. For example, compiler 160 may verify that all three load operations inp1[index], inp1[index+width], and inp1[index+2*width] have the same step size in inner loop X, e.g., Step(X-Loop)=VectorSize, and have the same step size in outer loop Y, e.g., Step(Y-Loop)=width.
[0524] In some exemplary aspects, compiler 160 may be configured to recognize that the offsets of three load operations inp1[index], inp1[index+width], and inp1[index+2*width] are different in the Y dimension. For example, the load operation inp1[index] has an offset of 0(*width) in the Y dimension, the load operation inp1[index+width] has an offset of 1(*width) in the Y dimension, and the load operation inp1[index+2*width] has an offset of 2(*width) in the Y dimension.
[0525] In some exemplary aspects, compiler 160 may be configured to verify whether predefined conditions for implementing the time-shifting scheme are satisfied with respect to the loop nest of Example 4.
[0526] In some exemplary aspects, compiler 160 may be configured to verify that all three memory access operations of Example 4 are load operations.
[0527] In some exemplary aspects, compiler 160 may be configured to verify that the differences between the offsets of the three load operations of Example 4 are constant.
[0528] In some exemplary aspects, compiler 160 may be configured to verify that any bounds of the three load operations of Example 4 are the same.
[0529] In some exemplary aspects, compiler 160 may be configured to verify that the loop nest of Example 4 does not include side effects.
[0530] In some exemplary aspects, compiler 160 may be configured to verify that the three load operations of Example 4 are not masked.
[0531] In some exemplary aspects, compiler 160 may be configured to compile source code 112 , eg, according to the time-shifting scheme, eg, based on a determination that all prerequisites for the time-shifting scheme are satisfied.
[0532] In some exemplary aspects, compiler 160 may be configured to determine whether loop swapping is necessary and / or possible for a loop nest.
[0533] In some exemplary aspects, compiler 160 may determine that loop swapping is needed, for example, based on a determination that a time shift is to be performed in a Y dimension that is not initially a dimension of the innermost loop.
[0534] In some exemplary aspects, compiler 160 may determine that loop swapping is possible, for example, based on a determination that the loop nest comprises a perfect loop nest and that there are no side effects in the loop nest.
[0535] In some exemplary aspects, compiler 160 may be configured to set AGU configuration code for multiple AGUs for a target processor (e.g., processor 180), for example, to set an AGU configuration, such as a vertical time-shift ("ShortColumn") configuration, which may indicate that a time-shifting scheme is to be implemented, for example, as described below.
[0536] In some exemplary aspects, compiler 160 may be configured, for example, to set the same count parameter, denoted as Count(Dim), in a time-shifted dimension (eg, Y dimension) for all of the plurality of AGUs, for example, as described below.
[0537] In some exemplary aspects, compiler 160 may be configured to set the same count parameter for the time-shifted dimension (e.g., Y dimension) based on the count parameter of the loop of the time-shifted dimension, the maximum offset of three load operations to the same memory pointer, and the minimum offset of the three load operations, e.g., as follows:
[0538] Count(Ydim)=(MaxYOffset–MinYOffset)+Count(LoopYDim)
[0539] Where MaxYOffset represents the maximum offset in the time-shifted dimension (eg, Y dimension) among all three load operations covered by the AGU, and where MinYOffst represents the minimum offset in the time-shifted dimension (eg, Y dimension) among all loads covered by the AGU.
[0540] In some exemplary aspects, compiler 160 may be configured to set, for example, all loop swap instructions in the AGU of target processor 180, such as set_y_innerloop(AGU), which is used to indicate loop swapping for the Y dimension, for example, if loop swapping is required.
[0541] In some exemplary aspects, compiler 160 may be configured to generate loop code based on the loop, and pre-header code to be executed before a header of the loop code, eg, as described below.
[0542] In some exemplary aspects, compiler 160 may be configured to configure loop code, for example, to load only one value within the loop through the same AGU utilized to implement multiple load operations. For example, the loop code may include only one load operation, for example, as described below.
[0543] In some exemplary aspects, compiler 160 may be configured to generate pre-header code to load initial values by the same AGU, e.g., before execution of a loop. For example, the pre-header code may include multiple AGU load instructions to be executed by the same AGU, e.g., before execution of a loop, e.g., as described below.
[0544] In some exemplary aspects, compiler 160 may be configured to remember and / or retain (eg, store in registers) other values from previous AGU load iterations in a loop, eg, as described below.
[0545] In some exemplary aspects, the first iteration in sequence of the loop may utilize values loaded by the pre-header code, eg, initial values, eg, as described below.
[0546] In some exemplary aspects, compiler 160 may be configured to configure pre-header code and loop code for a time-shifting scheme, for example, based on load operations performed by the same AGU, e.g., as follows:
[0547] In the pre-header of the loop, add MaxTS loads / bumps: %p_i = agu.load();
[0548] agu.bump(); (0 <= I <MaxTS)
[0549] In the loop, add a phi chain of length MaxTS: %q_i=phi(%p_i:Pre-header,%q_(i+1):
[0550] Loop)(0<=i <MaxTS)
[0551] In the loop, after the phi chain, add a load and a push: %q_MaxTS = agu.load();
[0552] agu.bump();
[0553] Use %q_i (0<=i<=MaxTS) in a loop.
[0554] where MaxTS represents the number of iterations whose results are stored and loaded within the loop, e.g. in order to use the result later.
[0555] For example, the value of MaxTS may be based on a count of iterations to be performed between a load operation corresponding to a minimum offset and a load operation corresponding to a maximum offset of load operations to be performed by the same AGU.
[0556] For example, the value of MaxTS may indicate, for a value used in a current iteration, how many iterations ago the value was loaded from a buffer, such as by a load instruction.
[0557] For example, the value of MaxTS may use the maximum value of all such values.
[0558] For example, the code %val=phi(%a:label1,%b:label2) (e.g., %q_i=phi(%p_i:Pre-header,%q_(i+1):Loop)) may represent a conditional instruction (also called a "conditional specification instruction"), where, for example, %val takes the value %a if the instruction is reached from the code under label1, or, for example, takes the value %b if the instruction is reached from the code under label2.
[0559] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, pre-header code, and loop code for the loop nest of Example 4, for example, according to a time-shifting scheme, for example, as follows:
[0560] AGU configuration code:
[0561] agu1=allocate_agu("load");
[0562] set_base(agu1,inp1);
[0563] set_y_minmax(agu1,0,(height–2)*width);
[0564] set_y_count_stride(agu1,height,width);
[0565] set_y_innerloop(agu1); / / Loop interchange
[0566] set_x_minmax(agu1,0,width);
[0567] VectorizationFactor = 32;
[0568] VectorSize = VectorizationFactor * 1;
[0569] xCount = AlignUpTo(width,VectorizationFactor) / VectorizationFactor;
[0570] set_x_count_stride(agu1,xCount,VectorSize);
[0571] agu2 = allocate_agu(“load”);
[0572] set_base(agu2,inp2);
[0573] set_y_minmax(agu2,0,(height–2)*width);
[0574] set_y_count_stride(agu2,height,width);
[0575] set_y_innerloop(agu2); / / Loop interchange
[0576] set_x_minmax(agu2,0,width);
[0577] set_x_count_stride(agu2,xCount,VectorSize);
[0578] agu_out = allocate_agu(“store”);
[0579] set_base(agu_out,out);
[0580] set_y_minmax(agu_out,0,(height–2)*width);
[0581] set_y_count_stride(agu_out,height,width);
[0582] set_y_innerloop(agu_out); / / loop exchange
[0583] set_x_minmax(agu_out,0,width);
[0584] set_x_count_stride(agu_out,xCount,VectorSize);
[0585] Pre-header:
[0586] %p0=agu1.load();agu1.bump();
[0587] %p1=agu1.load();agu1.bump();
[0588] cycle:
[0589] %q0=phi(%p0:Pre-header,%q1:Loop);
[0590] %q1=phi(%p1:Pre-header,%q2:Loop);
[0591] %q2=agu1.load();agu1.bump();
[0592] %val_inp2=agu2.load();agu2.bump();
[0593] %temp=mul%q0,q1;
[0594] %temp1=add temp,%q2;
[0595] %result=sub%temp1,%val_inp2;
[0596] agu_out.store(%result); agu_out.bump();
[0597] br(Loop);
[0598] Example (6)
[0599] In some exemplary aspects, as shown in Example 6, the AGU configuration code may include, for example, loop swap instructions for the Y dimension for all of the AGUs.
[0600] For example, as shown in Example 6, the AGU configuration code may include a loop exchange instruction set_y_innerloop(agu1) for agu1, a loop exchange instruction set_y_innerloop(agu2) for agu2, and a loop exchange instruction set_y_innerloop(agu_out) for agu_out.
[0601] In some exemplary aspects, as shown in Example 6, the AGU configuration code may include the same count parameter in the Y dimension for all of the multiple AGUs.
[0602] For example, as shown in Example 6, the AGU configuration code may include instructions set_y_count_stride(agu1, height, width) for agu1, instructions set_y_count_stride(agu2, height, width) for agu2, and instructions set_y_count_stride(agu_out, height, width) for agu_out.
[0603] For example, as shown in Example 6, the same count parameter in the Y dimension may be determined based on the count parameter of the Y cycle, the maximum offset of the three load operations, and the minimum offset of the three load operations, for example, as follows:
[0604] Count(yDim)=(MaxYOffset–MinYOffset)+Count(LoopYDim)=
[0605] =(2-0)+(height-2)=height
[0606] In some exemplary aspects, the count of pre-head AGU load instructions may be based on a value of MaxTS, for example, eg, MaxTS=2.
[0607] For example, as shown in Example 6, the pre-header code may include two pre-header AGU load instructions, such as %p0=agu1.load(); agu1.bump(); and %p1=agu1.load(); agu1.bump(), which will be executed by the same agu1.
[0608] In some exemplary aspects, as shown in Example 6, two pre-header AGU load instructions (eg, %p0=agul.load(); agul.bump(); and %p1=agul.load(); agul.bump()) may be used to load initial values for an AGU load iteration.
[0609] In some exemplary aspects, as shown in Example 6, the loop code may include two conditional designation instructions, e.g., %q0=phi(%p0:Pre-header, %q1:Loop); %q1=phi(%p1:Pre-header, %q2:Loop), for example, to designate (store) the result of an AGU load iteration performed by the same agu1 to a register based on a phi operation.
[0610] In some exemplary aspects, as shown in Example 6, the loop code may include a single specify (store) instruction, e.g., %q2=agu1.load(); agu1.bump(), e.g., to load and specify (store) the result of the current AGU load iteration by the same agu1 into a register.
[0611] For example, according to the code of Example 6, the pre-header instruction %p0=agu1.load(); agu1.bump() can be configured to assign (store) the first value loaded by agu1 to register %p0, for example, corresponding to the load operation inp1[index].
[0612] For example, according to the code of Example 6, the pre-header instruction %p1=agu1.load(); agu1.bump() can be configured to assign (store) the second value loaded by agu1 to register %p1, for example, corresponding to the load operation inp1[index+width].
[0613] For example, according to the code of Example 6, the loop instruction %q0=phi(%p0:Pre-header, %q1:Loop) can be configured to specify (store) the first value loaded by agu1 to register %q0, for example, corresponding to the load operation inp1[index], for example, in the first iteration in the order of the loop; or the value from register %q1, for example, corresponding to the load operation inp1[index], for example, in a non-first iteration in the order of the loop, for example, it is inp1[index+width] relative to the previous iteration.
[0614] For example, according to the code of Example 6, the loop instruction %q1=phi(%p1:Pre-header, %q2:Loop) can be configured to specify (store) the second value loaded by agu1 to register %q1, for example, corresponding to the load operation inp1[index+width], for example, in the first iteration in the order of the loop; or the value from register %q2, for example, corresponding to the load operation inp1[index+width], for example, in an iteration other than the first in the order of the loop, for example, it is inp1[index+2*width] relative to the previous iteration.
[0615] For example, according to the code of Example 6, the loop instruction %q2=agu1.load(); agu1.bump() can be configured to specify (store) the value loaded by agu1 in the current loop to register %q2, for example, corresponding to the load operation inp1[index+2*width].
[0616] In some exemplary aspects, compiler 160 may be configured to compile source code 112 of another executed program to be executed on a vector processor that includes multiple AGUs, eg, as described below. For example, a vector processor may include four AGUs.
[0617] For example, compiler 160 may identify a loop nest including five memory access operations based on source code 112, e.g., as follows:
[0618]
[0619]
[0620] Example (7)
[0621] For example, as shown in Example 7, the loop nest may include an outer loop (y loop) along a dimension (Y dimension) based on a variable height and an inner loop (x loop) along a dimension (X dimension) based on a variable width.
[0622] For example, as shown in Example 7, the loop nest may include two load operations to the same memory pointer denoted as inp1, the two load operations having different two-dimensional (2D) offsets, eg, different offsets in two dimensions.
[0623] For example, the loop nest may include a first load operation (e.g., inp1[index]) on the memory pointer inp1, which has a first 2D offset (e.g., index), which may have an offset of 0 in the Y dimension and an offset of 0 in the X dimension.
[0624] For example, the loop nest may include a second load operation on the memory pointer inp1 (e.g., inp1[index+2*width+3]), which has a second 2D offset (e.g., index+2*width+3), which may have an offset of 2(*width) in the Y dimension and an offset of 3 in the X dimension.
[0625] refer to Figure 5, which schematically illustrates a memory access scheme 500 to load data from a memory relative to a memory pointer, according to some exemplary aspects.
[0626] In one example, the memory access scheme 500 can demonstrate memory access relative to the same memory pointer denoted as inp1, such as the load operations of inp1[index] and inp1[index+2*width+3] according to Example 7.
[0627] like Figure 5 As shown, the two load operations of inp1[index] and inp1[index+2*width+3] of Example 7 may load data with different 2D offsets relative to the same memory pointer inp1.
[0628] like Figure 5 As shown, the first load operation (eg, inp1[index]) may have an offset of 0 in the Y dimension and an offset of 0 in the X dimension.
[0629] like Figure 5 As shown, the second load operation (eg, inp1[index+2*width+3]) may have an offset of 2(*width) in the Y dimension and an offset of 3 in the X dimension.
[0630] For example, an implementation that requires using a different AGU for each memory access with a different offset may require, for example, a total of 5 AGUs to process the memory access operation of instance 7. For example, a first AGU may be used to process the storage operation out1[index], a second AGU may be used to process the storage operation out2[index], a third AGU may be used to process the load operation inp2[index], and two additional AGUs may be used to process two load operations relative to the memory pointer inp1, for example, inp1[index] and inp1[index+2*width+3]. According to this example, for example, if the processor's available AGU count is less than 5, the processor may not be able to process the memory access operation of instance 7.
[0631] Return to reference Figure 1 In some exemplary aspects, compiler 160 may be configured to configure the loop nest of Example 7, for example, according to a combination of an extra dimension configuration scheme and a time shifting scheme, for example, as described below.
[0632] In some exemplary aspects, an extra dimension configuration scheme may be utilized to configure loop nests, for example, based on an offset in the Y dimension (eg, 2*width), eg, as described below.
[0633] In some exemplary aspects, a time-shifting scheme can be utilized to configure loop nests, for example, based on an offset in the X dimension (eg, +3), eg, as described below.
[0634] In some exemplary aspects, compiler 160 may, for example, based on the loop nesting of Example 7, identify two load operations with different 2D offsets to the same memory pointer inp1, eg, load operations inp1[index] and inp1[index+2*width+3].
[0635] In some exemplary aspects, compiler 160 may be configured to verify that two load operations to the same memory pointer have the same amplitude in all loops of the loop nest of Example 7.
[0636] For example, compiler 160 may verify that all two load operations inp1[index] and inp1[index+2*width+3] have the same step size in inner loop X (e.g., Step(X-Loop)=VectorSize) and have the same step size in outer loop Y (e.g., Step(Y-Loop)=width).
[0637] In some exemplary aspects, compiler 160 may be configured to recognize that the offsets of operations inp1[index] and inp1[index+2*width+3] are different in the Y dimension and the X dimension.
[0638] For example, the load operation inp1[index] may have an offset of 0(*width) in the Y dimension and an offset of 0 in the X dimension; and the load operation inp1[index+2*width+3] may have an offset of 2(*width) in the Y dimension and an offset of 3 in the X dimension.
[0639] In some exemplary aspects, compiler 160 may be configured to verify whether a predefined condition for the time-shifting scheme is satisfied with respect to the X dimension of the loop nest of Example 7.
[0640] In some exemplary aspects, compiler 160 may determine that a predefined condition for a time-shifting scheme is satisfied with respect to the X dimension of the loop nest of Example 7.
[0641] For example, compiler 160 may determine that both memory access operations are load operations; the difference between offsets for all load operations is constant, e.g., 3; any bounds for both load operations are the same; the loop nest does not include side effects; and / or both load operations are not masked.
[0642] In some exemplary aspects, compiler 160 may be configured to verify whether a predefined condition for an additional dimension configuration scheme is satisfied with respect to the Y dimension of the loop nest of Example 7.
[0643] In some exemplary aspects, compiler 160 may determine that the predefined condition for the extra dimension configuration scheme is satisfied with respect to the Y dimension of the loop nest of Example 7.
[0644] For example, compiler 160 may determine that the number of dimensions of the loop (e.g., two dimensions) satisfies conditions on the number of dimensions, e.g., #loop-dimensions=2<=3; the differences between offsets are all equal, e.g., because there are only two load operations; and / or there are no bounds or fallthroughs.
[0645] In some exemplary aspects, compiler 160 may be configured to set AGU configuration code for multiple AGUs of a target processor, for example, to set an AGU configuration, such as a vertical-horizontal shift ("NarrowRectangle") configuration, for example, to indicate that a combination of a horizontal time-shifting scheme and a vertical extra dimension configuration scheme may be implemented, for example, as described below.
[0646] In some exemplary aspects, compiler 160 may be configured to set AGU configuration code for the AGU, for example, according to an extra dimension configuration scheme, e.g., as described below.
[0647] In some exemplary aspects, compiler 160 may be configured to set an access count parameter, denoted C, corresponding to an additional dimension, for example based on a total count of memory accesses to which the extra dimension configuration scheme is to be applied.
[0648] For example, compiler 160 may recognize that the loop nest of Example 7 includes two load operations inp1[index] and inp1[index+2*width+3]. Accordingly, compiler 160 may set the count parameter C to two, for example, C=Count(Extra-dim)=#AccessesInTheGroup=2.
[0649] In some exemplary aspects, compiler 160 may be configured to set a stride parameter of an additional dimension of the AGU, for example, based on a difference between consecutive offsets of memory accesses to which the extra dimension configuration scheme is to be applied.
[0650] For example, compiler 160 may recognize that the difference between two "consecutive" offsets of two load operations is equal to 2(*width), e.g., as described above. Accordingly, compiler 160 may set the stride parameter of the additional dimension of the AGU to 2*width, e.g., Step(Extra-dim)=DifferenceBetween2ClosestOffsets*Stride=2*width.
[0651] In some exemplary aspects, compiler 160 may be configured to set AGU configuration code for multiple AGUs, for example, according to a time-shifting scheme, eg, as described below.
[0652] In some exemplary aspects, compiler 160 may be configured, for example, to set the same count parameter on a time-shifted dimension (eg, X dimension), denoted as Count(xDim), for all in the AGU, for example, as described below.
[0653] In some exemplary aspects, compiler 160 may be configured to set the same count parameter for the time-shifted dimension based on, for example, a count parameter of a loop of the time-shifted dimension, a maximum offset of a load operation to the same memory pointer, and a minimum offset of a load operation, for example, as follows:
[0654] Count(xDim)=AlignUpTo((MaxXOffset–MinXOffset)+
[0655] Count(LoopXDim),VF) / VF
[0656] where MaxXOffset represents the maximum offset of a memory access operation in the time-shifted dimension (e.g., X dimension) among all of the load operations covered by the AGU,
[0657] where MinXOffst represents the minimum offset of a memory access operation in the time-shifted dimension (e.g., X dimension) among all of the load operations covered by the AGU,
[0658] Where VF represents the vectorization factor (VectorizationFactor), and
[0659] The operation AlignUpTo(a,b) can be configured to align a upward to an integer multiple of b.
[0660] For example, VF may be similar to the VectorSize parameter. For example, VF may indicate how many elements each vector includes. For example, an element may be a char, a short (two bytes), an int, etc.
[0661] For example, VF may be based on the product of the number of elements and the element size (bytes), for example, VectorSize = numberof elements*element size (bytes) = VF*element size. For example, when the element size is 1, VF may be the same as VectorSize.
[0662] For example, the operation alignup(a,b) may include aligning the value of a upward to the nearest integer multiple of b. For example, the operation alignup(17,5) may produce the value 20, e.g., the smallest integer greater than or equal to 17 that is an integer multiple of 5. For example, the operation AlignUpTo((MaxXOffset–MinXOffset)+Count(LoopXDim), VF) / VF may be utilized to ensure that division by VF results in an integer value, e.g., to ensure that the count value Count(xDim) is an integer.
[0663] In some exemplary aspects, compiler 160 may be configured to generate loop code, for example based on the loop nest of Example 7, and pre-header code to be executed before a header of the loop code.
[0664] In some exemplary aspects, compiler 160 may be configured to configure loop code to load only current values, e.g., C current values, within the loop, e.g., based on a total count of memory accesses by the same AGU utilized to implement multiple load operations. For example, the loop code may include only two load operations, e.g., as described below.
[0665] In some exemplary aspects, compiler 160 may be configured to generate pre-header code to load initial values, for example, prior to execution of a loop, through the same AGU that is utilized to implement multiple load operations.
[0666] For example, the pre-header code may include multiple AGU load instructions to be executed by the same AGU that is utilized to implement multiple load operations, such as prior to execution of a loop, for example, as described below.
[0667] In some exemplary aspects, compiler 160 may be configured to remember and / or retain (eg, assign (store) to registers) other values from previous AGU load iterations in the loop, eg, as described below.
[0668] In some exemplary aspects, the first iteration in sequence of the loop may utilize values loaded by the pre-header code, eg, initial values, eg, as described below.
[0669] In some exemplary aspects, compiler 160 may be configured to shuffle between initial and current values, e.g., in an iteration that is first in order in a loop, and shuffle adjacent values of two consecutive iterations, e.g., when needed, e.g., for all other iterations.
[0670] In some exemplary aspects, compiler 160 may be configured to configure pre-header code and loop code for a time-shifting scheme, for example, based on load operations performed by the same AGU, e.g., as follows:
[0671] In the pre-header of the loop, add MaxTS*C loads / bumps: % p_i = agu.load(); agu.bump();
[0672] (0<=i <MaxTS*C)
[0673] In the loop, add a phi chain of length MaxTS*C: %q_i=phi(%p_i:Pre-header,
[0674] %q_(i+C):Loop)(0<=i <MaxTS*C)
[0675] In the loop, after the phi chain, add C loads / bumps: %q_i = agu.load(); agu.bump();
[0676] (MaxTS*C<=i<(MaxTS+1)*C)
[0677] Use shuffle(%q_i,%q_(i+C),Elements)(0<=i <MaxTS*C)。
[0678] where MaxTS*C represents the product of the number of iterations multiplied by the number of loads in each iteration, the result of which is stored and loaded within the loop, for example, to use the result later, and where the code shuffle(%q_i,%q_(i+C),Elements) (e.g., %val1=shuffle(%q1,%q3,[3,4,5,…,33,34])) can be configured to shuffle between adjacent results of AGU load iterations.
[0679] For example, the value of MaxTS*C may indicate, for the value used in the current iteration, how many iterations ago the value was loaded from the buffer, for example, by a load instruction. For example, the value of MaxTS*C may use the maximum value of all such values.
[0680] In some exemplary aspects, compiler 160 may be configured to generate AGU configuration code, pre-header code, and loop code for the loop nest of Example 7, for example, according to a vertical-horizontal shifting scheme, for example, as follows:
[0681] AGU configuration code:
[0682] agu1=allocate_agu("load");
[0683] set_base(agu1,inp1);
[0684] set_y_minmax(agu1,0,(height–2)*width);
[0685] set_y_count_stride(agu1,height-2,width);
[0686] set_x_minmax(agu1,0,width);
[0687] VF=VectorizationFactor=32;
[0688] VectorSize = VF * 1;
[0689] XCount=AlignUpTo((width–3)+3,VF) / VF;
[0690] set_x_count_stride(agu1,XCount,VectorSize);
[0691] set_z_count_stride(agu1,2,2*width)
[0692] agu2=allocate_agu("load");
[0693] set_base(agu2,inp2);
[0694] set_y_minmax(agu2,0,(height–2)*width);
[0695] set_y_count_stride(agu2,height-2,width);
[0696] set_x_minmax(agu2,0,width);
[0697] set_x_count_stride(agu2,XCount,VectorSize);
[0698] agu_out1 = allocate_agu(“store”);
[0699] set_base(agu_out1,out1);
[0700] set_y_minmax(agu_out1,0,(height–2)*width);
[0701] set_y_count_stride(agu_out1,height-2,width);
[0702] set_x_minmax(agu_out1,0,width);
[0703] set_x_count_stride(agu_out1,XCount,VectorSize);
[0704] agu_out2 = allocate_agu(“store”);
[0705] set_base(agu_out2,out2);
[0706] set_y_minmax(agu_out2,0,(height–2)*width);
[0707] set_y_count_stride(agu_out2,height-2,width);
[0708] set_x_minmax(agu_out2,0,width);
[0709] set_x_count_stride(agu_out2,XCount,VectorSize);
[0710] Pre-header:
[0711] %p0 = agu1.load(); agu1.bump();
[0712] %p1 = agu1.load(); agu1.bump();
[0713] Loop:
[0714] %q0=phi(%p0:Preheader,%q2:Loop);
[0715] %q1=phi(%p1:Preheader,%q3:Loop);
[0716] %q2=agu1.load();agu1.bump();
[0717] %q3=agu1.load();agu1.bump();
[0718] %val_inp2=agu2.load();agu2.bump();
[0719] %val0=%q0;
[0720] / / Look at the 32-element vectors %q1, %q3 as if they were a 64-element vector, and build new 32-element vectors
[0721] / / Take elements 3, 4, 5, ..., 33, 34.
[0722] %val1=shuffle(%q1,%q3,[3,4,5,…,33,34])
[0723] %result=add%val0,%val1;
[0724] agu_out1.store(%result); agu_out1.bump();
[0725] agu_out2.store(%val_inp2);agu_out2.bump();
[0726] br(Loop);
[0727] Example (8)
[0728] In some exemplary aspects, as shown in Example 8, the AGU configuration code for agul can include AGU configuration code for an additional dimension (eg, dimension z).
[0729] In some exemplary aspects, as shown in Example 8, for example, in addition to dimension x that can be configured based on an X cycle and dimension y that can be configured based on a Y cycle, additional dimensions can be configured, such as dimension z.
[0730] In some exemplary aspects, AGU configuration code for agul may include, for example, AGU configuration code to set count parameters and step size (stride) parameters for additional dimension z according to a vertical extra dimension configuration for dimension y.
[0731] For example, as shown in Example 8, the AGU configuration code for agu1 may include, for example, AGU configuration code to set a count parameter for dimension z to 2 and a stride (step length) parameter for dimension z to 2*width, for example, set_z_count_stride(agu1,2,2*width).
[0732] In some exemplary aspects, compiler 160 may be configured to configure the AGU configuration code to set the same count parameter in the X dimension for all of the AGUs, eg, according to horizontal time shifting for the X dimension.
[0733] For example, as shown in Example 8, the AGU configuration code may include instructions to set the same count parameters in the X dimension for all of the AGUs. For example, the AGU configuration code may include instructions set_x_count_stride(agu1,XCount,width) for agu1, set_x_count_stride(agu2,XCount,width) for agu2, and set_x_count_stride(agu_out,XCount,width) for agu_out.
[0734] For example, as shown in Example 8, the same count parameter in the x dimension may be determined, for example, based on the count parameter of the x loop; the maximum offset of all load operations (e.g., two load operations) covered by the AGU; the minimum offset of the two load operations; and the vectorization factor, for example, as follows:
[0735] XCount=Count(XDim)=AlignUpTo((MaxXOffset–MinXOffset)+
[0736] Count(LoopXDim),VF) / VF=AlignUpTo((width–3)+3,VF) / VF
[0737] In some exemplary aspects, as shown in Example 8, the pre-header code may include MaxTS*C=2 AGU load instructions to be executed by the same agul, for example, according to a horizontal time shift for the X dimension.
[0738] For example, two pre-header AGU load instructions (eg, %p0=agul.load(); agul.bump(); and %p1=agul.load(); agul.bump()) may be used, for example, to load initial values for an AGU load iteration by agul.
[0739] In some exemplary aspects, as shown in Example 8, the loop code may include two conditional designation instructions, e.g., %q0=phi(%p0:Pre-header, %q2:Loop); %q1=phi(%p1:Pre-header, %q3:Loop), e.g., to designate (store) the result of an AGU load iteration performed by the same agu1 to a register.
[0740] In some exemplary aspects, as shown in Example 8, two conditional designation instructions (eg, %q0=phi(%p0:Pre-header, %q2:Loop); %q1=phi(%p1:Pre-header, %q3:Loop)) may be based on, for example, a phi operation.
[0741] In some exemplary aspects, as shown in Example 8, the loop code may include two specifying instructions, e.g., %q2=agu1.load(); agu1.bump(,) and %q3=agu1.load(); agu1.bump(,), e.g., to load and specify (store) the result of the current loop iteration performed by the same agu1 into a register.
[0742] For example, instruction %q2=agu1.load(); agu1.bump(,) may be configured to assign (store) to register %q2 the value loaded by agu1 in the first iteration over dimension z, eg, corresponding to a load operation inp1[index].
[0743] For example, instruction %q3=agu1.load(); agu1.bump() may be configured to assign (store) to register %q3 the value loaded by agu1 in the second iteration over dimension z, eg, corresponding to a load operation inp1[index+2*width+3].
[0744] For example, according to the code of Example 8, the pre-header instruction %p0=agu1.load(); agu1.bump() can be configured to assign (store) the first value loaded by agu1 to register %p0, for example, corresponding to the load operation inp1[index], for example, according to the horizontal time shift for the X dimension.
[0745] For example, according to the code of Example 8, the pre-header instruction %p1=agu1.load(); agu1.bump() can be configured to specify (store) the second value loaded by agu1 to register %p1, for example, corresponding to the load operation inp1[index+2*width], for example, according to the horizontal time shift for the X dimension.
[0746] For example, according to the code of Example 8, the loop instruction %q0=phi(%p0:Pre-header,%q2:Loop) can be configured to specify (store) the first value loaded by agu1 to register %q0, for example, corresponding to the load operation inp1[index], for example, in the first iteration in the order of the loop; or the value from register %q2, for example, corresponding to the load operation inp1[index], for example, in the non-first iteration in the order of the loop.
[0747] For example, according to the code of Example 8, the loop instruction %q1=phi(%p1:Pre-header,%q3:Loop) can be configured to specify (store) the second value loaded by agu1 to register %q1, for example, corresponding to the load operation inp1[index+2*width+3], for example, in the first iteration in the order of the loop; or the value from register %q3, for example, corresponding to the load operation inp1[index+2*width], for example, in the non-first iteration in the order of the loop.
[0748] In some exemplary aspects, as shown in Example 8, the loop code may include a shuffle instruction, e.g., %val1=shuffle(%q1,%q3,[3,4,5,…,33,34]), e.g., to shuffle the results of a previous AGU load iteration and a current AGU load iteration, e.g., with respect to the load operation inp1[index+2*width+3].
[0749] In one example, the size of register %q1 and the size of register %q3 may be 32 elements.
[0750] For example, register %q1 can represent Figure 5 The first vector in the third row of the scheme of , for example, covers blocks 0-4 of the third row, for example, where VectorSize=VF=4. For example, register %q3 may represent Figure 5 The second vector in the third row of the scheme, for example, covers blocks 4-8 of the third row.
[0751] For example, the load operation inp1[index+2*width+3] may include the end portion of register %q1 (eg, corresponding to Figure 5 4) with the beginning of register %q3 (e.g., corresponding to Figure 5Accordingly, the shuffle operation can be configured to concatenate the vectors of two registers (e.g., register %q1 and register %q3) to a 64-element vector, and use the required 32-element portion of the 64-element vector, for example, corresponding to the load operation inp1[index+2*width+3].
[0752] refer to Figure 6 , which schematically illustrates a method of compiling code for a processor. For example, Figure 6 One or more operations of the method may be performed by: a system, for example, system 100 ( Figure 1 ); devices, for example, device 102 ( Figure 1 ); a server, for example, server 170 ( Figure 1 ); and / or a compiler, for example, compiler 160 ( Figure 1 ) and / or compiler 200( Figure 2 ).
[0753] In some exemplary aspects, as indicated at block 602, the method may include identifying a plurality of memory access operations in a loop based on source code to be compiled into target code to be executed by a target processor. For example, the plurality of memory access operations may include at least a first memory access operation and a second memory access operation for the same memory pointer. For example, the first memory access operation may have a first offset, and the second memory access operation may have a second offset, for example, different from the first offset. For example, compiler 160( Figure 1 ) may be configured, for example, based on source code 112 ( Figure 1 ) to identify multiple memory access operations in a loop, for example, as described above.
[0754] In some exemplary aspects, as indicated at block 604, the method may include configuring the AGU configuration code to configure multiple memory access operations performed by the same AGU. For example, the compiler 160 ( Figure 1 ) can be configured to configure AGU configuration code to configure multiple memory access operations performed by the same AGU, for example, as described above.
[0755] In some exemplary aspects, as indicated at block 606, the method may include generating target code based on compiling the source code. For example, the target code may be based on the AGU configuration code. For example, the compiler 160 ( Figure 1 ) may be configured, for example, based on the source code 112 ( Figure 1 ) to generate the target code 115( Figure 1 ), the target code is based on AGU configuration code, for example, as described above.
[0756] refer to Figure 7 , which schematically illustrates an article of manufacture 700 according to some exemplary aspects. Article 700 may include one or more tangible computer-readable ("machine-readable") non-transitory storage media 702, which may include computer-executable instructions implemented, for example, by logic 704, which are operable to enable at least one computer processor to perform operations on a device 102 ( Figure 1 )、Server 170( Figure 1 ) and / or compiler 160( Figure 1 ) to implement one or more operations to enable device 102 ( Figure 1 )、Server 170( Figure 1 ) and / or compiler 160( Figure 1 ) performs, triggers and / or implements one or more operations and / or functionalities, and / or performs, triggers and / or implements reference Figures 1 to 6 One or more operations and / or functionalities described, and / or one or more operations described herein. The phrases "non-transitory machine-readable medium" and "computer-readable non-transitory storage medium" may be directed to include all computer-readable media, with the sole exception of transitory propagating signals.
[0757] In some exemplary aspects, the product 700 and / or the machine-readable storage medium 702 may include one or more types of computer-readable storage media capable of storing data, including volatile memory, non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. For example, the machine-readable storage medium 702 may include RAM, DRAM, double data rate DRAM (DDR-DRAM), SDRAM, static RAM (SRAM), ROM, programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory (e.g., NOR or NAND flash memory), content addressable memory (CAM), polymer memory, phase change memory, ferroelectric memory, silicon oxide nitride oxide silicon (SONOS) memory, disk, hard drive, etc. The computer-readable storage medium may include any suitable medium involved in downloading or transferring a computer program from a remote computer to a requesting computer via a communication link (e.g., a modem, radio, or network connection), the computer program being carried by a data signal embedded in a carrier wave or other propagation medium.
[0758] In some exemplary aspects, logic 704 may include instructions, data, and / or code that, if executed by a machine, may cause the machine to perform methods, processes, and / or operations as described herein. The machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, etc., and may be implemented using any suitable combination of hardware, software, firmware, etc.
[0759] In some exemplary aspects, logic 704 may include or may be implemented as software, a software module, an application, a program, a subroutine, an instruction, an instruction set, a computing code, a word, a value, a symbol, etc. Instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Instructions may be implemented according to a predefined computer language, manner, or syntax for instructing a processor to perform a specific function. Instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language, machine code, etc.
[0760] Examples
[0761] The following examples relate to further aspects.
[0762] Example 1 includes a product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions that are operable to, when executed by at least one processor, enable the at least one processor to cause a compiler to: identify multiple memory access operations in a loop based on source code to be compiled into target code to be executed by a target processor, the multiple memory access operations including at least a first memory access operation and a second memory access operation to the same memory pointer, wherein the first memory access operation has a first offset and the second memory access operation has a second offset different from the first offset; configure an address generation unit (AGU) configuration code to configure the multiple memory access operations performed by the same AGU; and generate target code based on compiling the source code, wherein the target code is based on the AGU configuration code.
[0763] Example 2 includes the subject matter of claim 1, and optionally wherein the instructions, when executed, cause a compiler to configure the AGU configuration code to configure multiple memory access operations based on a dimension of the loop, wherein the first offset comprises a first integer multiple of the dimension of the loop, and the second offset comprises a second integer multiple of the dimension of the loop.
[0764] Example 3 includes the subject matter of claim 1 or 2, and optionally wherein the instructions, when executed, cause a compiler to configure the AGU configuration code to configure multiple memory access operations based on the difference between the first offset and the second offset based on a determination that the difference between the first offset and the second offset is independent of the dimensions of the loop.
[0765] Example 4 includes the subject matter of any of claims 1 to 3, and optionally wherein the instructions, when executed, cause a compiler to configure the AGU configuration code to configure a first dimension of the AGU based on the dimension of the loop and to configure a second dimension of the AGU based on the irrelevant value based on determining that a difference between the first offset and the second offset is based on an integer multiple of a dimension of the loop and based on a dimension-independent value that is independent of the dimension of the loop.
[0766] Example 5 includes the subject matter of any one of claims 1 to 4, and optionally wherein the multiple memory access operations include a third memory access operation to the same memory pointer, the third memory access operation having a third offset different from the first offset and the second offset, wherein the AGU configuration code is used to configure the multiple memory access operations performed by the same AGU based on the third offset.
[0767] Example 6 includes the subject matter of any one of claims 1 to 5, and optionally wherein the loop comprises a first loop, wherein the first loop is nested within a second loop, wherein the instruction when executed causes the compiler to configure the AGU configuration code based on the second loop.
[0768] Example 7 includes the subject matter of claim 6, and optionally wherein the instructions, when executed, cause a compiler to configure AGU configuration code, the AGU configuration code comprising: first dimension code to configure a first dimension of the AGU based on a first cycle, second dimension code to configure a second dimension of the AGU based on a plurality of memory access operations, and third dimension code to configure a third dimension of the AGU based on the second cycle.
[0769] Example 8 includes the subject matter of claim 6, and optionally wherein the instructions, when executed, cause a compiler to configure AGU configuration code based on a determination of a difference between the first offset and the second offset based on a dimension of the second loop, the AGU configuration code comprising a loop swap instruction to indicate a loop swap between a dimension of the AGU corresponding to the first loop and a dimension of the AGU corresponding to the second loop.
[0770] Example 9 includes the subject matter of any of claims 1 to 8, and optionally wherein the AGU configuration code comprises: a first dimension code to configure a first dimension of the AGU based on cycles, and a second dimension code to configure a second dimension of the AGU based on multiple memory access operations.
[0771] Example 10 includes the subject matter of claim 9, and optionally wherein the first dimension code is configured to set a count parameter of a first dimension of the AGU based on a count parameter of a loop, and to set a step parameter of the first dimension of the AGU based on a step parameter of the loop, and wherein the second dimension code is configured to set a count parameter of a second dimension of the AGU based on a total count of memory access operations in a plurality of memory access operations, and to set a step parameter of the second dimension of the AGU based on a difference between a first offset and a second offset.
[0772] Example 11 includes the subject matter of claim 9 or 10, and optionally wherein the instructions when executed cause a compiler to generate loop code based on the loop, wherein the target code is based on the loop code, wherein the loop code includes multiple memory access instructions to be executed by the same AGU, wherein the multiple memory access instructions correspond to multiple memory access operations respectively.
[0773] Example 12 includes the subject matter of any of claims 9 to 11, and optionally wherein the plurality of memory access operations include a third memory access operation to the same memory pointer, the third memory access operation having a third offset different from the first offset and the second offset, wherein the instruction when executed causes a compiler to generate loop code including first, second, and third memory access instructions to be executed by the same AGU based on the first, second, and third memory access instructions, respectively, wherein the target code is based on the loop code.
[0774] Example 13 includes the subject matter of claim 12, and optionally wherein the instructions when executed cause a compiler to configure a second dimension of the AGU based on the second difference based on a determination that a first difference between the first offset and the second offset is equal to a second difference between the second offset and the third offset.
[0775] Example 14 includes the subject matter of any one of claims 9 to 13, and optionally wherein the loop comprises a first loop, wherein the first loop is nested within a second loop, wherein the AGU configuration code comprises a third dimension code to configure a third dimension of the AGU based on the second loop.
[0776] Example 15 includes the subject matter of any of claims 9 to 14, and optionally wherein the instructions, when executed, cause a compiler to configure a second dimension of the AGU based on a plurality of memory access operations based on a determination that a dimension count of the first loop is less than a dimension count supported by the same AGU.
[0777] Example 16 includes the subject matter of any of claims 9 to 15, and optionally wherein the instructions when executed cause a compiler to configure a second dimension of the AGU based on a plurality of memory access operations based on a determination that the plurality of memory access operations have no bounds and no fall-through.
[0778] Example 17 includes the subject matter of any of claims 9 to 16, and optionally wherein the instructions, when executed, cause a compiler to configure a second dimension of the AGU based on multiple memory access operations based on a determination that differences between consecutive offsets of the multiple memory access operations are the same.
[0779] Example 18 includes the subject matter of any one of claims 9 to 17, and optionally wherein the instructions, when executed, cause a compiler to configure a second dimension of the AGU based on the plurality of memory access operations based on a determination that: a dimension count of the first loop is less than a dimension count supported by the same AGU, the plurality of memory access operations have no bounds and no fall-through, and differences between consecutive offsets of the plurality of memory access operations are the same.
[0780] Example 19 includes the subject matter of any one of claims 1 to 18, and optionally wherein the AGU configuration code includes code to configure the same AGU according to a time-shifting scheme, the time-shifting scheme being configured to implement multiple memory access operations as multiple time-shifted memory access operations of the same AGU during multiple corresponding loop iterations.
[0781] Example 20 includes the subject matter of claim 19, and optionally wherein the AGU configuration code includes a loop swap instruction for a dimension of the AGU corresponding to the loop, wherein the loop swap instruction is configured to indicate a loop swap between the dimension of the AGU corresponding to the loop and another dimension of the AGU corresponding to the outer loop.
[0782] Example 21 includes the subject matter of claim 19, and optionally wherein the instructions when executed cause a compiler to configure AGU configuration code based on a determination of a difference between the first offset and the second offset based on a dimension of another loop including the loop, the AGU configuration code including a loop swap instruction for a dimension of the AGU corresponding to the loop, wherein the loop swap instruction is configured to indicate a loop swap between the dimension of the AGU corresponding to the loop and another dimension of the AGU corresponding to the other loop.
[0783] Example 22 includes the subject matter of any one of claims 19 to 21, and optionally wherein the AGU configuration code is configured to set the same count parameter for all AGUs to perform memory access operations during a loop, wherein the same count parameter is based on a count parameter of the loop, a maximum offset of multiple memory access operations, and a minimum offset of multiple memory access operations.
[0784] Example 23 includes the subject matter of claim 22, and optionally wherein the AGU configuration code is to set the same count parameter denoted as Count as follows:
[0785] Count=AlignUpTo((MaxXOffset–MinXOffset)+
[0786] Count(Loop),VF) / VF
[0787] Where MaxXOffset represents the maximum offset of multiple memory access operations.
[0788] Where MinXOffst represents the minimum offset of multiple memory access operations.
[0789] Where VF represents the vectorization factor,
[0790] Where Count(Loop) represents the counting parameter of the loop, and
[0791] The operation AlignUpTo(a,b) is used to align a upward to an integer multiple of b.
[0792] Example 24 includes the subject matter of any one of claims 19 to 23, and optionally wherein the instructions, when executed, cause a compiler to generate loop code based on the loop, and pre-header code to be executed before a header of the loop code, wherein the target code is based on the loop code and the pre-header code, wherein the pre-header code includes multiple AGU load instructions to be executed by the same AGU, wherein a count of the AGU load instructions in the multiple AGU load instructions is based on a count of memory access operations in the multiple memory access operations.
[0793] Example 25 includes the subject matter of claim 24, and optionally wherein the loop code includes a plurality of store instructions to store results of a corresponding plurality of AGU load iterations performed by the same AGU, wherein a count of the plurality of store instructions is based on a count of memory access operations in the plurality of memory access operations.
[0794] Example 26 includes the subject matter of claim 25 and optionally wherein the plurality of store instructions are configured to cause the target processor to store a plurality of load results of the plurality of AGU load instructions in the pre-header code at an iteration that is first in order among the plurality of AGU load iterations.
[0795] Example 27 includes the subject matter of any one of claims 19 to 26, and optionally wherein the instructions, when executed, cause a compiler to configure AGU configuration code based on a determination that the multiple memory access operations include only load operations, the AGU configuration code including code to configure the same AGU according to a time-shifting scheme.
[0796] Example 28 includes the subject matter of any one of claims 19 to 27, and optionally wherein the instructions, when executed, cause a compiler to configure AGU configuration code based on a determination that a difference between successive offsets of a plurality of memory access operations is independent of the dimensions of the loop, the AGU configuration code being used to configure code of the same AGU according to a time-shifting scheme.
[0797] Example 29 includes the subject matter of any one of claims 19 to 28, and optionally wherein the instructions, when executed, cause a compiler to configure AGU configuration code based on a determination that there are no bounds and no cut-through on multiple memory access operations, the AGU configuration code comprising code to configure the same AGU according to a time-shifting scheme.
[0798] Example 30 includes the subject matter of any one of claims 19 to 29, and optionally wherein the instructions when executed cause a compiler to configure AGU configuration code based on a determination that multiple memory access operations are not masked, the AGU configuration code comprising code to configure the same AGU according to a time-shifting scheme.
[0799] Example 31 includes the subject matter of any of claims 19 to 30, and optionally wherein the instructions when executed cause a compiler to configure AGU configuration code based on a determination that the loop does not include any side effects, the AGU configuration code comprising code to configure the same AGU according to a time-shifting scheme.
[0800] Example 32 includes the subject matter of any one of claims 19 to 31, and optionally wherein the instructions, when executed, cause a compiler to configure AGU configuration code, the AGU configuration code comprising code to configure the same AGU according to a time-shifting scheme based on a determination that: the multiple memory access operations include only load operations, the difference between consecutive offsets of the multiple memory access operations is independent of the dimensions of the loop, the multiple memory access operations have no bounds and no fall-through, the multiple memory access operations are not masked, and the loop does not include any side effects.
[0801] Example 33 includes the subject matter of any of claims 1-32, and optionally wherein the plurality of memory access operations comprises a plurality of load operations.
[0802] Example 34 includes the subject matter of any one of claims 1 to 33, and optionally wherein the source code comprises Open Computing Language (OpenCL) code.
[0803] Example 35 includes the subject matter of any one of claims 1 to 34, and optionally wherein the computer-executable instructions, when executed, cause a compiler to compile source code into target code according to a low-level virtual machine (LLVM)-based compilation scheme.
[0804] Example 36 includes the subject matter of any one of claims 1 to 35, and optionally wherein the target code is configured for execution by a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor.
[0805] Example 37 includes the subject matter of any one of claims 1 to 36, and optionally wherein the object code is configured for execution by a target vector processor.
[0806] Example 38 includes a compiler configured to perform any of the operations described in any of Examples 1 to 37.
[0807] Example 39 includes a computing device configured to perform any of the operations described in any of Examples 1 to 37.
[0808] Example 40 includes a computing system comprising: at least one memory for storing instructions; and at least one processor for retrieving instructions from the memory and executing the instructions to cause the computing system to perform any of the operations described in any of Examples 1 to 37.
[0809] Example 41 includes a computing system comprising: a compiler configured to generate target code according to any of the operations described in any of Examples 1 to 37; and a processor configured to execute the target code.
[0810] Example 42 includes a device comprising means for performing any of the operations described in any of Examples 1-37.
[0811] Example 43 includes an apparatus comprising: a memory interface; and a processing circuit system configured to: perform any of the operations described in any of Examples 1-37.
[0812] Example 44 includes a method comprising any of the operations described in any of Examples 1 to 37.
[0813] The functions, operations, components, and / or features described herein with reference to one or more aspects may be combined with or utilized in combination with one or more other functions, operations, components, and / or features described herein with reference to one or more other aspects, or vice versa.
[0814] While certain features have been illustrated and described herein, numerous modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the disclosure.
Claims
1. A product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions operable to, when executed by at least one processor, enable the at least one processor to cause a compiler to: identifying, based on source code to be compiled into target code to be executed by a target processor, a plurality of memory access operations in a loop, the plurality of memory access operations comprising at least a first memory access operation and a second memory access operation to a same memory pointer, wherein the first memory access operation has a first offset and the second memory access operation has a second offset different from the first offset; configuring an address generation unit (AGU) configuration code to configure the plurality of memory access operations performed by the same AGU; as well as The object code is generated based on compiling the source code, wherein the object code is based on the AGU configuration code.
2. The product of claim 1 , wherein the instructions, when executed, cause the compiler to configure the AGU configuration code to configure the plurality of memory access operations based on a dimension of the loop, wherein the first offset comprises a first integer multiple of the dimension of the loop, and the second offset comprises a second integer multiple of the dimension of the loop.
3. The product of claim 1 , wherein the instructions, when executed, cause the compiler to configure the AGU configuration code to configure the plurality of memory access operations based on a difference between the first offset and the second offset based on a determination that the difference between the first offset and the second offset is independent of a dimension of the loop.
4. The article of manufacture of claim 1 , wherein the instructions, when executed, cause the compiler to configure the AGU configuration code to configure a first dimension of the AGU based on the dimension of the loop and to configure a second dimension of the AGU based on the irrelevant value based on determining that the difference between the first offset and the second offset is based on an integer multiple of a dimension of the loop and based on a dimension-independent value that is independent of the dimension of the loop.
5. The product of claim 1 , wherein the plurality of memory access operations include a third memory access operation to the same memory pointer, the third memory access operation having a third offset different from the first offset and the second offset, wherein the AGU configuration code is used to configure the plurality of memory access operations performed by the same AGU based on the third offset.
6. The article of manufacture of claim 1, wherein the loop comprises a first loop, wherein the first loop is nested within a second loop, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code based on the second loop.
7. The product of claim 6, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code, the AGU configuration code comprising: A first dimension code is used to configure a first dimension of the AGU based on the first cycle, a second dimension code is used to configure a second dimension of the AGU based on the plurality of memory access operations, and a third dimension code is used to configure a third dimension of the AGU based on the second cycle.
8. The article of manufacture of claim 6, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code based on a determination of a difference between the first offset and the second offset based on a dimension of the second loop, the AGU configuration code comprising a loop swap instruction to indicate a loop swap between a dimension of the AGU corresponding to the first loop and a dimension of the AGU corresponding to the second loop.
9. The product of claim 1, wherein the AGU configuration code comprises: A first dimension code is used to configure a first dimension of the AGU based on the cycle, and a second dimension code is used to configure a second dimension of the AGU based on the plurality of memory access operations.
10. The product of claim 9, wherein the first dimension code is configured to set a count parameter of the first dimension of the AGU based on a count parameter of the loop, and to set a step parameter of the first dimension of the AGU based on a step parameter of the loop, and wherein the second dimension code is configured to set a count parameter of the second dimension of the AGU based on a total count of memory access operations in the plurality of memory access operations, and to set a step parameter of the second dimension of the AGU based on a difference between the first offset and the second offset.
11. The product of claim 9, wherein the instructions, when executed, cause the compiler to generate loop code based on the loop, wherein the target code is based on the loop code, wherein the loop code includes a plurality of memory access instructions to be executed by the same AGU, wherein the plurality of memory access instructions correspond to the plurality of memory access operations, respectively.
12. The article of manufacture of claim 9, wherein the plurality of memory access operations include a third memory access operation to the same memory pointer, the third memory access operation having a third offset different from the first offset and the second offset, wherein the instruction when executed causes the compiler to generate loop code including the first, second, and third memory access instructions to be executed by the same AGU based on the first, second, and third memory access instructions, respectively, wherein the target code is based on the loop code.
13. The article of manufacture of claim 9, wherein the loop comprises a first loop, wherein the first loop is nested within a second loop, wherein the AGU configuration code comprises a third dimension code to configure a third dimension of the AGU based on the second loop.
14. The article of claim 9, wherein the instructions, when executed, cause the compiler to configure the second dimension of the AGU based on the plurality of memory access operations based on a determination that a dimension count of the first loop is less than a dimension count supported by the same AGU.
15. The article of claim 9, wherein the instructions, when executed, cause the compiler to configure the second dimension of the AGU based on the plurality of memory access operations based on a determination that the plurality of memory access operations have no bounds and no fall-through.
16. The article of claim 9, wherein the instructions, when executed, cause the compiler to configure the second dimension of the AGU based on the plurality of memory access operations based on a determination that differences between consecutive offsets of the plurality of memory access operations are the same.
17. The article of manufacture of claim 9, wherein the instructions, when executed, cause the compiler to configure the second dimension of the AGU based on the plurality of memory access operations based on a determination that a dimension count of the first loop is less than a dimension count supported by the same AGU, the plurality of memory access operations have no bounds and no fall-through, and differences between consecutive offsets of the plurality of memory access operations are the same.
18. The product of any one of claims 1 to 17, wherein the AGU configuration code comprises code for configuring the same AGU according to a time-shifting scheme, wherein the time-shifting scheme is configured to implement the multiple memory access operations as multiple time-shifted memory access operations of the same AGU during multiple corresponding loop iterations.
19. The product of claim 18, wherein the AGU configuration code comprises a loop swap instruction for a dimension of the AGU corresponding to the loop, wherein the loop swap instruction is configured to indicate a loop swap between the dimension of the AGU corresponding to the loop and another dimension of the AGU corresponding to an outer loop.
20. The product of claim 18, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code to include a loop swap instruction for a dimension of the AGU corresponding to the loop based on a determination of a difference between the first offset and the second offset based on a dimension of another loop including the loop, wherein the loop swap instruction is configured to indicate a loop swap between the dimension of the AGU corresponding to the loop and another dimension of the AGU corresponding to the other loop.
21. The product of claim 18, wherein the AGU configuration code is configured to set the same counting parameter for all AGUs to perform memory access operations during the cycle, wherein the same counting parameter is based on a counting parameter of the cycle, a maximum offset of the plurality of memory access operations, and a minimum offset of the plurality of memory access operations.
22. The article of manufacture of claim 21, wherein the AGU configuration code is to set the same count parameter denoted as Count as follows: Count=AlignUpTo((MaxXOffset–MinXOffset)+ Count(Loop),VF) / VF wherein MaxXOffset represents the maximum offset of the plurality of memory access operations, wherein MinXOffst represents the minimum offset of the plurality of memory access operations, Where VF represents the vectorization factor, wherein Count(Loop) represents the counting parameter of the loop, and The operation AlignUpTo(a,b) is used to align a upward to an integer multiple of b.
23. The article of claim 18, wherein the instructions, when executed, cause the compiler to generate loop code based on the loop, and pre-header code to be executed before a header of the loop code, wherein the target code is based on the loop code and the pre-header code, wherein the pre-header code includes a plurality of AGU load instructions to be executed by the same AGU, wherein a count of AGU load instructions in the plurality of AGU load instructions is based on a count of memory access operations in the plurality of memory access operations.
24. The article of manufacture of claim 23, wherein the loop code comprises a plurality of store instructions to store results of a corresponding plurality of AGU load iterations performed by the same AGU, wherein a count of the plurality of store instructions is based on a count of the memory access operations in the plurality of memory access operations.
25. The article of claim 24, wherein the plurality of store instructions are configured to cause the target processor to store a plurality of load results of the plurality of AGU load instructions in the pre-header code at an iteration that is first in order among the plurality of AGU load iterations.
26. The article of claim 18, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code based on a determination that the plurality of memory access operations include only load operations, the AGU configuration code comprising code to configure the same AGU according to the time-shifting scheme.
27. The product of claim 18, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code based on a determination that a difference between successive offsets of the plurality of memory access operations is independent of a dimension of the loop, the AGU configuration code comprising code to configure the same AGU according to the time-shifting scheme.
28. The product of claim 18, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code based on a determination that the plurality of memory access operations are unbounded and have no fall-through, the AGU configuration code comprising code to configure the same AGU according to the time-shifting scheme.
29. The product of claim 18, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code based on a determination that the plurality of memory access operations are not masked, the AGU configuration code comprising code to configure the same AGU according to the time-shifting scheme.
30. The article of claim 18, wherein the instructions, when executed, cause the compiler to configure the AGU configuration code based on a determination that the loop does not include any side effects, the AGU configuration code comprising code to configure the same AGU according to the time-shifting scheme.
31. The product of any one of claims 1 to 17, wherein the plurality of memory access operations comprises a plurality of load operations.
32. The product of any one of claims 1 to 17, wherein the source code comprises Open Computing Language (OpenCL) code.
33. The product of any one of claims 1 to 17, wherein the computer executable instructions, when executed, cause the compiler to compile the source code into the target code according to a Low Level Virtual Machine (LLVM)-based compilation scheme.
34. The product of any one of claims 1 to 17, wherein the object code is configured for execution by a Very Long Instruction Word (VLIW) Single Instruction / Multiple Data (SIMD) target processor.
35. A product according to any one of claims 1 to 17, wherein the object code is configured for execution by a target vector processor.
36. A computing system comprising: at least one memory for storing instructions; as well as at least one processor to retrieve the instructions from the memory and to execute the instructions to cause the computing system to: identifying, based on source code to be compiled into target code to be executed by a target processor, a plurality of memory access operations in a loop, the plurality of memory access operations comprising at least a first memory access operation and a second memory access operation to a same memory pointer, wherein the first memory access operation has a first offset and the second memory access operation has a second offset different from the first offset; configuring an address generation unit (AGU) configuration code to configure the plurality of memory access operations performed by the same AGU; as well as The object code is generated based on compiling the source code, wherein the object code is based on the AGU configuration code.
37. The computing system of claim 36, comprising the target processor to execute the target code.
38. A method comprising: identifying, based on source code to be compiled into target code to be executed by a target processor, a plurality of memory access operations in a loop, the plurality of memory access operations comprising at least a first memory access operation and a second memory access operation to a same memory pointer, wherein the first memory access operation has a first offset and the second memory access operation has a second offset different from the first offset; configuring an address generation unit (AGU) configuration code to configure the plurality of memory access operations performed by the same AGU; as well as The object code is generated based on compiling the source code, wherein the object code is based on the AGU configuration code.
39. The method of claim 38, wherein the AGU configuration code comprises first dimension code to configure a first dimension of the AGU based on the cycle, and second dimension code to configure a second dimension of the AGU based on the plurality of memory access operations.