Apparatus, system and method for compiling code for processor
By introducing VMP OpenCL compiler and LLVM compilation scheme into the compiler, the compilation process is optimized to automatically vectorize and downgrade instructions, and the problem of difficulty in generating efficient target code in the existing technology is solved, and high-performance code generation in the vector processor environment is achieved.
Patent Information
- Application Number
- CN202380072155.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-16
Smart Images

Figure CN120019359A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 415,309, filed on October 12, 2022, entitled “APPARATUS, SYSTEM, AND METHOD OF VECTOR PROCESSING,” the entire disclosure of which is incorporated herein by reference. Background Art
[0003] The compiler may be configured to compile source code into object code configured for execution by the processor.
[0004] There is a need to provide technical solutions to support efficient processing functionality. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity of presentation. In addition, reference numerals may be repeated in the drawings to indicate corresponding or similar elements. The drawings are listed below.
[0006] Figure 1 is a schematic block diagram illustration of a system according to some exemplary aspects.
[0007] Figure 2 is a schematic illustration of a compiler according to some exemplary aspects.
[0008] Figure 3 is a schematic illustration of a vector processor according to some exemplary aspects.
[0009] Figure 4 is a schematic flow chart illustration of a method of compiling code for a processor according to some exemplary aspects.
[0010] Figure 5 is a schematic illustration of a product according to some exemplary aspects. DETAILED DESCRIPTION
[0011] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of some aspects. However, it will be appreciated by those of ordinary skill in the art that some aspects may be practiced without these specific details. In other cases, well-known methods, procedures, components, units and / or circuits are not described in detail to avoid obscuring the discussion.
[0012] Some portions of the following detailed description are presented in terms of algorithms and symbolic representations of operations on data bits or binary digital signals within a computer memory. These algorithmic descriptions and representations may be techniques used by those skilled in the data processing arts to convey the substance of their work to others skilled in the art.
[0013] An algorithm is here and generally considered to be a self-consistent sequence of acts or operations leading to a desired result. These include physical manipulations of physical quantities. Typically, but not necessarily, these quantities capture forms of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for common sense reasons, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. It should be understood, however, that all of these terms and similar terms should be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities.
[0014] Discussions herein utilizing terms such as, for example, "process," "compute," "calculate," "determine," "create," "analyze," "verify," and the like may refer to the operations and / or processes of a computer, computing platform, computing system, or other electronic computing device that manipulates data represented as physical (e.g., electronic) quantities within the computer's registers and / or memories and / or transforms that data into other data similarly represented as physical quantities within the computer's registers and / or memories or other information storage media that may store instructions for performing operations and / or processes.
[0015] As used herein, the terms "plurality" and "a plurality" include, for example, "a plurality" or "two or more." For example, "a plurality of items" includes two or more items.
[0016] References to "one aspect," "an aspect," "exemplary aspect," "various aspects," etc. indicate that the aspects so described may include particular features, structures, or characteristics, but not every aspect necessarily includes the particular features, structures, or characteristics. Furthermore, repeated use of the phrase "in one aspect" does not necessarily refer to the same aspect, although it may.
[0017] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe common objects merely indicates that different instances of the same object are being referenced and is not intended to imply that the objects so described must be in a given sequence in time, space, ranking, or in any other manner.
[0018] For example, some aspects may capture the form of entirely hardware aspects, entirely software aspects, or aspects including both hardware and software elements.Some aspects may be implemented in software, including but not limited to firmware, resident software, microcode, etc.
[0019] Furthermore, some aspects may be captured in the form of a computer program product that can be accessed from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For example, a computer-usable or computer-readable medium can be or can include any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with an instruction execution system, device, or apparatus.
[0020] In some exemplary aspects, the medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium.
[0021] In some exemplary aspects, a data processing system suitable for storing and / or executing program code may include at least one processor coupled directly or indirectly to a memory element, for example, via a system bus. The memory element may include, for example, local memory employed during actual execution of the program code, a mass storage device, and a cache memory that may provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from a mass storage device during execution.
[0022] In some exemplary aspects, input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system directly or through an intermediate I / O controller. In some exemplary aspects, a network adapter can be coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices, such as through an intermediate private or public network. In some exemplary aspects, modems, cable modems, and Ethernet cards are exemplary examples of network adapter types. Other suitable components can be used.
[0023] Some aspects may be used in connection with various devices and systems, such as computing devices, computers, mobile computers, non-mobile computers, server computers, and the like.
[0024] As used herein, the term "circuitry" may refer to, be a part of, or include an application specific integrated circuit (ASIC), an integrated circuit, an electronic circuit, a processor (shared, dedicated, or grouped), and / or memory (shared. Dedicated or grouped) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable hardware components that provide the described functionality. In some aspects, some functions associated with the circuitry may be implemented by one or more software or firmware modules. In some aspects, the circuitry may include logic that is at least partially operable in hardware.
[0025] The term "logic" may refer to, for example, computing logic embedded in the circuit system of a computing device and / or computing logic stored in the memory of a computing device. For example, the logic may be accessed by a processor of a computing device to execute the computing logic to perform computing functions and / or operations. In one example, the logic may be embedded in various types of memory and / or firmware, such as silicon blocks of various chips and / or processors. The logic may be included in various circuit systems and / or implemented as part of various circuit systems, such as processor circuit systems, control circuit systems, and / or the like. In one example, the logic may be embedded in volatile memory and / or non-volatile memory, including random access memory, read-only memory, programmable memory, magnetic memory, flash memory, persistent memory, etc. The logic may be executed by one or more processors using memory (e.g., registers, lags, buffers, and / or the like) coupled to one or more processors, for example, executing the logic as needed.
[0026] Reference now Figure 1 , which schematically illustrates a block diagram of a system 100 according to some exemplary aspects.
[0027] like Figure 1 As shown, in some demonstrative aspects, system 100 may include a computing device 102 .
[0028] In some demonstrative aspects, device 102 may be implemented using suitable hardware components and / or software components, such as processors, controllers, memory units, storage units, input units, output units, communication units, operating systems, applications, and the like.
[0029] In some demonstrative aspects, device 102 may comprise, for example, a computer, a mobile computing device, a non-mobile computing device, a laptop computer, a notebook computer, a tablet computer, a handheld computer, a personal computer (PC), or the like.
[0030] In some exemplary aspects, device 102 may include, for example, one or more of the following: processor 191, input unit 192, output unit 193, memory unit 194, and / or storage unit 195. Device 102 may optionally include other suitable hardware components and / or software components. In some exemplary aspects, some or all components of one or more of devices in device 102 may be enclosed in a common housing or packaging and may be interconnected or operably associated using one or more wired or wireless links. In other aspects, components of one or more of devices in device 102 may be distributed in multiple or separate devices.
[0031] In some exemplary aspects, processor 191 may include, for example, a central processing unit (CPU), a digital signal processor (DSP), one or more processor cores, a single core processor, a dual core processor, a multi-core processor, a microprocessor, a host processor, a controller, multiple processors or controllers, a chip, a microchip, one or more circuits, a circuit system, a logic unit, an integrated circuit (IC), an application specific IC (ASIC), or any other suitable general-purpose or specific processor or controller. Processor 191 may execute, for example, instructions of an operating system (OS) of device 102 and / or instructions of one or more suitable applications.
[0032] In some exemplary aspects, input unit 192 may include, for example, a keyboard, a keypad, a mouse, a touch screen, a touch pad, a trackball, a stylus, a microphone, or other suitable pointing device or input device. Output unit 193 may include, for example, a monitor, a screen, a touch screen, a flat panel display, a light emitting diode (LED) display unit, a liquid crystal display (LCD) display unit, a plasma display unit, one or more audio speakers or headphones, or other suitable output devices.
[0033] In some exemplary aspects, memory unit 194 includes, for example, random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous DRAM (SD-RAM), flash memory, volatile memory, non-volatile memory, cache memory, buffer, short-term memory unit, long-term memory unit, or other suitable memory unit. Storage unit 195 may include, for example, a hard disk drive, a solid-state drive (SSD), or other suitable removable or non-removable storage unit. Memory unit 194 and / or storage unit 195 may, for example, store data processed by device 102.
[0034] In some demonstrative aspects, device 102 may be configured to communicate with one or more other devices via at least one network 103 (eg, a wireless and / or wired network).
[0035] In some exemplary aspects, network 103 may include a wired network, a local area network (LAN), a wireless network, a wireless LAN (WLAN) network, a radio network, a cellular network, a WiFi network, an IR network, a Bluetooth (BT) network, and the like.
[0036] In some demonstrative aspects, device 102 may be configured to perform and / or execute one or more operations, modules, processes, procedures, and / or the like, e.g., as described herein.
[0037] In some demonstrative aspects, device 102 may include compiler 160, which may be configured to generate object code 115 based on source code 112, for example, as described below.
[0038] In some exemplary aspects, compiler 160 may be configured to translate source code 112 into target code 115, eg, as described below.
[0039] In some demonstrative aspects, compiler 160 may include or may be implemented as software, a software module, an application, a program, a subroutine, instructions, an instruction set, computing code, words, values, symbols, and / or the like.
[0040] In some exemplary aspects, source code 112 may include computer code written in a source language.
[0041] In some exemplary aspects, the source language may include a programming language. For example, the source language may include a high-level programming language, such as, for example, C language, C++ language and / or the like.
[0042] In some exemplary aspects, target code 115 may include computer code written in a target language.
[0043] In some exemplary aspects, the target language can include a low-level language such as, for example, assembly language, object code, machine code, or the like.
[0044] In some exemplary aspects, object code 115 may include one or more purpose files, which may, for example, create and / or form an executable program.
[0045] In some exemplary aspects, the executable program may be configured to be executed on a target computer. For example, the target computer may include specific computer hardware, a specific machine and / or a specific operating system.
[0046] In some exemplary aspects, the executable program may be configured to be executed on processor 180, eg, as described below.
[0047] In some demonstrative aspects, processor 180 may include a vector processor 180, eg, as described below. In other aspects, processor 180 may include any other type of processor.
[0048] Some exemplary aspects are described herein with respect to a compiler (e.g., compiler 160) that is configured to compile source code 112 into target code 115 that is configured to be executed by a vector processor 180, e.g., as described below. In other aspects, a compiler (e.g., compiler 160) is configured to compile source code 112 into target code 115 that is configured to be executed by any other type of processor 180.
[0049] In some demonstrative aspects, processor 180 may be implemented as part of device 102 .
[0050] In other aspects, processor 180 may be implemented as part of any other device separate from device 102 , for example.
[0051] In some demonstrative aspects, vector processor 180 (also referred to as an "array processor") may include a processor that may be configured to process an entire vector in one instruction, eg, as described below.
[0052] In other aspects, the executable program may be configured to be executed on any other additional or alternative type of processor.
[0053] In some exemplary aspects, vector processor 180 may be designed to support high-performance image and / or vector processing. For example, vector processor 180 may be configured to process 1 / 2 / 3 / 4D arrays and / or floating point arrays of fixed point data very quickly and / or efficiently.
[0054] In some exemplary aspects, vector processor 180 can be configured to process arbitrary data, such as structures with pointers to structures. For example, vector processor 180 can include a scalar processor to calculate non-vector data, such as assuming that the non-vector data is minimal.
[0055] In some demonstrative aspects, compiler 160 may be implemented as a native application to be executed by device 102. For example, memory unit 194 and / or storage unit 195 may store instructions generated in compiler 160, and / or processor 191 may be configured to execute instructions generated in compiler 160 and / or perform one or more calculations and / or processes of compiler 160, e.g., as described below.
[0056] In other aspects, compiler 160 may comprise a remote application to be executed by any suitable computing system (eg, server 170 ).
[0057] In some exemplary aspects, server 170 may include at least a remote server, a network-based server, a cloud server, and / or any other server.
[0058] In some exemplary aspects, server 170 may include a suitable memory and / or storage unit 174 having stored thereon instructions generated in compiler 160 and a suitable processor 171 to execute the instructions, e.g., as described below.
[0059] In some exemplary aspects, compiler 160 may include a combination of remote applications and local applications.
[0060] In one example, compiler 160 may be downloaded and / or received by a user of device 102 from another computing system (e.g., server 170) such that compiler 160 may be executed locally by the user of device 102. For example, instructions may be received and stored temporarily in a memory or any suitable short-term storage or buffer of device 102, e.g., prior to execution by processor 191 of device 102.
[0061] In another example, compiler 160 may include a client module to be executed locally by device 102 and a server module to be executed by server 170. For example, the client module may include and / or may be implemented as a local application, a web application, a website, a web client, e.g., a hypertext markup language (HTML) web application, etc.
[0062] For example, one or more first operations of compiler 160 may be performed locally, such as by device 102 , and / or one or more second operations of compiler 160 may be performed remotely, such as by server 170 .
[0063] In other aspects, compiler 160 may include or be implemented by any other suitable computing arrangement and / or scheme.
[0064] In some demonstrative aspects, system 100 may include an interface 110 (eg, a user interface) to interface between a user of device 102 and one or more elements of system 100 (eg, compiler 160).
[0065] In some demonstrative aspects, interface 110 may be implemented using any suitable hardware components and / or software components, such as a processor, a controller, a memory unit, a storage unit, an input unit, an output unit, a communication unit, an operating system, and / or an application.
[0066] In some aspects, interface 110 may be implemented as part of any suitable module, system, device, or component of system 100 .
[0067] In other aspects, interface 110 may be implemented as a separate element of system 100 .
[0068] In some demonstrative aspects, interface 110 may be implemented as part of device 102. For example, interface 110 may be associated with device 102 and / or included as part of the device.
[0069] In one example, interface 110 can be implemented as part of any suitable application, such as middleware and / or device 102. For example, interface 110 can be implemented as part of compiler 160 and / or part of the OS of device 102.
[0070] In some demonstrative aspects, interface 110 may be implemented as part of server 170. For example, interface 110 may be associated with server 170 and / or included as part of the server.
[0071] In one example, interface 110 may include or be part of: a web-based application, a website, a web page, a plug-in, an ActiveX control, a rich content component (eg, a Flash or Shockwave component), or the like.
[0072] In some exemplary aspects, interface 110 may be associated therewith and / or may include, for example, a gateway (GW) 113 and / or an application programming interface (API) 114, for example, to transmit information and / or communicate between elements of system 100 and / or to one or more other parties (e.g., internal or external parties), users, applications and / or systems.
[0073] In some aspects, interface 110 may include any suitable graphical user interface (GUI) 116 and / or any other suitable interface.
[0074] In some demonstrative aspects, interface 110 may be configured to receive source code 112 from, for example, a user of device 102 via GUI 116 and / or API 114 .
[0075] In some exemplary aspects, interface 110 may be configured to transfer source code 112 to, for example, compiler 160 , for example, to generate object code 115 , for example, as described below.
[0076] refer to Figure 2 , which schematically illustrates a compiler 200 according to some exemplary aspects. For example, the compiler 160 ( Figure 1 ) may implement one or more elements of compiler 200 and / or may perform one or more operations and / or functionalities of compiler 200.
[0077] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may be configured to generate target code 233, for example, by compiling source code 212 in a source language.
[0078] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may include a front end 210 configured to receive and analyze source code 212 in a source language.
[0079] In some exemplary aspects, front end 210 may be configured to generate intermediate code 213 , for example, based on source code 212 .
[0080] In some exemplary aspects, intermediate code 213 may comprise a lower-level representation of source code 212 .
[0081] In some exemplary aspects, front end 210 can be configured to perform, for example, lexical analysis, syntactic analysis, semantic analysis, and / or any other additional or alternative types of analysis of source code 212 .
[0082] In some exemplary aspects, front end 210 can be configured to identify errors and / or problems using the results of the analysis of source code 212. For example, front end 210 can be configured to generate error information, e.g., including error and / or warning messages, which can identify a location in source code 212, e.g., where an error or problem is detected.
[0083] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may include a middle end 220 configured to receive and process intermediate code 213 and generate adjusted (eg, optimized) intermediate code 223 .
[0084] In some demonstrative aspects, middle end 220 may be configured to perform one or more adjustments (eg, optimizations) to intermediate code 213 , eg, to generate adjusted intermediate code 223 .
[0085] In some exemplary aspects, middle end 220 can be configured to perform one or more optimizations on intermediate code 213 , eg, independent of the type of target computer used to execute target code 233 .
[0086] In some exemplary aspects, middle end 220 can be implemented to support the use of optimized intermediate code 223, eg, for different machine types.
[0087] In some exemplary aspects, middle end 220 may be configured to optimize the intermediate representation of intermediate code 223 , for example, to improve the performance and / or quality of the generated target code 233 .
[0088] In some exemplary aspects, one or more optimizations of intermediate code 213 may include, for example, inline expansion, dead code elimination, constant propagation, loop transformation, parallelization, and / or the like.
[0089] In some exemplary aspects, such as Figure 2 As shown, the compiler 200 may include a back end 230 configured to receive and process the adjusted intermediate code 213 , and generate a target code 233 based on the adjusted intermediate code 213 .
[0090] In some exemplary aspects, backend 230 may be configured to perform one or more operations and / or processes that may be specific to a target computer used to execute target code 233. For example, backend 230 may be configured to process optimized intermediate code 213 by applying analysis, transformation, and / or optimization operations to adjusted intermediate code 213, which operations may be configured, for example, based on a target computer used to execute target code 233.
[0091] In some exemplary aspects, the one or more analysis, transformation, and / or optimization operations applied to the adjusted intermediate code 213 may include, for example, resource and storage decisions, such as register allocation, instruction scheduling, and / or the like.
[0092] In some exemplary aspects, the object code 233 may include target-dependent assembly code that may be specific to a target computer used to execute the object code 233 and / or a target operating system of the target computer.
[0093] In some exemplary aspects, the object code 233 may include code for a processor (e.g., vector processor 180 ( Figure 1 ))'s target-dependent assembly code.
[0094] In some exemplary aspects, compiler 200 may include a vector microcode processor (VMP) open computing language (OpenCL) compiler, for example, as described below. In other aspects, compiler 200 may include any other type of vector processor compiler, or may be implemented as part of any other type of vector processor compiler.
[0095] In some exemplary aspects, the VMP OpenCL compiler may include a low-level virtual machine (LLVM)-based compiler that may be configured according to an LLVM-based compilation scheme, for example, to reduce OpenCL C code to VMP accelerator assembly code, for example, suitable for use by vector processor 180 ( Figure 1 )implement.
[0096] In some exemplary aspects, compiler 200 may include one or more techniques that may be required to compile code into a format suitable for a VMP architecture, for example, in addition to an open source LLVM compiler pass.
[0097] In some exemplary aspects, FE 210 may be configured to parse OpenCL C code and translate it, for example, via an abstract syntax tree (AST), into, for example, an LLVM intermediate representation (IR).
[0098] In some exemplary aspects, compiler 200 may include a dedicated API, for example, to detect the correct pattern for compiler pattern matching, e.g., a pattern suitable for VMP. For example, VMP may be configured as a complex instruction set computer (CISC) machine that implements a very complex instruction set architecture (ISA) that may be difficult to target from standard C code. Accordingly, compiler pattern matching may not be able to easily detect the correct pattern, and for such cases, the compiler may require a dedicated API.
[0099] In some exemplary aspects, FE 210 may implement one or more vendor extension builtins that may target a VMP-specific ISA, for example, in addition to standard OpenCL builtins that may be optimized for VMP machines.
[0100] In some exemplary aspects, FE 210 may be configured to implement OpenCL constructs and / or work-item functionality.
[0101] In some exemplary aspects, ME 220 may be configured to process LLVM IR code, which may be generic and target-independent, e.g., although it may include one or more hooks for a specific target architecture.
[0102] In some demonstrative aspects, ME 220 may perform one or more custom passes, for example, to support a VMP architecture, for example, as described below.
[0103] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of control flow graph (CFG) linearization analysis, e.g., as described below.
[0104] In some exemplary aspects, the CFG linearization analysis can be configured to linearize the code, for example, by converting if statements to select modes, for example, where the VMP vector code does not support standard control flow.
[0105] In one example, ME 220 may receive a given code, for example, as follows:
[0106]
[0107] According to this example, ME 220 may be configured to apply CFG linearization analysis to a given code, for example, as follows:
[0108]
[0109]
[0110] Example (1)
[0111] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of automatic vectorization analysis, e.g., as described below.
[0112] In some exemplary aspects, the auto-vectorization analysis may be configured to vectorize (eg, auto-vectorize) a given code, for example, to exploit the vector capabilities of the VMP.
[0113] In some exemplary aspects, ME 220 may be configured to perform automatic vectorization analysis, e.g., to vectorize code into scalar form. For example, some or all operations of automatic vectorization analysis may not be performed, such as when the code is already provided in vectorized form.
[0114] In some exemplary aspects, for example, in some use cases and / or scenarios, a compiler may not always be able to auto-vectorize code, for example, due to data dependencies between loop iterations.
[0115] In one example, ME 220 may receive a given code, for example, as follows:
[0116]
[0117] According to this example, ME 220 may be configured to perform CFG auto-vectorization analysis by applying a first transformation, for example, as follows:
[0118]
[0119] Example (2a)
[0120] For example, ME 220 may be configured to perform CFG auto-vectorization analysis by applying a second transformation, e.g., after a first transformation, e.g., as follows:
[0121]
[0122] Example (2b)
[0123] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of scratch pad memory cycle access analysis (SPMLAA), e.g., as described below.
[0124] In some exemplary aspects, the SPMLAA may define a processing block (PB), for example, that should later be outlined and compiled for VMP.
[0125] In some exemplary aspects, a processing block may include an accelerated loop that may be executed by a vector unit of a VMP.
[0126] In some exemplary aspects, a PB (eg, each PB) may include memory references. For example, some or all memory accesses may refer to a local memory bank.
[0127] In some exemplary aspects, the VMP may enable the AGU (e.g., as described below with reference to Figure 3 AGU 320) and scatter-gather unit (SG) described above are used to access the memory bank.
[0128] In some exemplary aspects, the AGU can be pre-configured, for example, before a loop is executed. For example, a loop trip count can be calculated, for example, before running a processing block.
[0129] In some exemplary aspects, image references can be created at this stage, eg, some or all image references, and strides and offsets can then be calculated, eg, per-dimension strides and offsets for each reference.
[0130] In some exemplary aspects, ME 220 may be configured to perform one or more operations of AGU planner analysis, e.g., as described below.
[0131] In some exemplary aspects, the AGU planner analysis can include an iterator specification that can cover image references from an entire processing block, eg, all image references.
[0132] In some exemplary aspects, an iterator may cover a single reference or a group of references.
[0133] In some exemplary aspects, one or more memory references may be combined via a shuffle instruction and / or reuse the same access, and / or preserve values read from a previous iteration.
[0134] In some exemplary aspects, other memory references, such as those without a linear access pattern, may be processed using a scatter-gather (SG) unit, which may have a performance penalty, such as because it may need to maintain indexes and / or masks.
[0135] In some exemplary aspects, a plan may be configured as an arrangement of iterators in a processing block. For example, a processing block may, for example, theoretically have multiple plans.
[0136] In some exemplary aspects, the AGU planner analysis can be configured to construct all possible plans for all PBs and select a combination, eg, the best combination, from among all valid combinations.
[0137] In some exemplary aspects, the total number of iterators in a valid combination may be limited, eg, not to exceed the number of available AGUs on the VMP.
[0138] In some exemplary aspects, one or more parameters may be defined for an iterator (e.g., for each iterator), e.g., including stride, width, and / or cardinality, e.g., as part of an AGU planner analysis. For example, a minimum-maximum range for an iterator may be defined dimensionally, e.g., in each dimension, e.g., as part of an AGU planner analysis.
[0139] In some exemplary aspects, the AGU planner analysis can be configured to track and evaluate memory references to the image, eg, each memory reference, eg, to understand its access pattern.
[0140] In one example, according to Example 2a / 2b, image "a" as a base address can be accessed with 64 iterations using a step size of 32 bytes.
[0141] In some exemplary aspects, LLVM can include scalar evaluation analysis (SCEV) that can compute access patterns, for example, to understand each image reference.
[0142] In some exemplary aspects, ME 220 may exploit the masking capabilities of the AGU, eg, to avoid maintaining induction variables, which may have a performance penalty.
[0143] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of rewrite analysis, e.g., as described below.
[0144] In some exemplary aspects, the rewrite analysis may be configured to transform the code of a processing block, for example, when setting up iterators and / or modifying memory access instructions.
[0145] In some exemplary aspects, the setup of iterators (e.g., all iterators) can be implemented in the IR in a target-specific intrinsic function. For example, the setup of iterators can reside in the pre-header of the outermost loop.
[0146] In some exemplary aspects, the rewrite analysis may include a loop completion analysis, eg, as described below.
[0147] In some exemplary aspects, the code may be compiled with the goal that substantially all computations should be performed within the innermost loop.
[0148] For example, loop finishing analysis may promote instructions, for example, to move operations performed after the last iteration of the loop into the loop.
[0149] For example, loop finishing analysis may sink instructions, eg, to move operations performed before the first iteration of the loop into the loop.
[0150] For example, loop finishing analysis may hoist instructions and / or sink instructions, eg, such that substantially all instructions from an outer loop are moved to an innermost loop.
[0151] For example, loop completion analysis may be configured to provide technical solutions to support VMP iterators, for example, to work only on perfectly nested loops.
[0152] For example, loop completion analysis may lead to a situation where there are no instructions between "for" statements that make up a loop, e.g., to support VMP iterators, which cannot emulate such a situation.
[0153] In some exemplary aspects, loop refinement analysis may be configured to collapse nested loops into a single collapsed loop.
[0154] In one example, ME 220 may receive a given code, for example, as follows:
[0155]
[0156] According to this example, ME 220 may be configured to perform loop completion analysis to collapse nested loops in the code into a single folded loop, for example, as follows:
[0157]
[0158] Example (3)
[0159] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of vector loop delimitation analysis, eg, as described below.
[0160] In some exemplary aspects, the vector loop demarcation analysis can be configured to partition the code between the scalar subsystem and the vector subsystem, for example, as described below with reference to Figure 3 The vector processing block 310 ( Figure 3 ) and scalar processor 330( Figure 3 )between.
[0161] In some exemplary aspects, a VMP accelerator may include scalar and / or vector subsystems, e.g., as described below. For example, each of the subsystems may have a different computational unit / processor. Accordingly, scalar code may be compiled on a scalar compiler (e.g., an SSC compiler), and / or accelerated vector code may run on a VMP vector processor.
[0162] In some exemplary aspects, vector loop demarcation analysis can be configured to create separate functions for accelerating loop bodies of vector code. For example, these functions can be marked for VMP and / or can proceed to the VMP backend, for example, while the rest of the code can be compiled by the SSC compiler.
[0163] In some exemplary aspects, one or more portions of a vector loop (e.g., configuration of a vector unit and / or initialization of vector registers) may be performed by a scalar unit. However, these portions may be performed at a later stage, e.g., by backfilling the scalar code, e.g., because the scalar code may still be in LLVM IR before being processed by the SSC compiler.
[0164] In some exemplary aspects, BE 230 may be configured to translate LLVM IR into machine instructions. For example, BE 230 may not be target agnostic and may be familiar with target specific architectures and optimizations, for example, compared to ME 220 which may be agnostic to target specific architectures.
[0165] In some exemplary aspects, BE 230 may be configured to perform one or more analyses that may be specific to the target machine (eg, a VMP machine) to which the code is being downgraded, for example, even though BE 230 may use a general-purpose LLVM.
[0166] In some exemplary aspects, BE 230 may be configured to perform one or more operations of instruction degradation analysis, eg, as described below.
[0167] In some exemplary aspects, instruction degradation analysis may be configured to translate LLVM IR into target-specific instruction machine IR (MIR), for example, by translating LLVM IR into a directed acyclic graph (DAG).
[0168] In some exemplary aspects, the DAG may undergo a legalization process for instructions, such as based on data types and / or VMP instructions, which may be supported by the VMP HW.
[0169] In some exemplary aspects, instruction demotion analysis may be configured to, for example, perform a pattern matching process after a legalization process of instructions, for example, to demotion nodes (eg, each node) in a DAG to, for example, VMP-specific machine instructions.
[0170] In some exemplary aspects, instruction degradation analysis may be configured to generate a MIR, for example, after a pattern matching process.
[0171] In some exemplary aspects, instruction demotion analysis may be configured to degrade instructions according to a machine application binary interface (ABI) and / or calling convention.
[0172] In some exemplary aspects, BE 230 can be configured to perform one or more operations of a cell balance analysis, eg, as described below.
[0173] In some exemplary aspects, the unit balancing analysis may be configured to balance instructions among VMP computing units, for example, as described below with reference to Figure 3 The data processing unit 316 ( Figure 3 )between.
[0174] In some exemplary aspects, the cell balance analysis may be aware of some or all available arithmetic transformations, and / or may perform transformations according to an optimal algorithm.
[0175] In some exemplary aspects, BE 230 may be configured to perform one or more operations of a modulo scheduler (pipeliner) analysis, eg, as described below.
[0176] In some exemplary aspects, the pipeliner may be configured to schedule instructions according to one or more constraints (e.g., data dependencies, resource bottlenecks, and / or any other constraints), for example using a swing modulo scheduling (SMS) heuristic and / or any other additional and / or alternative heuristics.
[0177] In some exemplary aspects, the pipeliner can be configured to schedule a set of very long instruction word (VLIW) instructions (eg, of initiation intervals (II)) over which a program will iterate, such as during a steady state.
[0178] In some exemplary aspects, a performance metric may be measured, which may be based on the number of cycles a typical loop may execute, for example, as follows:
[0179] (input data size in bytes)*II / (bytes consumed / produced per iteration)
[0180] In some exemplary aspects, the pipeliner can attempt to minimize II as much as possible, for example, to improve performance.
[0181] In some exemplary aspects, the pipeliner can be configured to calculate a minimum II and schedule accordingly. For example, if the pipeliner fails to schedule, the pipeliner can attempt to increase the II and retry scheduling, for example, until a predefined II threshold is violated.
[0182] In some demonstrative aspects, BE 230 may be configured to perform one or more operations of register allocation analysis, eg, as described below.
[0183] In some exemplary aspects, register allocation analysis can be configured to attempt to assign registers in an efficient (eg, optimal) manner.
[0184] In some exemplary aspects, register allocation analysis may assign values to bypass vector registers, general purpose vector registers, and / or scalar registers.
[0185] In some exemplary aspects, the values may include private variables, constants, and / or values that rotate across iterations.
[0186] In some exemplary aspects, register allocation analysis may implement an optimal heuristic that fits one or more VMP register file (regfile) constraints. For example, in some use cases, register allocation analysis may not use standard LLVM register allocation.
[0187] In some exemplary aspects, in some cases, register allocation analysis may fail, which may mean that the loop cannot be compiled. Accordingly, register allocation analysis may implement a retry mechanism that may return to the modulo scheduler and may attempt to reschedule the loop, e.g., with an increased launch interval. For example, in many cases, increasing the launch interval may reduce register starvation and / or may support compilation of vector loops.
[0188] In some demonstrative aspects, BE 230 may be configured to perform one or more operations of SSC configuration analysis, eg, as described below.
[0189] In some exemplary aspects, the SSC configuration analysis may be configured to set a configuration for executing a kernel, such as an AGU configuration.
[0190] In some exemplary aspects, SSC configuration analysis may be performed at a later stage, such as due to configurations being calculated after legalization, register allocation analysis, and / or modulo scheduling analysis.
[0191] In some exemplary aspects, the SSC configuration analysis can include a zero overhead loop (ZOL) mechanism in a vector loop. For example, the ZOL mechanism can configure loop trip counts based on access patterns of memory references in the loop, e.g., to avoid running instructions that check loop exit conditions for each iteration.
[0192] In some exemplary aspects, a VMP compilation flow may include one or more (e.g., a small number) of steps that may be called during the compilation flow in a test library (testlib) (e.g., a wrapper script for compiling, executing, and / or testing a program). For example, these steps may be performed outside of the LLVM compiler.
[0193] In some exemplary aspects, a PCB Hardware Description Language (PHDL) simulator can be implemented to perform one or more roles of an assembler, an encoder, and / or a linker.
[0194] In some exemplary aspects, compiler 200 can be configured to provide technical solutions to support robustness, which can enable compilation of a wide range of loop selections with HW limitations. For example, compiler 200 can be configured to support technical solutions that may not generate verification errors.
[0195] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support programmability, which may provide users with the ability to express code in a variety of ways that may compile correctly to a VMP architecture.
[0196] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support an improved user experience, which may allow a user to debug and / or profile code. For example, the improved user experience may provide informative error messages, reporting tools, and / or profiling tools.
[0197] In some exemplary aspects, compiler 200 can be configured to provide technical solutions to support improved performance, for example, to optimize VMP assembly code and / or iterator access, which can lead to faster execution. For example, improved performance can be achieved through high utilization computing units and using their complex CISC.
[0198] refer to Figure 3 , which schematically illustrates a vector processor 300 according to some exemplary aspects. For example, the vector processor 180 ( Figure 1 ) may implement one or more elements of the vector processor 300 and / or may perform one or more operations and / or functionalities of the vector processor 300.
[0199] In some exemplary aspects, vector processor 300 may comprise a vector microcode processor (VMP).
[0200] In some exemplary aspects, vector processor 300 may include a wide vector machine, eg, supporting a very long instruction word (VLIW) architecture and / or a single instruction / multiple data (SIMD) architecture.
[0201] In some exemplary aspects, vector processor 300 may be configured to provide a technical solution to support high performance for short integer types, which may be common in, for example, computer vision and / or deep learning algorithms.
[0202] In other aspects, the vector processor 300 may include any other type of vector processor, and / or may be configured to support any other additional or alternative functionality.
[0203] In some exemplary aspects, such as Figure 3 As shown, the vector processor 300 may include a vector processing block (vector processor) 310, a scalar processor 330, and a direct memory access (DMA) 340, for example, as described below.
[0204] In some exemplary aspects, such as Figure 3As shown, the vector processing block 310 may be configured to process (eg, efficiently process) image data and / or vector data. For example, the vector processing block 310 may be configured to use a vector computing unit, for example, to accelerate computation.
[0205] In some exemplary aspects, scalar processor 330 may be configured to perform scalar calculations. For example, scalar processor 330 may be used as "glue logic" for a program that includes vector calculations. For example, some (e.g., even most) of the calculations of a program may be performed by vector processing block 310. However, several tasks (e.g., some basic tasks) (e.g., scalar calculations) may be performed by scalar processor 330.
[0206] In some demonstrative aspects, DMA 340 may be configured to interface with one or more memory elements in a chip including vector processor 300 .
[0207] In some demonstrative aspects, DMA 340 may be configured to read input from main memory, and / or write output to main memory.
[0208] In some exemplary aspects, scalar processor 330 and vector processing block 310 may use respective local memories to process data.
[0209] In some exemplary aspects, such as Figure 3 As shown, the vector processor 300 may include an extractor and decoder 350 , which may be configured to control the scalar processor 330 and / or the vector processing block 310 .
[0210] In some exemplary aspects, operations of scalar processor 330 and / or vector processing block 310 may be triggered by instructions stored in program memory 352 .
[0211] In some demonstrative aspects, DMA 340 may be configured to transfer data in parallel with the execution of program instructions in memory 352, for example.
[0212] In some exemplary aspects, DMA 340 may be controlled by software, such as via configuration registers, rather than instructions, for example, and accordingly may be considered a second “thread” of execution in vector processor 300 .
[0213] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include one or more data processing units, e.g., a group of data processing units, e.g., as described below.
[0214] In some exemplary aspects, a data processing unit may include hardware configured to perform calculations, such as an arithmetic logic unit (ALU).
[0215] In one example, the data processing unit may be configured to add numbers and / or store numbers in memory.
[0216] In some exemplary aspects, the data processing unit may be controlled by commands encoded in, for example, program memory 352 and / or configuration registers. For example, the configuration registers may be memory mapped and writeable by memory storage commands of scalar processor 330.
[0217] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include a state configuration including a set of registers and memory, eg, as described below.
[0218] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a set of vector memories 312 , which may be configured, for example, to store data to be processed by the vector processor block 310 .
[0219] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a set of vector registers 314 that may be configured for use, for example, in data processing performed by the vector processor block 310 .
[0220] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may be associated with a set of memory maps.
[0221] In some exemplary aspects, a memory map may include a set of addresses accessible by a data processing unit that may load data from / to registers and memory and / or store data.
[0222] In some exemplary aspects, such as Figure 3 As shown, the vector processing block 310 may include a plurality of address generation units (AGUs) 320 , which may include addresses accessible to them, for example, in one or more memories in the memory 312 .
[0223] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a plurality of data processing units 316, for example, as described below.
[0224] In some exemplary aspects, the data processing unit 316 may be configured to process commands, for example, including a number of digits at a time. In one example, the command may include 8 digits. In another example, the command may include 4 digits, 16 digits, or any other count of digits.
[0225] In some exemplary aspects, two or more data processing units 316 can be used simultaneously. In one example, data processing unit 316 can process and execute multiple different commands, for example, 3 different commands, for example, including 8 numbers, in a single cycle.
[0226] In some exemplary aspects, the data processing units 316 may be asymmetric. For example, the first and second data processing units 316 may support different commands. For example, addition may be performed by the first data processing unit 316, and / or multiplication may be performed by the second data processing unit 316. For example, both operations may be performed by one or more additional data processing units 316.
[0227] In some demonstrative aspects, data processing unit 316 may be configured to support arithmetic operations for many combinations of input and output data types.
[0228] In some exemplary aspects, data processing unit 316 may be configured to support one or more operations, which may be less common. For example, processing unit 316 may support operations to work with a lookup table (LUT) of vector processor 300 and / or any other operations.
[0229] In some exemplary aspects, data processing unit 316 may be configured to support efficient computation of nonlinear functions, histograms, and / or random data access, which may, for example, facilitate implementation of algorithms like image scaling, Hough transform, and / or any other algorithm.
[0230] In some exemplary aspects, vector memory 312 may include a memory bank having a size of 16K, for example, or any other size, which may be accessed in the same cycle.
[0231] In one example, the maximum memory access size may be 64 bits. According to this example, the peak throughput may be 256 bits, for example, 64×4=256. For example, a high memory bandwidth may be achieved to utilize the computational power of the data processing unit 316 .
[0232] In one example, two data processing units 316 may support 16 8-bit multiply and accumulate operations (MACs) per cycle. According to this example, two data processing units 316 may not be useful, for example, if the input numbers are not extracted at that speed, and / or there is no input of exactly 256 bits, for example, 16x8x2=256.
[0233] In some exemplary aspects, AGU 320 may be configured to perform memory access operations, such as loading and storing data from / to vector memory 314 .
[0234] In some exemplary aspects, AGU 320 may be configured to calculate addresses of input and output data items, for example, to handle I / O in situations where high bandwidth is not sufficient to utilize data processing unit 316 .
[0235] In some exemplary aspects, AGU 320 may be configured to calculate addresses of input and / or output data items, for example, based on configuration registers written by scalar processor 330, prior to entering a vector command block (eg, a loop).
[0236] For example, the AGU 320 may be configured to write an image base pointer, width, height, and / or stride to configuration registers, for example, to iterate over an image.
[0237] In some exemplary aspects, the AGU 320 may be configured to handle addressing (e.g., all addressing), for example, to provide a technical solution in which the data processing unit 316 may not have the burden of incrementing a pointer or counter in a loop and / or the burden of checking a line end condition, for example, to zero a counter in a loop.
[0238] In some exemplary aspects, such as Figure 3 As shown, the AGU 320 may include four AGUs, and accordingly, four memories 312 may be accessed in the same cycle. In other aspects, any other count of AGUs 32 may be implemented.
[0239] In some exemplary aspects, AGUs 320 may not be "bound" to memory banks 312. For example, an AGU 320 (e.g., each AGU 320) may access a memory bank 312 (e.g., each memory bank 312), e.g., as long as two or more AGUs 320 do not attempt to access the same memory bank 312 in the same cycle.
[0240] In some demonstrative aspects, vector registers 314 may be configured to support communications between data processing unit 316 and AGU 320 .
[0241] In one example, the total number of vector registers 314 may be 28, which may be divided into several subsets, for example, based on their functions. For example, a first subset of vector registers 314 may be used for input / output of, for example, all data processing units 316 and / or AGU 320; and / or a second subset of vector registers 314 may not be used for output of some operations (e.g., most operations) and may be used for one or more other operations, for example, to store loop-invariant inputs.
[0242] In some exemplary aspects, a data processing unit 316 (e.g., each data processing unit 316) may have one or more registers to host the output of the last performed operation, e.g., which may be fed as input to other data processing units 316. For example, these registers may "bypass" vector registers 314 and may operate faster than writing these outputs to the first set of vector registers 314.
[0243] In some exemplary aspects, the extractor and decoder 350 may be configured to support low-overhead vector loops, e.g., very low-overhead vector loops (also referred to as "zero-overhead vector loops"), e.g., where a termination (exit) condition of the vector loop may not need to be checked during execution of the vector loop.
[0244] For example, the AGU 320 may signal a termination (exit) condition, such as when the AGU 320 completes iterations over a configured memory region.
[0245] For example, the fetcher and decoder 350 may exit the loop when, for example, the AGU 320 signals a termination condition.
[0246] For example, the scalar processor 330 may be utilized to configure loop parameters, such as the first and last instructions and / or exit conditions.
[0247] In one example, vector loops may be utilized, for example, together with high memory bandwidth and / or cheap addressing, for example, to solve control and data flow problems, for example, to provide a technical solution to allow data processing unit 316 to process data with substantially no additional overhead.
[0248] In some exemplary aspects, scalar processor 330 may be configured to provide one or more functionalities that may be complementary to the functionality of vector processing block 310. For example, a large portion (e.g., most) of the work in a vector program may be performed by data processing unit 316. For example, scalar processor 330 may be utilized, for example, to "glue" together various vector code blocks of a vector program.
[0249] In some exemplary aspects, the scalar processor 330 may be implemented separately from the vector processing block 310. In other aspects, the scalar processor 330 may be configured to share one or more components and / or functionality with the vector processing block 310.
[0250] In some exemplary aspects, scalar processor 330 may be configured to perform operations that may not be suitable for execution on vector processing block 310 .
[0251] For example, the scalar processor 330 may be utilized to execute a 32-bit C program. For example, the scalar processor 330 may be configured to support 1, 2, and / or 4-byte data types of the C code and / or some or all arithmetic operators of the C code.
[0252] For example, the scalar processor 330 may be configured to provide a technical solution to perform operations that cannot be performed on the vector processing block 310 without, for example, using a full CPU.
[0253] In some exemplary aspects, scalar processor 330 may include a scalar data memory 332 , for example, having a size of 16K or any other size, which may be configured to store data, such as variables used by a scalar portion of a program.
[0254] For example, scalar processor 330 may store local and / or global variables declared by portable C code, which may be compiled by a compiler (e.g., compiler 200 ( Figure 2 )) is allocated to scalar data memory.
[0255] In some exemplary aspects, such as Figure 3 As shown, the scalar processor 330 may include or may be associated with a set of vector registers 334 that may be used for data processing by the scalar processor 330 .
[0256] In some exemplary aspects, the scalar processor 330 can be associated with a scalar memory map that can enable the scalar processor 330 to access substantially all states of the vector processor 300. For example, the scalar processor 330 can configure a vector unit and / or a DMA channel via the scalar memory map.
[0257] In some exemplary aspects, the scalar processor 330 may not be allowed to access one or more block control registers that may be used by an external processor to run and debug a vector program.
[0258] In some exemplary aspects, DMA 340 can be configured to communicate, for example, via main memory, with one or more other components of a chip implementing vector processor 300. For example, DMA 340 can be configured to transfer blocks of data, for example, large, contiguous blocks of data, for example, to support scalar processor 330 and / or vector processing blocks that can manipulate data stored in local memory. For example, a vector program may be able to use DMA 340 to read data from main chip memory.
[0259] In some exemplary aspects, DMA 340 may be configured to communicate with other elements of the chip, for example, via a plurality of DMA channels (e.g., 8 DMA channels or any other count of DMA channels). For example, a DMA channel (e.g., each DMA channel) may be able to transfer a rectangular patch from a local memory to a main chip memory, or vice versa. In other aspects, a DMA channel may transfer any other type of data block between a local memory and a main chip memory.
[0260] In some exemplary aspects, a rectangular tile may be defined by a base pointer, a width, a height, and a stride.
[0261] For example, at peak throughput, 8 bytes may be transferred per cycle, however, there may be an overhead for each tile and / or for each row in a tile.
[0262] In some exemplary aspects, DMA 340 can be configured to transfer data in parallel with computations, such as via multiple DMA channels, for example, as long as the executed commands do not access local memory involved in the transfer.
[0263] In one example, since all channels can access the same memory bus, using several channels to implement a transfer may not save I / O cycles, for example, compared to when a single channel is used. However, multiple DMA channels can be utilized to schedule several transfers and execute them in parallel with the calculation. For example, this may be advantageous compared to a single channel, which may not allow a second transfer to be scheduled before the first transfer is completed.
[0264] In some exemplary aspects, DMA 340 can be associated with a memory map that can support DMA channels accessing vector memory and / or scalar data. For example, access to vector memory can be performed in parallel with computation. For example, access to scalar data may not generally allow for parallelism, e.g., because scalar processor 330 may be involved in almost any reasonable program and may access its local variables while performing a transfer, which may result in memory contention with active DMA channels.
[0265] In some exemplary aspects, DMA 340 can be configured to provide a technical solution to support parallelization of I / O and computation. For example, a program performing computations may not have to wait for I / O, for example, when these computations can be run quickly by vector processing block 310.
[0266] In some exemplary aspects, an external processor (eg, a CPU) may be configured to initiate execution of a program on vector processor 300. For example, vector processor 300 may remain idle, for example, as long as program execution is not initiated.
[0267] In some exemplary aspects, the external processor may be configured to debug the program, for example, to execute a single step at a time, to stop when the program reaches a breakpoint, and / or to examine the contents of registers and memory storing program variables.
[0268] In some exemplary aspects, external memory mapping may be implemented to enable an external processor to control the vector processor 300 and / or a debugger, for example, by writing to control registers of the vector processor 300 .
[0269] In some exemplary aspects, the external memory map may be implemented by a superset of the scalar memory map. For example, the implementation may make all registers and memories defined by the architecture of the vector processor 300 accessible to a debugger backend running on an external processor.
[0270] In some exemplary aspects, the vector processor 300 may issue an interrupt signal, such as when the vector processor 300 terminates a program.
[0271] In some exemplary aspects, the interrupt signal may be used, for example, to implement a driver to maintain a queue of programs scheduled for execution by vector processor 300 and / or may be used to start a new program, for example, by an external processor, when a previously executed program completes.
[0272] Return to reference Figure 1 In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on one or more loops, which may be based on, for example, source code 112, for example, as described below.
[0273] In some exemplary aspects, compiler 160 may be configured to compile one or more operations according to a compilation scheme that may be configured to provide a technical solution to reduce or even eliminate the use of induction variables, such as in one or more masked memory access operations in a loop, e.g., as described below.
[0274] In some exemplary aspects, the one or more masked memory access operations in the loop may include a masked load operation, a masked store operation, a masked select operation, and / or any other masked operation to access memory, eg, as described below.
[0275] In some exemplary aspects, a masked memory access instruction may include a conditional operation based on a mask expression including a logical condition, eg, as described below.
[0276] In some exemplary aspects, conditional operation of masked memory access instructions may be performed, such as based on whether a mask expression is true or false.
[0277] In some demonstrative aspects, a mask expression may include one or more Boolean conditions that may be based on and / or may represent, for example, a result of a condition.
[0278] In some demonstrative aspects, a mask expression may include one or more mask leaves corresponding to one or more Boolean conditions.
[0279] In some exemplary aspects, the mask leaves may correspond to respective Boolean conditions, eg, as described below.
[0280] In one example, a loop may include a mask expression, which may be defined based on one or more Boolean conditions, for example, as follows:
[0281] char Mask = (a>x)|(b<10);
[0282] For example, the mask expression may include a first leaf and a second leaf. For example, the first leaf may include a first Boolean condition (a>x), and the second leaf may include a second Boolean condition (b<10).
[0283] According to this example, the result of the mask expression may be true, for example, when a first Boolean condition (a>x) of the first leaf is true and / or a second Boolean condition (b<10) of the second leaf is true.
[0284] In some exemplary aspects, for example, in some use cases, scenarios, and / or implementations, when implementing mask operations in a loop, one or more technical issues may need to be addressed, for example, as described below.
[0285] In some exemplary aspects, eg, in some use cases, scenarios, and / or implementations, calculation of some masks in a loop may be computationally complex and / or expensive.
[0286] For example, computation of a mask based on an induction variable (IV) of a loop that includes a mask (also referred to as "induction-based masking") may require maintaining the IV for the loop, comparing the IV to a bound, preserving a mask register, and / or one or more additional or alternative operations based on the IV.
[0287] In one example, an induction variable of a loop may include a variable that may, for example, increase or decrease by a fixed amount at each iteration of the loop. In another example, an induction variable of a loop may be a function of another induction variable of the loop, for example, a linear function.
[0288] For example, in some use cases, scenarios, and / or implementations, some loop transformations (eg, loop vectorization transformations) may introduce one or more induction-based masks, eg, to filter out outliers / computations.
[0289] For example, the filtering may be performed, for example, by a masking operation that may be configured to select, for example, between a loaded / computed value and some default value according to a mask based on the IV.
[0290] For example, in some use cases, scenarios, and / or implementations, computing mask operations may be computationally "painful", such as when implemented by one or more target processor architectures that may not have efficient means to maintain generalizations.
[0291] In one example, some target architectures may have other hardware (HW) mechanisms to control the execution of loops, and to control bounds on memory accesses ("bounded loads / stores").
[0292] In some exemplary aspects, compiler 160 may be configured to identify one or more masked memory access operations based on source code 112 and compile the identified masked memory access operations according to a mask operation compilation scheme, eg, as described below.
[0293] In some exemplary aspects, compiler 160 may be configured to generate target code 115 by compiling source code 112 , for example, according to a mask operation compilation scheme, e.g., as described below.
[0294] In some exemplary aspects, the identified masked memory access operations may be based on a mask expression, eg, as described below.
[0295] In some exemplary aspects, the mask operation compilation scheme can be configured to provide a technical solution to reduce, eliminate, optimize, and / or exclude the use of induction variables in identified masked memory access operations, eg, as described below.
[0296] In some exemplary aspects, the mask operation coding scheme may be configured to provide a technical solution to exclude one or more mask leaves from a mask expression, which one or more mask leaves may be based on IV (IV-based mask leaves), for example, as described below.
[0297] In some exemplary aspects, the mask operation coding scheme may be configured to provide a technical solution to maintain one or more mask leaves in a mask expression that may not be based on IV (non-IV based mask leaves), for example, as described below.
[0298] In some exemplary aspects, the mask operation compilation scheme can be configured to provide a technical solution to generate target code 115, for example, based on masked memory access operations that may not be based on and / or may not require processing of induction variables, for example, as described below.
[0299] In some exemplary aspects, the mask operation compilation scheme may be configured to provide a technical solution that may improve the performance of an executed program (e.g., an image processing program), for example, by excluding one or more (e.g., some or all) IV-based mask leaves from a mask expression, for example, as described below.
[0300] In some exemplary aspects, the mask operation coding scheme may be configured to convert a first mask expression into a second mask expression, for example, by reconfiguring the first mask expression, eg, as described below.
[0301] In some demonstrative aspects, the second mask expression may be configured to simplify or exclude one or more mask leaves of the first mask expression, eg, as described below.
[0302] In some exemplary aspects, the mask operation coding scheme may be configured to convert the first mask expression into a logical form including a first logical expression portion and a second logical expression portion, eg, as described below.
[0303] In some exemplary aspects, the first logical expression portion may be based on one or more mask leaves of the first mask expression, eg, as described below.
[0304] In some exemplary aspects, the first logical expression portion may include an additional expression (Expr1) that may be obtained from the mask expression, for example, by replacing one or more mask leaves of the first mask expression and / or one or more don't care leaves of the first mask expression with, for example, constants 0 or 1, for example, as described below.
[0305] In one example, the mask operation coding scheme may be configured to convert the original mask expression into, for example, a logical form, which may be represented, for example, as follows:
[0306] P1&~P2&P3&…&Expr1
[0307] Wherein P1, P2, P3 represent mask leaves of the original mask expression, and expression Expr1 can be obtained from the original mask expression, for example, by replacing mask leaves P1, P2 and / or P3 and / or one or more "don't care" mask leaves with, for example, constants "0" or "1".
[0308] In some exemplary aspects, the mask operation coding scheme may be configured to build a truth table for the first mask expression, eg, as described below.
[0309] In some exemplary aspects, the mask operation coding scheme may be configured to build a truth table for the first mask expression, for example, based on mask leaves of the first mask expression, eg, as described below.
[0310] In some demonstrative aspects, the mask operation coding scheme may be configured to identify one or more mask leaves to be simplified and / or replaced, e.g., based on a truth table, e.g., as described below.
[0311] For example, the identified mask leaves may be represented by one or more AND leaves of a logical form corresponding to the first mask expression, eg, as described below.
[0312] In some exemplary aspects, the mask operation coding scheme may be configured to replace one or more of the identified mask leaves with iterator min / max parameters, eg, as described below.
[0313] For example, a mask operation compilation scheme may be configured to assign an identified mask leaf (eg, the identified mask leaf may be represented by a literal P1&~P2&P3 in logical form) to, for example, one or more AGUs, for example, as described below.
[0314] In some exemplary aspects, the mask operation compilation scheme may be configured to generate a second mask expression, for example, based on a second logical expression portion (e.g., expression Expr1 in logical form), which second logical expression portion may be retained, for example, after the identified mask leaf is assigned to the AGU, for example, as described below.
[0315] In some exemplary aspects, compiler 160 may be configured to identify, for example, a first mask memory access operation in a loop based on source code 112, for example, as described below.
[0316] In some exemplary aspects, the first mask memory access operation may be based on a first mask expression including one or more mask leaves, eg, as described below.
[0317] In some demonstrative aspects, compiler 160 may be configured to determine an identified mask leaf among the one or more mask leaves of the first mask expression, eg, based on at least one predefined criterion, eg, as described below.
[0318] In some exemplary aspects, compiler 160 may be configured to determine an identified mask leaf among one or more mask leaves of the first mask expression, for example based on criteria related to the impact of the identified mask leaf on the logical value of the first mask expression, for example, as described below.
[0319] In some exemplary aspects, compiler 160 may be configured to determine the second masked memory access operation by, for example, reconfiguring the first masked memory access operation based on identified mask leaves of the first mask expression, eg, as described below.
[0320] In some exemplary aspects, the second mask memory access operation can be based on a second mask expression, eg, as described below.
[0321] In some exemplary aspects, the second mask expression may be logically simplified, eg, as compared to the first mask expression, eg, as described below.
[0322] In some exemplary aspects, compiler 160 may be configured to configure the second mask memory access operation, e.g., such that the count of mask leaves in the second mask expression is less than the count of mask leaves in the first mask expression of the first mask memory access operation, e.g., as described below.
[0323] In some exemplary aspects, the second masked memory access operation may include a masked load operation to conditionally load a value from memory according to a second mask expression, eg, as described below.
[0324] In some exemplary aspects, the second masked memory access operation may include a masked store operation to conditionally load a value into memory according to a second mask expression, eg, as described below.
[0325] In some exemplary aspects, the second mask memory access operation can include any other additional or alternative types of mask operations.
[0326] In some exemplary aspects, compiler 160 may be configured to generate object code 115 , for example, based on compiling source code 112 , for example, as described below.
[0327] In some exemplary aspects, the target code 115 may perform memory access operations based on, for example, the second mask, for example, as described below.
[0328] In some exemplary aspects, compiler 160 may be configured to generate target code 115 configured for execution by, for example, a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor (eg, processor 180).
[0329] In other aspects, compiler 160 may be configured to generate object code 115 configured for execution by, for example, any other suitable type of processor.
[0330] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on source code 112 including Open Computing Language (OpenCL) code, for example.
[0331] In other aspects, compiler 160 may be configured to generate object code 115 based on source code 112 , including any other suitable type of code, for example.
[0332] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into target code 115 , for example, according to a low-level virtual machine (LLVM)-based compilation scheme.
[0333] In other aspects, compiler 160 may be configured to compile source code 112 into target code 115 according to any other suitable compilation scheme.
[0334] In some exemplary aspects, the identified mask leaves may be based on, for example, an induction variable of a loop, eg, as described below.
[0335] In some exemplary aspects, compiler 160 may be configured to determine the identified mask leaf, for example, based on a determination that the identified mask leaf is based on an IV of a loop that includes a first masked memory access operation, eg, as described below.
[0336] In some exemplary aspects, compiler 160 may be configured to configure an AGU instruction, for example, based on the identified mask leaf, eg, as described below.
[0337] In some exemplary aspects, the AGU instruction may include a memory access instruction of the AGU to perform a second mask memory access operation, eg, as described below.
[0338] In some exemplary aspects, compiler 160 may be configured to generate object code 115 based on AGU instructions, for example, as described below.
[0339] In some exemplary aspects, the AGU instruction may include at least one of a lower bound and / or an upper bound of a memory access range to be applied by the AGU for the second masked memory access operation, eg, as described below.
[0340] In some exemplary aspects, compiler 160 may be configured to configure AGU instructions to define memory access ranges, eg, based on identified mask leaves, eg, as described below.
[0341] In some exemplary aspects, compiler 160 may be configured to configure the AGU instruction to configure the AGU to apply the memory access range for the second mask memory access operation, eg, as described below.
[0342] In some exemplary aspects, compiler 160 may be configured to configure an AGU instruction to define a lower bound of a memory access range, eg, based on an identified mask leaf, eg, as described below.
[0343] In some exemplary aspects, compiler 160 may be configured to configure an AGU instruction to define an upper bound of a memory access range, eg, based on an identified mask leaf, eg, as described below.
[0344] In some exemplary aspects, the second mask expression may exclude the identified mask, eg, as described below.
[0345] In some exemplary aspects, compiler 160 may be configured to configure the second masked memory access operation to exclude one or more IV-based mask leaves of the first masked memory access operation, the one or more IV-based mask leaves being based on the IV of a loop that includes the first masked memory access operation, e.g., as described below.
[0346] In some exemplary aspects, compiler 160 may be configured to configure the second masked memory access operation to exclude any IV-based mask leaves of the first masked memory access operation that are based on an IV of a loop that includes the first masked memory access operation, e.g., as described below.
[0347] In some exemplary aspects, compiler 160 may be configured to configure the second masked memory access operation to include, for example, only non-IV based mask leaves that are not based on the IV of the loop including the first masked memory access operation, e.g., as described below.
[0348] In other aspects, compiler 160 may be configured to configure the second masked memory access operation to exclude only some of the IV-based mask leaves of the first masked memory access operation.
[0349] In some exemplary aspects, the compiler 160 may be configured to configure the second masked memory access operation to maintain, for example, one or more (e.g., some or all) non-IV-based mask leaves of the first masked memory access operation, where the one or more non-IV-based mask leaves are not based on the IV of the loop that includes the first masked memory access operation, e.g., as described below.
[0350] In some exemplary aspects, compiler 160 may be configured to determine the identified mask leaf of the first mask expression, for example based on criteria, which may include the requirement that all possibilities of true logical values of the first mask expression can be produced by the same logical value of the identified mask leaf, for example, as described below.
[0351] In some exemplary aspects, compiler 160 may be configured to determine the identified mask leaf of the first mask expression, for example based on criteria, which may include the requirement that all possibilities of true logical values of the first mask expression may be independent of the logical value of the identified mask leaf, for example, as described below.
[0352] In some demonstrative aspects, compiler 160 may be configured to determine the identified mask leaves of the first mask expression, eg, based on any other additional or alternative criteria.
[0353] In some exemplary aspects, the compiler 160 may be configured to determine the identified mask leaves based on, for example, a truth table corresponding to a first mask expression, as described below, for example.
[0354] In one instance, the compiler 160 may compile the source code 112 of a program to be executed by a target processor (e.g., processor 180, e.g., a target vector processor).
[0355] For example, the compiler 160 may identify a loop that includes a first masked memory access operation, as follows, for example:
[0356]
[0357]
[0358] Example (4a)
[0359] As shown in Example 4a, the first masked memory access operation may include a first masked load operation, e.g., charval = masked_load(inp2 + index, 0, Mask), which may be based on a mask, e.g., Mask.
[0360] As shown in Example 4a, the first masked load operation may be based on a first mask expression, e.g., char Mask = ~((x < a) | (~(s >= b) & (t <= 10))).
[0361] As shown in Example 4a, the first mask expression may include three mask leaves. For example, the first mask expression may include a first mask leaf (e.g., (x < a)), a second mask leaf (e.g., (s >= b)), and / or a third mask leaf (e.g., (t <= 10)).
[0362] As shown in Example 4a, the first mask leaf may include an IV-based mask leaf that may be based on an induction variable x of a loop that includes the first masked load operation.
[0363] As shown in Example 4a, the second and third mask leaves may include non-IV-based mask leaves, e.g., because they may not be based on any induction variable of a loop that includes the first masked load operation.
[0364] For example, as shown in Example 4a, the second mask leaf may be based on variables s and b, and / or the third mask leaf may be based on variable t.
[0365] In some exemplary aspects, the compiler 160 may be configured to identify the first masked load operation in a loop based on, for example, the source code 112.
[0366] In some demonstrative aspects, compiler 160 may be configured to determine the first mask leaf as an identified mask leaf of the first mask expression, eg, based on a determination that the first mask leaf is an IV-based leaf, eg, as described below.
[0367] In some exemplary aspects, the compiler 160 may be configured to determine the first mask leaf as the identified mask leaf of the first mask expression, for example based on a criterion requiring that all possibilities of true logical values of the first mask expression can be produced by the same logical value of the first mask leaf, for example, as described below.
[0368] In some demonstrative aspects, compiler 160 may be configured to determine whether a criterion is satisfied with respect to a mask leaf of the first mask expression, eg, based on a truth table corresponding to the first mask expression, eg, as described below.
[0369] In some exemplary aspects, the first mask expression may be represented by three mask leaves, e.g., as follows:
[0370] ~(P|(~Q&R))
[0371] Expression (1)
[0372] Where P = (x<a),Q=(s> =b), and R=(t<=10).
[0373] In some exemplary aspects, compiler 160 may be configured to determine a truth table corresponding to expression (1), for example, as follows:
[0374] P Q R Mask 0 0 0 1 0 1 0 1 0 1 1 1
[0375] Table (1)
[0376] As shown in truth table (1), the mask leaf P may be required to always be 0, for example, so that Mask = True. For example, as shown in truth table (1), the mask leaf P may be unchanged in the truth table, for example, the value of the mask leaf P may be 0 in all entries of the truth table.
[0377] For example, based on this determination, the mask leaf P may be extracted outside the logical expression with a minus sign (logical NOT), for example, while replacing the mask leaf P with a constant logical "0", for example, as follows:
[0378] ~P&~(0|(~Q&R)).
[0379] Expression (2)
[0380] As shown in truth table (1), the values of other mask leaves (e.g., leaves Q and R) may vary in truth table (1). For example, these mask leaves may not be constant or not care. Accordingly, these mask leaves Q and R may still be in the logic expression.
[0381] As shown in expression (2), expression (2) may include a constant value, such as 0, for example, instead of mask leaf P. For example, when the sign of mask leaf P is negative, the constant value zero may be selected.
[0382] For example, expression (2) may be logically equivalent to expression (1). For example, when P=0, the second part (~(0|(~Q&R))) of expression (2) may be completely equal to expression (1), and the first part (~P) of expression (2) may not change the second part of expression (2), for example, because (~P)=1. For example, expression (2) may be a logical AND operation of the first part of expression (2) and a logical value "1", for example, because ~P=1.
[0383] For example, when P=1, the logical NOT operation may be equal to zero, for example, ~P=0. Accordingly, expression (2) may be equal to zero. For example, as shown in truth table (1), when P=1, expression (1) may be zero.
[0384] In some exemplary aspects, compiler 160 may be configured to further simplify expression (2), for example, based on the following equation: “0|A=A”.
[0385] For example, compiler 160 may be configured to further simplify the expression (~(0|(~Q&R))) to the expression (Q&~R). For example, compiler 160 may be configured to optimize (e.g., standardly optimize) the other mask leaves to (Q|~R)=(s>=b)|~(t<=10).
[0386] For example, compiler 160 may be configured to further simplify expression (2), such as by an instcombine operation, after the pass.
[0387] In some exemplary aspects, compiler 160 may be configured to utilize truth tables to provide a technical solution to support determining a logical expression (e.g., expression (2)) based on a first mask expression (e.g., expression (1)) (e.g., even with respect to relatively complex mask expressions).
[0388] In some exemplary aspects, truth tables may be implemented to provide a technical solution to support determining a simplified logical expression based on a mask expression (e.g., even with respect to a mask expression that may not be easily simplified, e.g., using one or more transformations, e.g., according to de-Morgan's laws and / or any other transformation rules and / or laws).
[0389] In one example, the mask expression may include the following mask with additional mask leaves A and B:
[0390] Mask2=(~(P|(~Q&R))&A)|(~(P|(~Q&R))&~A&B)|(~(P|(~Q&R))&
[0391] ~A & ~B)
[0392] Expression (3)
[0393] For example, a truth table may be utilized to determine that the logic expression ˜P&˜(0|(˜Q&R)) may be logically equivalent to mask Mask2, for example, even though mask Mask2 may not be easily transformable into a simplified logic expression, for example based on De Morgan's rules.
[0394] In some exemplary aspects, compiler 160 may be configured to determine the second masked memory access operation by, for example, reconfiguring the first masked load operation based on the first mask leaf P.
[0395] In some exemplary aspects, the second mask memory access operation can be based on a second mask expression, which can be logically simplified, eg, compared to the first mask expression.
[0396] For example, the second mask expression may be based on a second portion of expression (2) which may be logically simplified, for example, compared to expression (1), eg, ˜(0|(˜Q&R)).
[0397] In some exemplary aspects, compiler 160 may be configured to configure AGU instructions, including memory access instructions of the AGU, to perform a second mask memory access operation, eg, as described below.
[0398] In some exemplary aspects, compiler 160 may configure the AGU instruction based on at least one mask leaf of the first mask memory access operation, eg, as described below.
[0399] In some exemplary aspects, compiler 160 may configure the AGU instruction, e.g., to selectively restrict memory access of a second masked memory access operation, e.g., based on at least one mask leaf of the first masked memory access operation to be excluded from the second masked memory access operation, e.g., as described below.
[0400] In some exemplary aspects, compiler 160 may configure the AGU instruction, for example, to selectively limit memory access of a second masked memory access operation, for example, in the following manner: at least one mask leaf that is logically equivalent to the first masked memory access operation is to be excluded from the second masked memory access operation, for example, as described below.
[0401] In some exemplary aspects, compiler 160 may configure the AGU instruction based on a first mask leaf of a first mask memory access operation, eg, as described below.
[0402] In some exemplary aspects, compiler 160 may configure an AGU instruction, for example, to selectively restrict memory access of a second masked memory access operation, for example, based on a first mask leaf of the first masked memory access operation, eg, as described below.
[0403] In some exemplary aspects, compiler 160 may be configured to configure the AGU instruction to define a memory access range to be accessed by the second masked memory access operation, eg, as described below.
[0404] In some exemplary aspects, compiler 160 may be configured to configure an AGU instruction to define a memory access range, for example, based on a first mask leaf of a first mask memory access operation, eg, as described below.
[0405] In some exemplary aspects, compiler 160 may be configured to configure the AGU instruction to selectively restrict the second masked memory access operation to a memory access range that may be based on a first mask leaf of the first masked memory access operation, eg, as described below.
[0406] In some exemplary aspects, compiler 160 may be configured to configure the AGU instruction to selectively restrict the second mask memory access operation to a memory access range, for example, in the following manner: the first mask leaf that may be logically equivalent to the first mask memory access operation, for example, as described below.
[0407] In some exemplary aspects, compiler 160 may be configured to configure the AGU instructions to define a lower bound and / or an upper bound of a memory access range to be applied by the AGU for the second masked memory access operation, eg, as described below.
[0408] For example, the compiler 160 may configure a lower bound and / or an upper bound of a memory access range, for example, based on the first mask leaf P, for example, as described below.
[0409] In some exemplary aspects, compiler 160 may be configured to identify a first mask leaf in a first mask expression of a first masked memory access operation in a loop, e.g., based on determining the first mask leaf based on an induction variable of the loop (e.g., induction variable x), e.g., as described below.
[0410] In some exemplary aspects, compiler 160 may be configured to configure the second masked memory access operation, for example, by reconfiguring the first masked memory access operation based on the first mask leaf, eg, as described below.
[0411] In some exemplary aspects, compiler 160 may be configured to configure AGU instructions to define a lower bound (e.g., lower bound xmin) and / or an upper bound (e.g., upper bound xmax) of a memory access range to be applied by the AGU for the second masked memory access operation, e.g., as described below.
[0412] In some exemplary aspects, compiler 160 may be configured to configure the AGU instructions to define a lower bound (e.g., lower bound xmin) and / or an upper bound (e.g., upper bound xmax), for example, based on a logical condition regarding an induction variable (e.g., induction variable x) that may be defined by a first mask leaf, for example, as described below.
[0413] In some exemplary aspects, compiler 160 may be configured to translate the first mask leaf ~P=(x>=a) in expression (2) into a bound, for example, for an AGU memory access instruction (eg, for a second mask memory access operation).
[0414] For example, compiler 160 may be configured to configure an AGU lower bound (xmin) for AGU memory access instructions in a compiled loop that may be based on the loop of Example 4a, eg, as described below.
[0415] For example, compiler 160 may be configured to configure an AGU lower bound (xmin) for AGU memory access instructions, for example, based on the first mask leaf ˜P=(x>=a) in expression (2).
[0416] For example, the compiler 160 may be configured to configure the AGU lower bound (xmin) of the AGU memory access instruction, for example by setting the AGU lower bound (xmin) to a value a (xmin=a), which may be logically equivalent to the condition of the first mask leaf ~P=(x>=a) in expression (2).
[0417] For example, compiler 160 may be configured to configure AGU instructions that may be based on AGU memory access instructions in a compiled loop of the loop of Example 4a, for example, to define an AGU lower bound (xmin=a), which may be based on the first mask leaf ˜P=(x>=a) in expression (2), for example, as follows:
[0418] agu1=allocate_agu("load");
[0419] set_base(agu1,inp1)
[0420] / / ...other agu1 parameters
[0421] agu2=allocate_agu("load");
[0422] set_base(agu2,inp2)
[0423] set_x_minmax(agu2,a,width);
[0424] / / ...other agu2 parameters
[0425] agu3=allocate_agu("store");
[0426] set_base(agu3,out);
[0427] / / ...other agu3 parameters
[0428] cycle:
[0429] char s = agu1.load();
[0430] char t = agu1.load();
[0431] char NewMask=(s>=b)|(t>10);
[0432] char val=agu2.masked_load(0,NewMask);
[0433] agu3.store(val+7);
[0434] Example (4b)
[0435] As shown in Example 4b, the compiled loop may include masked memory access operations that may include masked load operations based on a second mask (NewMask), eg, char val = agu2.masked_load(0, NewMask).
[0436] As shown in Example 4b, the masked load operation can be based on a second mask, such as NewMask, which can be defined based on a second mask expression, such as char NewMask=(s>=b)|(t>10).
[0437] As shown in Example 4b, the second mask expression may be based on a second portion of expression (2) which may be logically simplified, for example, compared to expression (1).
[0438] As shown in Example 4b, the second mask expression may include only two mask leaves, eg, two non-IV-based mask leaves, eg, ~(s>=b) and t(t<=10).
[0439] As shown in Example 4b, the second mask expression may include only non-IV-based mask leaves.
[0440] As shown in Example 4b, the second mask expression may exclude any IV-based mask leaves of mask expression (1).
[0441] As shown in Example 4b, the second mask expression may exclude any IV-based mask leaves.
[0442] As shown in Example 4b, the second mask expression may not include the first mask leaf P of expression (1). For example, the second mask expression may exclude the first mask leaf P, which may be based on the induction variable x.
[0443] In some exemplary aspects, as shown in Example 4b, compiler 160 may designate a first AGU (eg, agu1), for example, to load data from a first input pointer (eg, inp1).
[0444] In some exemplary aspects, as shown in Example 4b, compiler 160 may generate a compiled loop to include a first load instruction, eg, char s=agul.load(), eg, to load a character value based on pointer inp1[index] into a character variable s by agul.
[0445] In some exemplary aspects, as shown in Example 4b, compiler 160 may generate a compiled loop to include a second load instruction, eg, char t=agul.load(), eg, to load a character value based on pointer inp1[index+1] into a character variable t by agul.
[0446] In some exemplary aspects, as shown in Example 4b, compiler 160 may designate a second AGU (eg, agu2), for example, to perform a masked load operation, for example, to load data from a second input pointer (eg, inp2) according to a mask NewMask.
[0447] In some exemplary aspects, as shown in Example 4b, the compiler 160 may configure AGU instructions for the second AGU, e.g., to set the lower and upper bounds of the dimensions of agu2 corresponding to the induction variable x.
[0448] In some exemplary aspects, as shown in Example 4b, the compiler 160 may set the lower and / or upper bounds of the second AGU that is used to perform a masked load instruction, e.g., based on the first leaf of expression (1) (e.g., ~P = ~(x < a)).
[0449] In one example, setting the lower bound for the masked load instruction to be performed by the second AGU according to the condition (x >= a) may be logically equivalent to the leaf ~P = ~(x < a).
[0450] In some exemplary aspects, as shown in Example 4b, the compiler 160 may generate an AGU instruction, e.g., set_x_minmax(agu2,a,width), to set the lower bound of the dimensions of agu2 corresponding to the induction variable x to a.
[0451] For example, the AGU instruction (e.g., set_x_minmax(agu2,a,width)) may restrict agu2, e.g., to load data from the pointer inp2 only when the IV x is equal to or greater than a. For example, this restriction may be according to the condition of the first leaf of expression (1), e.g., ~P = ~(x < a) = (x >= a).
[0452] In some exemplary aspects, as shown in Example 4b, the compiler 160 may set the upper bound of the AGU instruction for agu2 to width, e.g., to load data from the pointer inp2, e.g., according to the condition of the IV x in the inner loop of Example 4a, e.g., for(int x = 0; x < width; x++).
[0453] In some exemplary aspects, as shown in Example 4b, for example, when agu2 is configured according to the instruction set_x_minmax(agu2,a,width), the first masked leaf of expression (1) may become redundant.
[0454] For example, it may not be necessary to mask NewMask to calculate the condition ~(x < a) = (x >= a), e.g., because this condition may already be maintained by setting the lower bound of agu2 (xmin = a).
[0455] In some exemplary aspects, as shown in Example 4b, the compiler 160 may exclude the masked leaf P from the mask NewMask.
[0456] For example, as shown in Example 4b, the masked load instruction char val=agu2.masked_load(0,NewMask) can be based on the mask NewMask.
[0457] In some exemplary aspects, as shown in Example 4b, the compiler 160 may generate a compiled loop that may be configured to define a second mask expression based on, for example, a mask NewMask, for example, char NewMask=(s>=b)|(t>10), which may exclude the first mask leaf P.
[0458] In some exemplary aspects, as shown in Example 4b, compiler 160 may generate a compiled loop to include a third load instruction (e.g., char val=agu2.masked_load(0,NewMask)), for example, to load a character value based on pointer inp2 into a character variable val by agu2, for example based on mask NewMask.
[0459] In some exemplary aspects, as shown in Example 4b, compiler 160 can designate a third AGU (eg, agu3), for example, to store data based on an output pointer (eg, out).
[0460] In some exemplary aspects, as shown in Example 4b, compiler 160 may generate a compiled loop to include a store instruction (eg, agu3.store(val+7)), eg, to store the result of summing val+7 by agu3 to output pointer out.
[0461] refer to Figure 4 , which schematically illustrates a method of compiling code for a processor. For example, Figure 4 One or more operations of the method may be performed by: a system, for example, system 100 ( Figure 1 ); devices, for example, device 102 ( Figure 1 ); a server, for example, server 170 ( Figure 1 ); and / or a compiler, for example, compiler 160 ( Figure 1 ) and / or compiler 200( Figure 2 ).
[0462] In some exemplary aspects, as indicated at block 402, the method may include identifying a first mask memory access operation based on the source code, wherein the first mask memory access operation is based on a first mask expression including one or more mask leaves. For example, the compiler 160( Figure 1 ) may be configured, for example, based on source code 112 ( Figure 1 ) to identify the first mask memory access operation in a loop operation, for example, as described above.
[0463] In some exemplary aspects, as indicated at block 404, the method may include determining a second masked memory access operation, for example, by reconfiguring the first masked memory access operation based on an identified mask leaf in the one or more mask leaves. For example, the second masked memory access operation may be based on a second mask expression that is logically simplified compared to the first mask expression. For example, compiler 160( Figure 1 ) may be configured to determine a second masked memory access operation by, for example, reconfiguring a first masked memory access operation based on an identified mask leaf, e.g., as described above.
[0464] In some exemplary aspects, as indicated at block 406, the method may include generating target code based on compiling the source code, wherein the target code is based on the second mask memory access operation. For example, compiler 160 ( Figure 1 ) can be configured to generate target code 115 ( Figure 1 ), for example, the target code is based on a second mask memory access operation, for example, as described above.
[0465] refer to Figure 5 , which schematically illustrates an article of manufacture 500 according to some exemplary aspects. Article 500 may include one or more tangible computer-readable ("machine-readable") non-transitory storage media 502, which may include, for example, computer-executable instructions implemented by logic 504, which are operable to enable at least one computer processor to perform operations on device 102 ( Figure 1 )、Server 170( Figure 1 ) and / or compiler 160( Figure 1 ) to implement one or more operations to enable device 102 ( Figure 1 )、Server 170( Figure 1 ) and / or compiler 160( Figure 1 ) performs, triggers and / or implements one or more operations and / or functionalities, and / or performs, triggers and / or implements reference Figures 1 to 4 One or more operations and / or functionalities described, and / or one or more operations described herein. The phrases "non-transitory machine-readable medium" and "computer-readable non-transitory storage medium" may be directed to include all computer-readable media, with the sole exception of transitory propagating signals.
[0466] In some exemplary aspects, the product 500 and / or the machine-readable storage medium 502 may include one or more types of computer-readable storage media capable of storing data, including volatile memory, non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. For example, the machine-readable storage medium 502 may include RAM, DRAM, double data rate DRAM (DDR-DRAM), SDRAM, static RAM (SRAM), ROM, programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory (e.g., NOR or NAND flash memory), content addressable memory (CAM), polymer memory, phase change memory, ferroelectric memory, silicon-oxide-nitride-oxide-silicon (SONOS) memory, disk, hard drive, etc. The computer-readable storage medium may include any suitable medium involved in downloading or transferring a computer program from a remote computer to a requesting computer via a communication link (e.g., a modem, radio, or network connection), the computer program being carried by a data signal embedded in a carrier wave or other propagation medium.
[0467] In some exemplary aspects, logic 504 may include instructions, data, and / or code that, if executed by a machine, may cause the machine to perform methods, processes, and / or operations as described herein. The machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, etc., and may be implemented using any suitable combination of hardware, software, firmware, etc.
[0468] In some exemplary aspects, logic 504 may include or may be implemented as software, a software module, an application, a program, a subroutine, an instruction, an instruction set, a computing code, a word, a value, a symbol, etc. Instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Instructions may be implemented according to a predefined computer language, manner, or syntax for instructing a processor to perform a specific function. Instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language, machine code, etc.
[0469] Examples
[0470] The following examples relate to further aspects.
[0471] Example 1 includes a product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions that are operable to, when executed by at least one processor, enable the at least one processor to cause a compiler to: identify a first mask memory access operation based on source code, wherein the first mask memory access operation is based on a first mask expression including one or more mask leaves; determine a second mask memory access operation by reconfiguring the first mask memory access operation based on an identified mask leaf from the one or more mask leaves, wherein the second mask memory access operation is based on a second mask expression that is logically simplified compared to the first mask expression; and generate target code based on compiling the source code, wherein the target code is based on the second mask memory access operation.
[0472] Example 2 includes the subject matter of Example 1, and optionally wherein the instructions, when executed, cause a compiler to determine the identified mask leaf based on criteria related to the effect of the identified mask leaf on the logical value of the first mask expression,
[0473] Example 3 includes the subject matter of example 2, and optionally wherein the criterion includes the requirement that all possibilities of true logical values of the first mask expression are produced by the same logical value of the identified mask leaf.
[0474] Example 4 includes the subject matter of example 2 or 3, and optionally wherein the criterion includes the requirement that all possibilities of true logical values of the first mask expression are independent of the logical values of the identified mask leaves.
[0475] Example 5 includes the subject matter of any of Examples 1 to 4, and optionally wherein the instructions, when executed, cause a compiler to determine the identified mask leaf based on a determination that the identified mask leaf is based on an induction variable (IV) of a loop that includes a first mask memory access operation.
[0476] Example 6 includes the subject matter of any one of Examples 1 to 5, and optionally wherein the instructions, when executed, cause a compiler to configure an address generation unit (AGU) instruction based on the identified mask leaf, wherein the AGU instructions include memory access instructions for the AGU to perform a second masked memory access operation, wherein the target code is based on the AGU instructions.
[0477] Example 7 includes the subject matter of Example 6, and optionally wherein the instructions, when executed, cause a compiler to configure the AGU instructions to define a memory access range to be applied by the AGU for a second masked memory access operation based on the identified mask leaf.
[0478] Example 8 includes the subject matter of example 7, and optionally wherein the instructions, when executed, cause a compiler to configure the AGU instructions to define at least one of a lower bound or an upper bound of a memory access range based on the identified mask leaf.
[0479] Example 9 includes the subject matter of any one of Examples 1 to 8, and optionally wherein the second mask expression excludes the identified mask leaves.
[0480] Example 10 includes subject matter according to any one of Examples 1 to 9, and optionally wherein the instructions, when executed, cause the compiler to configure the second masked memory access operation to exclude one or more induction variable (IV)-based mask leaves of the first masked memory access operation, the one or more IV-based mask leaves being based on the IV of the loop that includes the first masked memory access operation.
[0481] Example 11 includes subject matter according to any one of Examples 1 to 10, and optionally wherein the instructions, when executed, cause the compiler to configure the second masked memory access operation to exclude any induction variable (IV)-based (IV-based) mask leaves of the first masked memory access operation, any IV-based mask leaves being based on the IV of the loop that includes the first masked memory access operation.
[0482] Example 12 includes subject matter according to any one of Examples 1 to 11, and optionally wherein the instructions, when executed, cause the compiler to configure the second masked memory access operation to maintain one or more non-IV-based mask leaves of the first masked memory access operation, the one or more non-IV-based mask leaves not based on the IV of the loop including the first masked memory access operation.
[0483] Example 13 includes subject matter according to any one of Examples 1 to 12, and optionally wherein the instructions, when executed, cause the compiler to configure the second masked memory access operation to include only non-IV-based mask leaves that are not based on the IV of the loop that includes the first masked memory access operation.
[0484] Example 14 includes the subject matter of any of Examples 1 to 13, and optionally wherein the instructions, when executed, cause a compiler to determine the identified mask leaf based on a truth table corresponding to the first mask expression.
[0485] Example 15 includes the subject matter of any of Examples 1 to 14, and optionally wherein the count of mask leaves in the second mask expression is less than the count of mask leaves in the first mask expression.
[0486] Example 16 includes the subject matter of any of Examples 1 to 15, and optionally wherein the second masked memory access operation comprises a masked load operation to conditionally load a value from memory based on a second mask expression.
[0487] Example 17 includes the subject matter of any of Examples 1 to 16, and optionally wherein the second masked memory access operation comprises a masked store operation to conditionally store a value in memory based on a second mask expression.
[0488] Example 18 includes the subject matter of any of Examples 1-17, and optionally wherein the source code comprises Open Computing Language (OpenCL) code.
[0489] Example 19 includes the subject matter of any one of Examples 1 to 18, and optionally wherein the computer-executable instructions, when executed, cause a compiler to compile source code into target code according to a low-level virtual machine (LLVM-based) compilation scheme.
[0490] Example 20 includes the subject matter of any of Examples 1 to 19, and optionally wherein the target code is configured for execution by a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor.
[0491] Example 21 includes the subject matter of any of Examples 1 to 20, and optionally wherein the object code is configured for execution by a target vector processor.
[0492] Example 22 includes a compiler configured to perform any of the operations described in any of Examples 1 to 21.
[0493] Example 23 includes a computing device configured to perform any of the operations described in any of Examples 1 to 21.
[0494] Example 24 includes a computing system comprising: at least one memory for storing instructions; and at least one processor for retrieving instructions from the memory and executing the instructions to cause the computing system to perform any of the operations described in any of Examples 1 to 21.
[0495] Example 25 includes a computing system comprising: a compiler configured to generate a target code according to any of the operations described in any of Examples 1 to 21; and a processor configured to execute the target code.
[0496] Example 26 includes a device comprising means for performing any of the operations described in any of Examples 1-21.
[0497] Example 27 includes an apparatus comprising: a memory interface; and a processing circuit system configured to: perform any of the operations described in any of Examples 1-21.
[0498] Example 28 includes a method comprising any of the operations described in any of Examples 1 to 21.
[0499] The functions, operations, components, and / or features described herein with reference to one or more aspects may be combined with or utilized in combination with one or more other functions, operations, components, and / or features described herein with reference to one or more other aspects, or vice versa.
[0500] While certain features have been illustrated and described herein, numerous modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the disclosure.
Claims
1. A product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions operable to, when executed by at least one processor, enable the at least one processor to cause a compiler to: identifying a first masked memory access operation based on the source code, wherein the first masked memory access operation is based on a first mask expression including one or more mask leaves; determining a second masked memory access operation by reconfiguring the first masked memory access operation based on an identified mask leaf of the one or more mask leaves, wherein the second masked memory access operation is based on a second mask expression that is logically simplified compared to the first mask expression; as well as An object code is generated based on compiling the source code, wherein the object code is based on the second mask memory access operation.
2. The product of claim 1 , wherein the instructions, when executed, cause the compiler to determine the identified mask leaf based on criteria related to the effect of the identified mask leaf on the logical value of the first mask expression, 3. The product of claim 2, wherein the criterion comprises a requirement that all possibilities of true logical values of the first mask expression are produced by the same logical value of the identified mask leaf.
4. The product of claim 2, wherein the criterion includes a requirement that all possibilities of true logical values of the first mask expression be independent of the logical values of the identified mask leaves.
5. The article of manufacture of claim 1, wherein the instructions, when executed, cause the compiler to determine the identified mask leaf based on a determination that the identified mask leaf is based on an induction variable (IV) of a loop that includes the first masked memory access operation.
6. The product of claim 1 , wherein the instructions, when executed, cause the compiler to configure address generation unit (AGU) instructions based on the identified mask leaf, wherein the AGU instructions include memory access instructions for the AGU to perform the second masked memory access operation, wherein the target code is based on the AGU instructions.
7. The article of claim 6, wherein the instructions, when executed, cause the compiler to configure the AGU instructions to define a memory access range based on the identified mask leaf, the memory access range to be applied by the AGU for the second masked memory access operation.
8. The product of claim 7, wherein the instructions, when executed, cause the compiler to configure the AGU instructions to define at least one of a lower bound or an upper bound of the memory access range based on the identified mask leaf.
9. The product of claim 1, wherein the second mask expression excludes the identified mask leaf.
10. The product of any one of claims 1 to 9, wherein the instructions, when executed, cause the compiler to configure the second masked memory access operation to exclude one or more induction variable (IV)-based mask leaves of the first masked memory access operation, the one or more IV-based mask leaves being based on an IV of a loop including the first masked memory access operation.
11. The product of any one of claims 1 to 9, wherein the instruction, when executed, causes the compiler to configure the second masked memory access operation to exclude any induction variable (IV)-based (IV-based) mask leaves of the first masked memory access operation, wherein any IV-based mask leaves are based on an IV of a loop that includes the first masked memory access operation.
12. The product of any one of claims 1 to 9, wherein the instruction, when executed, causes the compiler to configure the second masked memory access operation to maintain one or more non-IV-based mask leaves of the first masked memory access operation, the one or more non-IV-based mask leaves not being based on an IV of a loop including the first masked memory access operation.
13. The product of any one of claims 1 to 9, wherein the instructions, when executed, cause the compiler to configure the second masked memory access operation to include only non-IV-based mask leaves, the only non-IV-based mask leaves not based on the IV of the loop including the first masked memory access operation.
14. The product of any one of claims 1 to 9, wherein the instructions, when executed, cause the compiler to determine the identified mask leaf based on a truth table corresponding to the first mask expression.
15. The product of any one of claims 1 to 9, wherein the count of mask leaves in the second mask expression is less than the count of mask leaves in the first mask expression.
16. The product of any one of claims 1 to 9, wherein the second masked memory access operation comprises a masked load operation to conditionally load a value from memory based on the second mask expression.
17. The product of any one of claims 1 to 9, wherein the second masked memory access operation comprises a masked store operation to conditionally store a value in memory based on the second mask expression.
18. The product of any one of claims 1 to 9, wherein the source code comprises Open Computing Language (OpenCL) code.
19. The product of any one of claims 1 to 9, wherein the computer executable instructions, when executed, cause the compiler to compile the source code into the target code according to a Low Level Virtual Machine (LLVM)-based compilation scheme.
20. The product of any one of claims 1 to 9, wherein the object code is configured for execution by a Very Long Instruction Word (VLIW) Single Instruction / Multiple Data (SIMD) target processor.
21. A product according to any one of claims 1 to 9, wherein the object code is configured for execution by a target vector processor.
22. A computing system comprising: at least one memory for storing instructions; as well as at least one processor to retrieve the instructions from the memory and to execute the instructions to cause the computing system to: identifying a first masked memory access operation based on the source code, wherein the first masked memory access operation is based on a first mask expression including one or more mask leaves; determining a second masked memory access operation by reconfiguring the first masked memory access operation based on an identified mask leaf of the one or more mask leaves, wherein the second masked memory access operation is based on a second mask expression that is logically simplified compared to the first mask expression; generating a target code based on compiling the source code, wherein the target code is based on the second mask memory access operation; as well as The object code is output.
23. The computing system of claim 22, wherein the instruction, when executed, causes the computing system to configure an address generation unit (AGU) instruction based on the identified mask leaf, wherein the AGU instruction comprises a memory access instruction of the AGU to perform the second masked memory access operation, wherein the target code is based on the AGU instruction.
24. The computing system of claim 22, comprising the target processor.
25. A method comprising: identifying a first masked memory access operation based on the source code, wherein the first masked memory access operation is based on a first mask expression including one or more mask leaves; determining a second masked memory access operation by reconfiguring the first masked memory access operation based on an identified mask leaf of the one or more mask leaves, wherein the second masked memory access operation is based on a second mask expression that is logically simplified compared to the first mask expression; as well as An object code is generated based on compiling the source code, wherein the object code is based on the second mask memory access operation.
26. The method of claim 25, comprising configuring an address generation unit (AGU) instruction based on the identified mask leaf, wherein the AGU instruction comprises a memory access instruction of the AGU to perform the second masked memory access operation, wherein the target code is based on the AGU instruction.