Apparatus, system and method for compiling code for processor

By introducing technical means related to vector microcode processors into the compiler, the problem that existing compilers are difficult to efficiently process functional code in the vector processor environment is solved, efficient code generation and execution are achieved, and performance and efficiency are improved.

CN120035813APending Publication Date: 2025-05-23MOBILEYE VISION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071504.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-10-12
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing compilers have difficulty in effectively supporting efficient processing of functional code, especially in vector processor environments, where there is a need to provide technical solutions to optimize code generation and execution.

Method used

By introducing technical means related to vector microcode processor (VMP) into the compiler, including front-end, mid-end and back-end optimization processing, the OpenCL compiler and LLVM compilation scheme are used to perform instruction scheduling and register allocation to generate efficient target code.

Benefits of technology

It realizes efficient compilation and execution of functional code in a vector processor environment, improves performance and efficiency, and supports the optimization of the complex instruction set computer (CISC) architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035813A_ABST
    Figure CN120035813A_ABST
Patent Text Reader

Abstract

For example, the present disclosure provides a compiler that may be configured to: identify a first data operation and a second data operation, the first data operation and the second data operation being executable in parallel according to a SIMD instruction to be executed by a single ALU of a target processor; a selected compilation scheme is determined from a first compilation scheme and a second compilation scheme based on predefined selection criteria, where the first compilation scheme includes compiling the first data operation and the second data operation into the SIMD instruction, where the second compilation scheme includes compiling the first data operation into a first ALU instruction, and compiling the second data operation into a second ALU instruction to be executed separately from the first ALU instruction; and generating a target code based on compilation of the first data operation and the second data operation according to the selected compilation scheme.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 415,305, filed on October 12, 2022, entitled “APPARATUS, SYSTEM, AND METHOD OF COMPILING CODE FOR A PROCESSOR,” the entire disclosure of which is incorporated herein by reference. Background Art

[0003] The compiler may be configured to compile source code into object code configured for execution by the processor.

[0004] There is a need to provide technical solutions to support efficient processing functionality. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity of presentation. In addition, reference numerals may be repeated in the drawings to indicate corresponding or similar elements. The drawings are listed below.

[0006] Figure 1 is a schematic block diagram illustration of a system according to some exemplary aspects.

[0007] Figure 2 is a schematic illustration of a compiler according to some exemplary aspects.

[0008] Figure 3 is a schematic illustration of a vector processor according to some exemplary aspects.

[0009] Figure 4 is a schematic flow chart illustration of a method of compiling code for a processor according to some exemplary aspects.

[0010] Figure 5 is a schematic flow chart illustration of a method of compiling code for a processor according to some exemplary aspects.

[0011] Figure 6 is a schematic illustration of a product according to some exemplary aspects. DETAILED DESCRIPTION

[0012] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of some aspects. However, it will be appreciated by those of ordinary skill in the art that some aspects may be practiced without these specific details. In other cases, well-known methods, procedures, components, units and / or circuits are not described in detail to avoid obscuring the discussion.

[0013] Some portions of the following detailed description are presented in terms of algorithms and symbolic representations of operations on data bits or binary digital signals within a computer memory. These algorithmic descriptions and representations may be techniques used by those skilled in the data processing arts to convey the substance of their work to others skilled in the art.

[0014] An algorithm is here and generally considered to be a self-consistent sequence of acts or operations leading to a desired result. These include physical manipulations of physical quantities. Typically, but not necessarily, these quantities capture forms of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for common sense reasons, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. It should be understood, however, that all of these terms and similar terms should be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities.

[0015] Discussions herein utilizing terms such as, for example, "process," "compute," "calculate," "determine," "create," "analyze," "verify," and the like may refer to the operations and / or processes of a computer, computing platform, computing system, or other electronic computing device that manipulates data represented as physical (e.g., electronic) quantities within the computer's registers and / or memories and / or transforms that data into other data similarly represented as physical quantities within the computer's registers and / or memories or other information storage media that may store instructions for performing operations and / or processes.

[0016] As used herein, the terms "plurality" and "a plurality" include, for example, "a plurality" or "two or more." For example, "a plurality of items" includes two or more items.

[0017] References to "one aspect," "an aspect," "exemplary aspect," "various aspects," etc. indicate that the aspects so described may include particular features, structures, or characteristics, but not every aspect necessarily includes the particular features, structures, or characteristics. Furthermore, repeated use of the phrase "in one aspect" does not necessarily refer to the same aspect, although it may.

[0018] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe common objects merely indicates that different instances of the same object are being referenced and is not intended to imply that the objects so described must be in a given sequence in time, space, ranking, or in any other manner.

[0019] For example, some aspects may capture the form of entirely hardware aspects, entirely software aspects, or aspects including both hardware and software elements.Some aspects may be implemented in software, including but not limited to firmware, resident software, microcode, etc.

[0020] Furthermore, some aspects may be captured in the form of a computer program product that can be accessed from a computer-usable or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. For example, a computer-usable or computer-readable medium can be or can include any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with an instruction execution system, device, or apparatus.

[0021] In some exemplary aspects, the medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium.

[0022] In some exemplary aspects, a data processing system suitable for storing and / or executing program code may include at least one processor coupled directly or indirectly to a memory element, for example, via a system bus. The memory element may include, for example, local memory employed during actual execution of the program code, a mass storage device, and a cache memory that may provide temporary storage of at least some program code in order to reduce the number of times code must be retrieved from a mass storage device during execution.

[0023] In some exemplary aspects, input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system directly or through an intermediate I / O controller. In some exemplary aspects, a network adapter can be coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices, such as through an intermediate private or public network. In some exemplary aspects, modems, cable modems, and Ethernet cards are exemplary examples of network adapter types. Other suitable components can be used.

[0024] Some aspects may be used in connection with various devices and systems, such as computing devices, computers, mobile computers, non-mobile computers, server computers, and the like.

[0025] As used herein, the term "circuitry" may refer to, be a part of, or include an application specific integrated circuit (ASIC), an integrated circuit, an electronic circuit, a processor (shared, dedicated, or grouped), and / or memory (shared. Dedicated or grouped) that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable hardware components that provide the described functionality. In some aspects, some functions associated with the circuitry may be implemented by one or more software or firmware modules. In some aspects, the circuitry may include logic that is at least partially operable in hardware.

[0026] The term "logic" may refer to, for example, computing logic embedded in the circuit system of a computing device and / or computing logic stored in the memory of a computing device. For example, the logic may be accessed by a processor of a computing device to execute the computing logic to perform computing functions and / or operations. In one example, the logic may be embedded in various types of memory and / or firmware, such as silicon blocks of various chips and / or processors. The logic may be included in various circuit systems and / or implemented as part of various circuit systems, such as processor circuit systems, control circuit systems, and / or the like. In one example, the logic may be embedded in volatile memory and / or non-volatile memory, including random access memory, read-only memory, programmable memory, magnetic memory, flash memory, persistent memory, etc. The logic may be executed by one or more processors using memory (e.g., registers, lags, buffers, and / or the like) coupled to one or more processors, for example, executing the logic as needed.

[0027] Reference now Figure 1 , which schematically illustrates a block diagram of a system 100 according to some exemplary aspects.

[0028] like Figure 1 As shown, in some demonstrative aspects, system 100 may include a computing device 102 .

[0029] In some demonstrative aspects, device 102 may be implemented using suitable hardware components and / or software components, such as processors, controllers, memory units, storage units, input units, output units, communication units, operating systems, applications, and the like.

[0030] In some demonstrative aspects, device 102 may comprise, for example, a computer, a mobile computing device, a non-mobile computing device, a laptop computer, a notebook computer, a tablet computer, a handheld computer, a personal computer (PC), or the like.

[0031] In some exemplary aspects, device 102 may include, for example, one or more of the following: processor 191, input unit 192, output unit 193, memory unit 194, and / or storage unit 195. Device 102 may optionally include other suitable hardware components and / or software components. In some exemplary aspects, some or all components of one or more of devices in device 102 may be enclosed in a common housing or packaging and may be interconnected or operably associated using one or more wired or wireless links. In other aspects, components of one or more of devices in device 102 may be distributed in multiple or separate devices.

[0032] In some exemplary aspects, processor 191 may include, for example, a central processing unit (CPU), a digital signal processor (DSP), one or more processor cores, a single core processor, a dual core processor, a multi-core processor, a microprocessor, a host processor, a controller, multiple processors or controllers, a chip, a microchip, one or more circuits, a circuit system, a logic unit, an integrated circuit (IC), an application specific IC (ASIC), or any other suitable general-purpose or specific processor or controller. Processor 191 may execute, for example, instructions of an operating system (OS) of device 102 and / or instructions of one or more suitable applications.

[0033] In some exemplary aspects, input unit 192 may include, for example, a keyboard, a keypad, a mouse, a touch screen, a touch pad, a trackball, a stylus, a microphone, or other suitable pointing device or input device. Output unit 193 may include, for example, a monitor, a screen, a touch screen, a flat panel display, a light emitting diode (LED) display unit, a liquid crystal display (LCD) display unit, a plasma display unit, one or more audio speakers or headphones, or other suitable output devices.

[0034] In some exemplary aspects, memory unit 194 includes, for example, random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous DRAM (SD-RAM), flash memory, volatile memory, non-volatile memory, cache memory, buffer, short-term memory unit, long-term memory unit, or other suitable memory unit. Storage unit 195 may include, for example, a hard disk drive, a solid-state drive (SSD), or other suitable removable or non-removable storage unit. Memory unit 194 and / or storage unit 195 may, for example, store data processed by device 102.

[0035] In some demonstrative aspects, device 102 may be configured to communicate with one or more other devices via at least one network 103 (eg, a wireless and / or wired network).

[0036] In some exemplary aspects, network 103 may include a wired network, a local area network (LAN), a wireless network, a wireless LAN (WLAN) network, a radio network, a cellular network, a WiFi network, an IR network, a Bluetooth (BT) network, and the like.

[0037] In some demonstrative aspects, device 102 may be configured to perform and / or execute one or more operations, modules, processes, procedures, and / or the like, e.g., as described herein.

[0038] In some demonstrative aspects, device 102 may include compiler 160, which may be configured to generate object code 115 based on source code 112, for example, as described below.

[0039] In some exemplary aspects, compiler 160 may be configured to translate source code 112 into target code 115, eg, as described below.

[0040] In some demonstrative aspects, compiler 160 may include or may be implemented as software, a software module, an application, a program, a subroutine, instructions, an instruction set, computing code, words, values, symbols, and / or the like.

[0041] In some exemplary aspects, source code 112 may include computer code written in a source language.

[0042] In some exemplary aspects, the source language may include a programming language. For example, the source language may include a high-level programming language, such as, for example, C language, C++ language and / or the like.

[0043] In some exemplary aspects, target code 115 may include computer code written in a target language.

[0044] In some exemplary aspects, the target language can include a low-level language such as, for example, assembly language, object code, machine code, or the like.

[0045] In some exemplary aspects, object code 115 may include one or more purpose files, which may, for example, create and / or form an executable program.

[0046] In some exemplary aspects, the executable program may be configured to be executed on a target computer. For example, the target computer may include specific computer hardware, a specific machine and / or a specific operating system.

[0047] In some exemplary aspects, the executable program may be configured to be executed on processor 180, eg, as described below.

[0048] In some demonstrative aspects, processor 180 may include a vector processor 180, eg, as described below. In other aspects, processor 180 may include any other type of processor.

[0049] Some exemplary aspects are described herein with respect to a compiler (e.g., compiler 160) that is configured to compile source code 112 into target code 115 that is configured to be executed by a vector processor 180, e.g., as described below. In other aspects, a compiler (e.g., compiler 160) is configured to compile source code 112 into target code 115 that is configured to be executed by any other type of processor 180.

[0050] In some demonstrative aspects, processor 180 may be implemented as part of device 102 .

[0051] In other aspects, processor 180 may be implemented as part of any other device separate from device 102 , for example.

[0052] In some demonstrative aspects, vector processor 180 (also referred to as an "array processor") may include a processor that may be configured to process an entire vector in one instruction, eg, as described below.

[0053] In other aspects, the executable program may be configured to be executed on any other additional or alternative type of processor.

[0054] In some exemplary aspects, vector processor 180 may be designed to support high-performance image and / or vector processing. For example, vector processor 180 may be configured to process 1 / 2 / 3 / 4D arrays and / or floating point arrays of fixed point data very quickly and / or efficiently.

[0055] In some exemplary aspects, vector processor 180 can be configured to process arbitrary data, such as structures with pointers to structures. For example, vector processor 180 can include a scalar processor to calculate non-vector data, such as assuming that the non-vector data is minimal.

[0056] In some demonstrative aspects, compiler 160 may be implemented as a native application to be executed by device 102. For example, memory unit 194 and / or storage unit 195 may store instructions generated in compiler 160, and / or processor 191 may be configured to execute instructions generated in compiler 160 and / or perform one or more calculations and / or processes of compiler 160, e.g., as described below.

[0057] In other aspects, compiler 160 may comprise a remote application to be executed by any suitable computing system (eg, server 170 ).

[0058] In some exemplary aspects, server 170 may include at least a remote server, a network-based server, a cloud server, and / or any other server.

[0059] In some exemplary aspects, server 170 may include a suitable memory and / or storage unit 174 having stored thereon instructions generated in compiler 160 and a suitable processor 171 to execute the instructions, e.g., as described below.

[0060] In some exemplary aspects, compiler 160 may include a combination of remote applications and local applications.

[0061] In one example, compiler 160 may be downloaded and / or received by a user of device 102 from another computing system (e.g., server 170) such that compiler 160 may be executed locally by the user of device 102. For example, instructions may be received and stored temporarily in a memory or any suitable short-term storage or buffer of device 102, e.g., prior to execution by processor 191 of device 102.

[0062] In another example, compiler 160 may include a client module to be executed locally by device 102 and a server module to be executed by server 170. For example, the client module may include and / or may be implemented as a local application, a web application, a website, a web client, e.g., a hypertext markup language (HTML) web application, etc.

[0063] For example, one or more first operations of compiler 160 may be performed locally, such as by device 102 , and / or one or more second operations of compiler 160 may be performed remotely, such as by server 170 .

[0064] In other aspects, compiler 160 may include or be implemented by any other suitable computing arrangement and / or scheme.

[0065] In some demonstrative aspects, system 100 may include an interface 110 (eg, a user interface) to interface between a user of device 102 and one or more elements of system 100 (eg, compiler 160).

[0066] In some demonstrative aspects, interface 110 may be implemented using any suitable hardware components and / or software components, such as a processor, a controller, a memory unit, a storage unit, an input unit, an output unit, a communication unit, an operating system, and / or an application.

[0067] In some aspects, interface 110 may be implemented as part of any suitable module, system, device, or component of system 100 .

[0068] In other aspects, interface 110 may be implemented as a separate element of system 100 .

[0069] In some demonstrative aspects, interface 110 may be implemented as part of device 102. For example, interface 110 may be associated with device 102 and / or included as part of the device.

[0070] In one example, interface 110 can be implemented as part of any suitable application, such as middleware and / or device 102. For example, interface 110 can be implemented as part of compiler 160 and / or part of the OS of device 102.

[0071] In some demonstrative aspects, interface 110 may be implemented as part of server 170. For example, interface 110 may be associated with server 170 and / or included as part of the server.

[0072] In one example, interface 110 may include or be part of: a web-based application, a website, a web page, a plug-in, an ActiveX control, a rich content component (eg, a Flash or Shockwave component), or the like.

[0073] In some exemplary aspects, interface 110 may be associated therewith and / or may include, for example, a gateway (GW) 113 and / or an application programming interface (API) 114, for example, to transmit information and / or communicate between elements of system 100 and / or to one or more other parties (e.g., internal or external parties), users, applications and / or systems.

[0074] In some aspects, interface 110 may include any suitable graphical user interface (GUI) 116 and / or any other suitable interface.

[0075] In some demonstrative aspects, interface 110 may be configured to receive source code 112 from, for example, a user of device 102 via GUI 116 and / or API 114 .

[0076] In some exemplary aspects, interface 110 may be configured to transfer source code 112 to, for example, compiler 160 , for example, to generate object code 115 , for example, as described below.

[0077] refer to Figure 2 , which schematically illustrates a compiler 200 according to some exemplary aspects. For example, the compiler 160 ( Figure 1 ) may implement one or more elements of compiler 200 and / or may perform one or more operations and / or functionalities of compiler 200.

[0078] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may be configured to generate target code 233, for example, by compiling source code 212 in a source language.

[0079] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may include a front end 210 configured to receive and analyze source code 212 in a source language.

[0080] In some exemplary aspects, front end 210 may be configured to generate intermediate code 213 , for example, based on source code 212 .

[0081] In some exemplary aspects, intermediate code 213 may comprise a lower-level representation of source code 212 .

[0082] In some exemplary aspects, front end 210 can be configured to perform, for example, lexical analysis, syntactic analysis, semantic analysis, and / or any other additional or alternative types of analysis of source code 212 .

[0083] In some exemplary aspects, front end 210 can be configured to identify errors and / or problems using the results of the analysis of source code 212. For example, front end 210 can be configured to generate error information, e.g., including error and / or warning messages, which can identify a location in source code 212, e.g., where an error or problem is detected.

[0084] In some exemplary aspects, such as Figure 2 As shown, compiler 200 may include a middle end 220 configured to receive and process intermediate code 213 and generate adjusted (eg, optimized) intermediate code 223 .

[0085] In some demonstrative aspects, middle end 220 may be configured to perform one or more adjustments (eg, optimizations) to intermediate code 213 , eg, to generate adjusted intermediate code 223 .

[0086] In some exemplary aspects, middle end 220 can be configured to perform one or more optimizations on intermediate code 213 , eg, independent of the type of target computer used to execute target code 233 .

[0087] In some exemplary aspects, middle end 220 can be implemented to support the use of optimized intermediate code 223, eg, for different machine types.

[0088] In some exemplary aspects, middle end 220 may be configured to optimize the intermediate representation of intermediate code 223 , for example, to improve the performance and / or quality of the generated target code 233 .

[0089] In some exemplary aspects, one or more optimizations of intermediate code 213 may include, for example, inline expansion, dead code elimination, constant propagation, loop transformation, parallelization, and / or the like.

[0090] In some exemplary aspects, such as Figure 2 As shown, the compiler 200 may include a back end 230 configured to receive and process the adjusted intermediate code 213 , and generate a target code 233 based on the adjusted intermediate code 213 .

[0091] In some exemplary aspects, backend 230 may be configured to perform one or more operations and / or processes that may be specific to a target computer used to execute target code 233. For example, backend 230 may be configured to process optimized intermediate code 213 by applying analysis, transformation, and / or optimization operations to adjusted intermediate code 213, which operations may be configured, for example, based on a target computer used to execute target code 233.

[0092] In some exemplary aspects, the one or more analysis, transformation, and / or optimization operations applied to the adjusted intermediate code 213 may include, for example, resource and storage decisions, such as register allocation, instruction scheduling, and / or the like.

[0093] In some exemplary aspects, the object code 233 may include target-dependent assembly code that may be specific to a target computer used to execute the object code 233 and / or a target operating system of the target computer.

[0094] In some exemplary aspects, the object code 233 may include code for a processor (e.g., vector processor 180 ( Figure 1 ))'s target-dependent assembly code.

[0095] In some exemplary aspects, compiler 200 may include a vector microcode processor (VMP) open computing language (OpenCL) compiler, for example, as described below. In other aspects, compiler 200 may include any other type of vector processor compiler, or may be implemented as part of any other type of vector processor compiler.

[0096] In some exemplary aspects, the VMP OpenCL compiler may include a low-level virtual machine (LLVM)-based compiler that may be configured according to an LLVM-based compilation scheme, for example, to reduce OpenCL C code to VMP accelerator assembly code, for example, suitable for use by vector processor 180 ( Figure 1 )implement.

[0097] In some exemplary aspects, compiler 200 may include one or more techniques that may be required to compile code into a format suitable for a VMP architecture, for example, in addition to an open source LLVM compiler pass.

[0098] In some exemplary aspects, FE 210 may be configured to parse OpenCL C code and translate it, for example, via an abstract syntax tree (AST), into, for example, an LLVM intermediate representation (IR).

[0099] In some exemplary aspects, compiler 200 may include a dedicated API, for example, to detect the correct pattern for compiler pattern matching, e.g., a pattern suitable for VMP. For example, VMP may be configured as a complex instruction set computer (CISC) machine that implements a very complex instruction set architecture (ISA) that may be difficult to target from standard C code. Accordingly, compiler pattern matching may not be able to easily detect the correct pattern, and for such cases, the compiler may require a dedicated API.

[0100] In some exemplary aspects, FE 210 may implement one or more vendor extension builtins that may target a VMP-specific ISA, for example, in addition to standard OpenCL builtins that may be optimized for VMP machines.

[0101] In some exemplary aspects, FE 210 may be configured to implement OpenCL constructs and / or work-item functionality.

[0102] In some exemplary aspects, ME 220 may be configured to process LLVM IR code, which may be generic and target-independent, e.g., although it may include one or more hooks for a specific target architecture.

[0103] In some demonstrative aspects, ME 220 may perform one or more custom passes, for example, to support a VMP architecture, for example, as described below.

[0104] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of control flow graph (CFG) linearization analysis, e.g., as described below.

[0105] In some exemplary aspects, the CFG linearization analysis can be configured to linearize the code, for example, by converting if statements to select modes, for example, where the VMP vector code does not support standard control flow.

[0106] In one example, ME 220 may receive a given code, for example, as follows:

[0107]

[0108] According to this example, ME 220 may be configured to apply CFG linearization analysis to a given code, for example, as follows:

[0109] tmpA=A+5;

[0110] tmpB = B * 2;

[0111] mask=x>0;

[0112] A=Select mask,tmpA,A

[0113] B=Select not mask,tmpB,B

[0114] Example (1)

[0115] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of automatic vectorization analysis, e.g., as described below.

[0116] In some exemplary aspects, the auto-vectorization analysis may be configured to vectorize (eg, auto-vectorize) a given code, for example, to exploit the vector capabilities of the VMP.

[0117] In some exemplary aspects, ME 220 may be configured to perform automatic vectorization analysis, e.g., to vectorize code into scalar form. For example, some or all operations of automatic vectorization analysis may not be performed, such as when the code is already provided in vectorized form.

[0118] In some exemplary aspects, for example, in some use cases and / or scenarios, a compiler may not always be able to auto-vectorize code, for example, due to data dependencies between loop iterations.

[0119] In one example, ME 220 may receive a given code, for example, as follows:

[0120]

[0121] According to this example, ME 220 may be configured to perform CFG auto-vectorization analysis by applying a first transformation, for example, as follows:

[0122]

[0123] For example, ME 220 may be configured to perform CFG auto-vectorization analysis by applying a second transformation, e.g., after a first transformation, e.g., as follows:

[0124]

[0125] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of scratch pad memory cycle access analysis (SPMLAA), e.g., as described below.

[0126] In some exemplary aspects, the SPMLAA may define a processing block (PB), for example, that should later be outlined and compiled for VMP.

[0127] In some exemplary aspects, a processing block may include an accelerated loop that may be executed by a vector unit of a VMP.

[0128] In some exemplary aspects, a PB (eg, each PB) may include memory references. For example, some or all memory accesses may refer to a local memory bank.

[0129] In some exemplary aspects, the VMP may enable the AGU (e.g., as described below with reference to Figure 3 AGU 320) and scatter-gather unit (SG) described above are used to access the memory bank.

[0130] In some exemplary aspects, the AGU can be pre-configured, for example, before a loop is executed. For example, a loop trip count can be calculated, for example, before running a processing block.

[0131] In some exemplary aspects, image references can be created at this stage, eg, some or all image references, and strides and offsets can then be calculated, eg, per-dimension strides and offsets for each reference.

[0132] In some exemplary aspects, ME 220 may be configured to perform one or more operations of AGU planner analysis, e.g., as described below.

[0133] In some exemplary aspects, the AGU planner analysis can include an iterator specification that can cover image references from an entire processing block, eg, all image references.

[0134] In some exemplary aspects, an iterator may cover a single reference or a group of references.

[0135] In some exemplary aspects, one or more memory references may be combined via a shuffle instruction and / or reuse the same access, and / or preserve values ​​read from a previous iteration.

[0136] In some exemplary aspects, other memory references, such as those without a linear access pattern, may be processed using a scatter-gather (SG) unit, which may have a performance penalty, such as because it may need to maintain indexes and / or masks.

[0137] In some exemplary aspects, a plan may be configured as an arrangement of iterators in a processing block. For example, a processing block may, for example, theoretically have multiple plans.

[0138] In some exemplary aspects, the AGU planner analysis can be configured to construct all possible plans for all PBs and select a combination, eg, the best combination, from among all valid combinations.

[0139] In some exemplary aspects, the total number of iterators in a valid combination may be limited, eg, not to exceed the number of available AGUs on the VMP.

[0140] In some exemplary aspects, one or more parameters may be defined for an iterator (e.g., for each iterator), e.g., including stride, width, and / or cardinality, e.g., as part of an AGU planner analysis. For example, a minimum-maximum range for an iterator may be defined dimensionally, e.g., in each dimension, e.g., as part of an AGU planner analysis.

[0141] In some exemplary aspects, the AGU planner analysis can be configured to track and evaluate memory references to the image, eg, each memory reference, eg, to understand its access pattern.

[0142] In one example, according to Example 2a / 2b, image "a" as a base address can be accessed with 64 iterations using a step size of 32 bytes.

[0143] In some exemplary aspects, LLVM can include scalar evaluation analysis (SCEV) that can compute access patterns, for example, to understand each image reference.

[0144] In some exemplary aspects, ME 220 may exploit the masking capabilities of the AGU, eg, to avoid maintaining induction variables, which may have a performance penalty.

[0145] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of rewrite analysis, e.g., as described below.

[0146] In some exemplary aspects, the rewrite analysis may be configured to transform the code of a processing block, for example, when setting up iterators and / or modifying memory access instructions.

[0147] In some exemplary aspects, the setup of iterators (e.g., all iterators) can be implemented in the IR in a target-specific intrinsic function. For example, the setup of iterators can reside in the pre-header of the outermost loop.

[0148] In some exemplary aspects, the rewrite analysis may include a loop completion analysis, eg, as described below.

[0149] In some exemplary aspects, the code may be compiled with the goal that substantially all computations should be performed within the innermost loop.

[0150] For example, loop finishing analysis may promote instructions, for example, to move operations performed after the last iteration of the loop into the loop.

[0151] For example, loop finishing analysis may sink instructions, eg, to move operations performed before the first iteration of the loop into the loop.

[0152] For example, loop finishing analysis may hoist instructions and / or sink instructions, eg, such that substantially all instructions from an outer loop are moved to an innermost loop.

[0153] For example, loop completion analysis may be configured to provide technical solutions to support VMP iterators, for example, to work only on perfectly nested loops.

[0154] For example, loop completion analysis may lead to a situation where there are no instructions between "for" statements that make up a loop, e.g., to support VMP iterators, which cannot emulate such a situation.

[0155] In some exemplary aspects, loop refinement analysis may be configured to collapse nested loops into a single collapsed loop.

[0156] In one example, ME 220 may receive a given code, for example, as follows:

[0157]

[0158]

[0159] According to this example, ME 220 may be configured to perform loop completion analysis to collapse nested loops in the code into a single folded loop, for example, as follows:

[0160]

[0161] In some demonstrative aspects, ME 220 may be configured to perform one or more operations of vector loop delimitation analysis, eg, as described below.

[0162] In some exemplary aspects, the vector loop demarcation analysis can be configured to partition the code between the scalar subsystem and the vector subsystem, for example, as described below with reference to Figure 3 The vector processing block 310 ( Figure 3 ) and scalar processor 330( Figure 3 )between.

[0163] In some exemplary aspects, a VMP accelerator may include scalar and / or vector subsystems, e.g., as described below. For example, each of the subsystems may have a different computational unit / processor. Accordingly, scalar code may be compiled on a scalar compiler (e.g., an SSC compiler), and / or accelerated vector code may run on a VMP vector processor.

[0164] In some exemplary aspects, vector loop demarcation analysis can be configured to create separate functions for accelerating loop bodies of vector code. For example, these functions can be marked for VMP and / or can proceed to the VMP backend, for example, while the rest of the code can be compiled by the SSC compiler.

[0165] In some exemplary aspects, one or more portions of a vector loop (e.g., configuration of a vector unit and / or initialization of vector registers) may be performed by a scalar unit. However, these portions may be performed at a later stage, e.g., by backfilling the scalar code, e.g., because the scalar code may still be in LLVM IR before being processed by the SSC compiler.

[0166] In some exemplary aspects, BE 230 may be configured to translate LLVM IR into machine instructions. For example, BE 230 may not be target agnostic and may be familiar with target specific architectures and optimizations, for example, compared to ME 220 which may be agnostic to target specific architectures.

[0167] In some exemplary aspects, BE 230 may be configured to perform one or more analyses that may be specific to the target machine (eg, a VMP machine) to which the code is being downgraded, for example, even though BE 230 may use a general-purpose LLVM.

[0168] In some exemplary aspects, BE 230 may be configured to perform one or more operations of instruction degradation analysis, eg, as described below.

[0169] In some exemplary aspects, instruction degradation analysis may be configured to translate LLVM IR into target-specific instruction machine IR (MIR), for example, by translating LLVM IR into a directed acyclic graph (DAG).

[0170] In some exemplary aspects, the DAG may undergo a legalization process for instructions, such as based on data types and / or VMP instructions, which may be supported by the VMP HW.

[0171] In some exemplary aspects, instruction demotion analysis may be configured to, for example, perform a pattern matching process after a legalization process of instructions, for example, to demotion nodes (eg, each node) in a DAG to, for example, VMP-specific machine instructions.

[0172] In some exemplary aspects, instruction degradation analysis may be configured to generate a MIR, for example, after a pattern matching process.

[0173] In some exemplary aspects, instruction demotion analysis may be configured to degrade instructions according to a machine application binary interface (ABI) and / or calling convention.

[0174] In some exemplary aspects, BE 230 can be configured to perform one or more operations of a cell balance analysis, eg, as described below.

[0175] In some exemplary aspects, the unit balancing analysis may be configured to balance instructions among VMP computing units, for example, as described below with reference to Figure 3 The data processing unit 316 ( Figure 3 )between.

[0176] In some exemplary aspects, the cell balance analysis may be aware of some or all available arithmetic transformations, and / or may perform transformations according to an optimal algorithm.

[0177] In some exemplary aspects, BE 230 may be configured to perform one or more operations of a modulo scheduler (pipeliner) analysis, eg, as described below.

[0178] In some exemplary aspects, the pipeliner may be configured to schedule instructions according to one or more constraints (e.g., data dependencies, resource bottlenecks, and / or any other constraints), for example using a swing modulo scheduling (SMS) heuristic and / or any other additional and / or alternative heuristics.

[0179] In some exemplary aspects, the pipeliner can be configured to schedule a set of very long instruction word (VLIW) instructions (eg, of initiation intervals (II)) over which a program will iterate, such as during a steady state.

[0180] In some exemplary aspects, a performance metric may be measured, which may be based on the number of cycles a typical loop may execute, for example, as follows:

[0181] (input data size in bytes)*II / (bytes consumed / produced per iteration)

[0182] In some exemplary aspects, the pipeliner can attempt to minimize II as much as possible, for example, to improve performance.

[0183] In some exemplary aspects, the pipeliner can be configured to calculate a minimum II and schedule accordingly. For example, if the pipeliner fails to schedule, the pipeliner can attempt to increase the II and retry scheduling, for example, until a predefined II threshold is violated.

[0184] In some demonstrative aspects, BE 230 may be configured to perform one or more operations of register allocation analysis, eg, as described below.

[0185] In some exemplary aspects, register allocation analysis can be configured to attempt to assign registers in an efficient (eg, optimal) manner.

[0186] In some exemplary aspects, register allocation analysis may assign values ​​to bypass vector registers, general purpose vector registers, and / or scalar registers.

[0187] In some exemplary aspects, the values ​​may include private variables, constants, and / or values ​​that rotate across iterations.

[0188] In some exemplary aspects, register allocation analysis may implement an optimal heuristic that fits one or more VMP register file (regfile) constraints. For example, in some use cases, register allocation analysis may not use standard LLVM register allocation.

[0189] In some exemplary aspects, in some cases, register allocation analysis may fail, which may mean that the loop cannot be compiled. Accordingly, register allocation analysis may implement a retry mechanism that may return to the modulo scheduler and may attempt to reschedule the loop, e.g., with an increased launch interval. For example, in many cases, increasing the launch interval may reduce register starvation and / or may support compilation of vector loops.

[0190] In some demonstrative aspects, BE 230 may be configured to perform one or more operations of SSC configuration analysis, eg, as described below.

[0191] In some exemplary aspects, the SSC configuration analysis may be configured to set a configuration for executing a kernel, such as an AGU configuration.

[0192] In some exemplary aspects, SSC configuration analysis may be performed at a later stage, such as due to configurations being calculated after legalization, register allocation analysis, and / or modulo scheduling analysis.

[0193] In some exemplary aspects, the SSC configuration analysis can include a zero overhead loop (ZOL) mechanism in a vector loop. For example, the ZOL mechanism can configure loop trip counts based on access patterns of memory references in the loop, e.g., to avoid running instructions that check loop exit conditions for each iteration.

[0194] In some exemplary aspects, a VMP compilation flow may include one or more (e.g., a small number) of steps that may be called during the compilation flow in a test library (testlib) (e.g., a wrapper script for compilation, execution, and / or program testing). For example, these steps may be performed outside of the LLVM compiler.

[0195] In some exemplary aspects, a PCB Hardware Description Language (PHDL) simulator may be implemented to perform one or more roles of an assembler, an encoder, and / or a linker.

[0196] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support robustness, which may enable compilation of a wide range of loop selections in the presence of HW limitations. For example, compiler 200 may be configured to support technical solutions that may not generate verification errors.

[0197] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support programmability, which may provide the user with the ability to express code in multiple ways that may be correctly compiled to the VMP architecture.

[0198] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support an improved user experience, which may allow the user to be able to debug and / or profile the code. For example, the improved user experience may provide informative error messages, reporting tools, and / or profiling tools.

[0199] In some exemplary aspects, compiler 200 may be configured to provide technical solutions to support improved performance, e.g., to optimize VMP assembly code and / or iterator access, which may result in faster execution. For example, improved performance may be achieved through high utilization of computing units and use of their complex CISC.

[0200] Reference Figure 3 , which schematically illustrates a vector processor 300 according to some exemplary aspects. For example, vector processor 180 ( Figure 1 ) may implement one or more elements of vector processor 300, and / or may perform one or more operations and / or functionality of vector processor 300.

[0201] In some exemplary aspects, vector processor 300 may include a Vector Microcode Processor (VMP).

[0202] In some exemplary aspects, vector processor 300 may include a wide vector machine, e.g., supporting a Very Long Instruction Word (VLIW) architecture and / or a Single Instruction / Multiple Data (SIMD) architecture.

[0203] In some exemplary aspects, vector processor 300 may be configured to provide technical solutions to support high performance for short integer types, which may be common in, for example, computer vision and / or deep learning algorithms.

[0204] In other aspects, the vector processor 300 may include any other type of vector processor, and / or may be configured to support any other additional or alternative functionality.

[0205] In some exemplary aspects, such as Figure 3 As shown, the vector processor 300 may include a vector processing block (vector processor) 310, a scalar processor 330, and a direct memory access (DMA) 340, for example, as described below.

[0206] In some exemplary aspects, such as Figure 3 As shown, the vector processing block 310 may be configured to process (eg, efficiently process) image data and / or vector data. For example, the vector processing block 310 may be configured to use a vector computing unit, for example, to accelerate computation.

[0207] In some exemplary aspects, scalar processor 330 may be configured to perform scalar calculations. For example, scalar processor 330 may be used as "glue logic" for a program that includes vector calculations. For example, some (e.g., even most) of the calculations of a program may be performed by vector processing block 310. However, several tasks (e.g., some basic tasks) (e.g., scalar calculations) may be performed by scalar processor 330.

[0208] In some demonstrative aspects, DMA 340 may be configured to interface with one or more memory elements in a chip including vector processor 300 .

[0209] In some demonstrative aspects, DMA 340 may be configured to read input from main memory, and / or write output to main memory.

[0210] In some exemplary aspects, scalar processor 330 and vector processing block 310 may use respective local memories to process data.

[0211] In some exemplary aspects, such as Figure 3 As shown, the vector processor 300 may include an extractor and decoder 350 , which may be configured to control the scalar processor 330 and / or the vector processing block 310 .

[0212] In some exemplary aspects, operations of scalar processor 330 and / or vector processing block 310 may be triggered by instructions stored in program memory 352 .

[0213] In some demonstrative aspects, DMA 340 may be configured to transfer data in parallel with the execution of program instructions in memory 352, for example.

[0214] In some exemplary aspects, DMA 340 may be controlled by software, such as via configuration registers, rather than instructions, for example, and accordingly may be considered a second “thread” of execution in vector processor 300 .

[0215] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include one or more data processing units, e.g., a group of data processing units, e.g., as described below.

[0216] In some exemplary aspects, a data processing unit may include hardware configured to perform calculations, such as an arithmetic logic unit (ALU).

[0217] In one example, the data processing unit may be configured to add numbers and / or store numbers in memory.

[0218] In some exemplary aspects, the data processing unit may be controlled by commands encoded in, for example, program memory 352 and / or configuration registers. For example, the configuration registers may be memory mapped and writeable by memory storage commands of scalar processor 330.

[0219] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may include a state configuration including a set of registers and memory, eg, as described below.

[0220] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a set of vector memories 312 , which may be configured, for example, to store data to be processed by the vector processor block 310 .

[0221] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a set of vector registers 314 that may be configured for use, for example, in data processing performed by the vector processor block 310 .

[0222] In some demonstrative aspects, scalar processor 330, vector processing block 310, and / or DMA 340 may be associated with a set of memory maps.

[0223] In some exemplary aspects, a memory map may include a set of addresses accessible by a data processing unit that may load data from / to registers and memory and / or store data.

[0224] In some exemplary aspects, such as Figure 3As shown, the vector processing block 310 may include a plurality of address generation units (AGUs) 320 , which may include addresses accessible to them, for example, in one or more memories in the memory 312 .

[0225] In some exemplary aspects, such as Figure 3 As shown, the vector processor block 310 may include a plurality of data processing units 316, for example, as described below.

[0226] In some exemplary aspects, the data processing unit 316 may be configured to process commands, for example, including a number of digits at a time. In one example, the command may include 8 digits. In another example, the command may include 4 digits, 16 digits, or any other count of digits.

[0227] In some exemplary aspects, two or more data processing units 316 can be used simultaneously. In one example, data processing unit 316 can process and execute multiple different commands, for example, 3 different commands, for example, including 8 numbers, in a single cycle.

[0228] In some exemplary aspects, the data processing units 316 may be asymmetric. For example, the first and second data processing units 316 may support different commands. For example, addition may be performed by the first data processing unit 316, and / or multiplication may be performed by the second data processing unit 316. For example, both operations may be performed by one or more additional data processing units 316.

[0229] In some demonstrative aspects, data processing unit 316 may be configured to support arithmetic operations for many combinations of input and output data types.

[0230] In some exemplary aspects, data processing unit 316 may be configured to support one or more operations, which may be less common. For example, processing unit 316 may support operations to work with a lookup table (LUT) of vector processor 300 and / or any other operations.

[0231] In some exemplary aspects, data processing unit 316 may be configured to support efficient computation of nonlinear functions, histograms, and / or random data access, which may, for example, facilitate implementation of algorithms like image scaling, Hough transform, and / or any other algorithm.

[0232] In some exemplary aspects, vector memory 312 may include a bank of memory having a size of 16K, for example, or any other size, that may be accessed in the same cycle.

[0233] In one example, the maximum memory access size may be 64 bits. According to this example, the peak throughput may be 256 bits, for example, 64×4=256. For example, a high memory bandwidth may be achieved to utilize the computational power of the data processing unit 316 .

[0234] In one example, two data processing units 316 may support 16 8-bit multiply and accumulate operations (MACs) per cycle. According to this example, two data processing units 316 may not be useful, for example, if the input numbers are not extracted at that speed, and / or there is no input of exactly 256 bits, for example, 16x8x2=256.

[0235] In some exemplary aspects, AGU 320 may be configured to perform memory access operations, such as loading and storing data from / to vector memory 314 .

[0236] In some exemplary aspects, AGU 320 may be configured to calculate addresses of input and output data items, for example, to handle I / O in situations where high bandwidth is not sufficient to utilize data processing unit 316 .

[0237] In some exemplary aspects, AGU 320 may be configured to calculate addresses of input and / or output data items, for example, based on configuration registers written by scalar processor 330, prior to entering a vector command block (eg, a loop).

[0238] For example, the AGU 320 may be configured to write an image base pointer, width, height, and / or stride to configuration registers, for example, to iterate over an image.

[0239] In some exemplary aspects, the AGU 320 may be configured to handle addressing (e.g., all addressing), for example, to provide a technical solution in which the data processing unit 316 may not have the burden of incrementing a pointer or counter in a loop and / or the burden of checking a line end condition, for example, to zero a counter in a loop.

[0240] In some exemplary aspects, such as Figure 3 As shown, the AGU 320 may include four AGUs, and accordingly, four memories 312 may be accessed in the same cycle. In other aspects, any other count of AGUs 32 may be implemented.

[0241] In some exemplary aspects, AGUs 320 may not be "bound" to memory banks 312. For example, an AGU 320 (e.g., each AGU 320) may access a memory bank 312 (e.g., each memory bank 312), e.g., as long as two or more AGUs 320 do not attempt to access the same memory bank 312 in the same cycle.

[0242] In some demonstrative aspects, vector registers 314 may be configured to support communications between data processing unit 316 and AGU 320 .

[0243] In one example, the total number of vector registers 314 may be 28, which may be divided into several subsets, for example, based on their functions. For example, a first subset of vector registers 314 may be used for input / output of, for example, all data processing units 316 and / or AGU 320; and / or a second subset of vector registers 314 may not be used for output of some operations (e.g., most operations) and may be used for one or more other operations, for example, to store loop-invariant inputs.

[0244] In some exemplary aspects, a data processing unit 316 (e.g., each data processing unit 316) may have one or more registers to host the output of the last performed operation, e.g., which may be fed as input to other data processing units 316. For example, these registers may "bypass" vector registers 314 and may operate faster than writing these outputs to the first set of vector registers 314.

[0245] In some exemplary aspects, the extractor and decoder 350 may be configured to support low-overhead vector loops, e.g., very low-overhead vector loops (also referred to as "zero-overhead vector loops"), e.g., where a termination (exit) condition of the vector loop may not need to be checked during execution of the vector loop.

[0246] For example, the AGU 320 may signal a termination (exit) condition, such as when the AGU 320 completes iterations over a configured memory region.

[0247] For example, the fetcher and decoder 350 may exit the loop when, for example, the AGU 320 signals a termination condition.

[0248] For example, the scalar processor 330 may be utilized to configure loop parameters, such as the first and last instructions and / or exit conditions.

[0249] In one example, vector loops may be utilized, for example, together with high memory bandwidth and / or cheap addressing, for example, to solve control and data flow problems, for example, to provide a technical solution to allow data processing unit 316 to process data with substantially no additional overhead.

[0250] In some exemplary aspects, scalar processor 330 may be configured to provide one or more functionalities that may be complementary to the functionality of vector processing block 310. For example, a large portion (e.g., most) of the work in a vector program may be performed by data processing unit 316. For example, scalar processor 330 may be utilized, for example, to "glue" together various vector code blocks of a vector program.

[0251] In some exemplary aspects, the scalar processor 330 may be implemented separately from the vector processing block 310. In other aspects, the scalar processor 330 may be configured to share one or more components and / or functionality with the vector processing block 310.

[0252] In some exemplary aspects, scalar processor 330 may be configured to perform operations that may not be suitable for execution on vector processing block 310 .

[0253] For example, the scalar processor 330 may be utilized to execute a 32-bit C program. For example, the scalar processor 330 may be configured to support 1, 2, and / or 4-byte data types of the C code and / or some or all arithmetic operators of the C code.

[0254] For example, the scalar processor 330 may be configured to provide a technical solution to perform operations that cannot be performed on the vector processing block 310 without, for example, using a full CPU.

[0255] In some exemplary aspects, scalar processor 330 may include a scalar data memory 332 , for example, having a size of 16K or any other size, which may be configured to store data, such as variables used by a scalar portion of a program.

[0256] For example, scalar processor 330 may store local and / or global variables declared by portable C code, which may be compiled by a compiler (e.g., compiler 200 ( Figure 2 )) is allocated to scalar data memory.

[0257] In some exemplary aspects, such as Figure 3 As shown, the scalar processor 330 may include or may be associated with a set of vector registers 334 that may be used for data processing by the scalar processor 330 .

[0258] In some exemplary aspects, the scalar processor 330 can be associated with a scalar memory map that can enable the scalar processor 330 to access substantially all states of the vector processor 300. For example, the scalar processor 330 can configure a vector unit and / or a DMA channel via the scalar memory map.

[0259] In some exemplary aspects, the scalar processor 330 may not be allowed to access one or more block control registers that may be used by an external processor to run and debug a vector program.

[0260] In some exemplary aspects, DMA 340 can be configured to communicate, for example, via main memory, with one or more other components of a chip implementing vector processor 300. For example, DMA 340 can be configured to transfer blocks of data, for example, large, contiguous blocks of data, for example, to support scalar processor 330 and / or vector processing blocks that can manipulate data stored in local memory. For example, a vector program may be able to use DMA 340 to read data from main chip memory.

[0261] In some exemplary aspects, DMA 340 may be configured to communicate with other elements of the chip, for example, via a plurality of DMA channels (e.g., 8 DMA channels or any other count of DMA channels). For example, a DMA channel (e.g., each DMA channel) may be able to transfer a rectangular patch from a local memory to a main chip memory, or vice versa. In other aspects, a DMA channel may transfer any other type of data block between a local memory and a main chip memory.

[0262] In some exemplary aspects, a rectangular tile may be defined by a base pointer, a width, a height, and a stride.

[0263] For example, at peak throughput, 8 bytes may be transferred per cycle, however, there may be an overhead for each tile and / or for each row in a tile.

[0264] In some exemplary aspects, DMA 340 can be configured to transfer data in parallel with computations, such as via multiple DMA channels, for example, as long as the executed commands do not access local memory involved in the transfer.

[0265] In one example, since all channels can access the same memory bus, using several channels to implement a transfer may not save I / O cycles, for example, compared to when a single channel is used. However, multiple DMA channels can be utilized to schedule several transfers and execute them in parallel with the calculation. For example, this may be advantageous compared to a single channel, which may not allow a second transfer to be scheduled before the first transfer is completed.

[0266] In some exemplary aspects, the DMA 340 may be associated with a memory map that may support DMA channel access to vector memory and / or scalar data. For example, access to vector memory may occur in parallel with computing. For example, access to scalar data may generally not be allowed in parallel, e.g., because the scalar processor 330 may be involved in almost any reasonable program and may access its local variables during a transfer, which may result in memory contention with an active DMA channel.

[0267] In some exemplary aspects, the DMA 340 may be configured to provide a technical solution to support the parallelization of I / O and computing. For example, a program performing computations may not have to wait for I / O, e.g., in cases where these computations can be run quickly by the vector processing block 310.

[0268] In some exemplary aspects, an external processor (e.g., a CPU) may be configured to initiate the execution of a program on the vector processor 300. For example, the vector processor 300 may remain idle, e.g., as long as program execution has not been initiated.

[0269] In some exemplary aspects, an external processor may be configured to debug a program, e.g., execute a single step at a time, stop when the program reaches a breakpoint, and / or examine the contents of registers and memory storing program variables.

[0270] In some exemplary aspects, an external memory map may be implemented to support an external processor in controlling the vector processor 300 and / or debugging a program, e.g., by writing to the control registers of the vector processor 300.

[0271] In some exemplary aspects, the external memory map may be implemented as a superset of the scalar memory map. For example, this implementation may make all registers and memory defined by the architecture of the vector processor 300 accessible to a debugger backend running on an external processor.

[0272] In some exemplary aspects, the vector processor 300 may issue an interrupt signal, e.g., when the vector processor 300 terminates a program.

[0273] In some exemplary aspects, the interrupt signal may be used, e.g., to implement a driver to maintain a queue of programs scheduled for execution by the vector processor 300, and / or may be used, e.g., to initiate a new program by an external processor when a previously executed program has completed.

[0274] Return reference Figure 1In some exemplary aspects, compiler 160 may be configured to generate object code 115 that is configured to utilize registers of a processor (e.g., a vector processor, such as vector processor 180), such as according to a register allocation scheme, such as described below.

[0275] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce usage of allocated registers for executing a program by a processor (eg, a vector processor), for example, as described below.

[0276] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to utilizing a reduced (eg, optimized and / or minimized) number of allocated registers for executing a program by a processor (eg, a vector processor), for example, as described below.

[0277] In one example, the register allocation scheme can be configured to provide a technical solution to utilize the multiple vector registers 314 ( Figure 3 ) allocates a reduced (eg, optimized and / or minimized) number of allocated registers for use by the vector processor 300 ( Figure 3 ) executes a procedure, for example, as described below.

[0278] In some exemplary aspects, a compiler (e.g., compiler 160) may be configured to generate target code (e.g., target code 115) that may be configured to utilize registers of a vector processor (e.g., vector processor 180) according to a register allocation scheme, for example, as described below.

[0279] In other aspects, a compiler (e.g., compiler 160) may be configured to generate object code (e.g., object code 115) that may be configured to utilize registers of any other suitable type of processor according to a register allocation scheme, e.g., as described below.

[0280] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to improve (eg, to optimize) allocation of registers used to execute an executable program, eg, as described below.

[0281] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to support improved allocation (eg, efficient allocation, eg, optimized allocation) of registers used to execute executable programs, for example, as described below.

[0282] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to execute an executable program with a reduced number (eg, an optimized number, eg, a minimum number) of allocated registers, eg, as described below.

[0283] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to improve performance of an executable program, such as by reducing the number of allocated registers used to execute the executable program, for example, as described below.

[0284] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to reduce usage of allocated registers, such as by reducing the range of an active interval of one or more variables, for example, as described below.

[0285] In some exemplary aspects, the active scope of a variable may include a range of cycles of an executable program, e.g., between a first cycle (e.g., which includes a first use and / or generation of the variable) and a second cycle (e.g., which includes a second use of the variable after, for example, the first use), e.g., as described below.

[0286] In some exemplary aspects, it may be desirable to provide technical solutions to efficiently allocate registers of a processor (e.g., a vector processor) for execution of a program, for example, so as to reduce usage of allocated registers (e.g., allocated vector registers), for example, as described below.

[0287] For example, the number of physical registers implemented by a chip including a processor (e.g., a vector processor or any other processor) may be limited, e.g., based on the design and / or layout of the chip. For example, the number of vector registers implemented by a vector processor may be limited by the number of physical registers on the chip implementing the vector processor.

[0288] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to reduce (e.g., optimize and / or minimize) the number of allocated registers, for example, for a CPU (e.g., a vector processor) with limited storage capabilities (e.g., a limited register pool).

[0289] In some exemplary aspects, the register allocation scheme may be configured to provide technical solutions to reduce (e.g., optimize and / or minimize) the number of allocated registers, for example, for CPUs (e.g., vector processors) that have limited or no support for memory spill / fill operations (e.g., to store active values).

[0290] For example, a processor without fill / overflow capabilities and / or limited storage capabilities may be forced to use computational resources instead.

[0291] In one example, a processor may identify an unsuccessful register allocation according to instruction scheduling, for example, due to a limited number of registers, which may not be able to support register allocation according to instruction scheduling. One option to address this situation may be to relax instruction scheduling, for example, to try to reduce the number of active variables that share the same execution cycle. However, this option may result in performance degradation.

[0292] In some exemplary aspects, the register allocation scheme can be configured to provide a technical solution to reduce the number of allocated registers, for example, to avoid or even eliminate the use of these additional computing resources.

[0293] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce register shortage for executing a program by a processor (eg, a vector processor and / or any other processor).

[0294] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce register starvation, such as by reducing the number of allocated registers for execution of a program, for example, as described below.

[0295] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to support efficient execution of programs (e.g., complex programs that may be sensitive to insufficient registers). For example, for some programs, insufficient registers may be a bottleneck, and accordingly, insufficient registers may affect performance.

[0296] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to reduce (e.g., optimize and / or minimize) the number of allocated registers for execution of a program, for example, while providing a suitable allocation of registers (e.g., vector registers) for execution of the program, for example, as described below.

[0297] In some exemplary aspects, a register allocation scheme may be configured to provide a technical solution to reduce (eg, minimize and / or optimize) the range of an active interval for one or more variables, eg, as described below.

[0298] In some exemplary aspects, the reduction in the range of the active interval may provide a technical solution to support reduced register usage, and accordingly, to support reduction (eg, minimization and / or optimization) of register starvation.

[0299] For example, a register allocation scheme may be implemented to provide a technical solution to support efficient execution of programs that may be subject to register starvation problems.

[0300] In some exemplary aspects, the register allocation scheme may be configured to provide a technical solution to support improved performance of programs that may be executed, for example, by an exposed pipeline processor and / or any other additional or alternative type of processor.

[0301] In some exemplary aspects, compiler 160 may be configured to process, for example, a given instruction schedule based on source code 112, and generate target code 115 that may be configured to utilize a reduced (e.g., optimized and / or minimized) number of allocated registers, for example, for successful register allocation, e.g., as described below.

[0302] In some exemplary aspects, compiler 160 may be configured to identify a pair of instructions in a given instruction schedule, eg, the pair of instructions including a first instruction and a second instruction that may be independent and / or similar.

[0303] In some exemplary aspects, compiler 160 may be configured to determine whether to perform the first instruction and the second instruction in parallel using, for example, a single SIMD instruction, or using two separate instructions, for example, as described below.

[0304] In one example, in some use cases, implementations, and / or scenarios, the output and / or input of a SIMD instruction may have asymmetric usage. For example, the execution of such a SIMD instruction may impose one or more constraints on the instruction scheduler. For example, these constraints may be reflected in the resulting register shortage, for example, as described below.

[0305] In some exemplary aspects, compiler 160 may be configured to selectively assign a pair of instructions to be executed as SIMD instructions, e.g., based on a criterion that may be configured to reduce (e.g., minimize) active intervals corresponding to the pair of instructions, e.g., as described below.

[0306] In some exemplary aspects, the criteria may be configured to provide technical solutions to reduce (eg, minimize and / or optimize) register usage, eg, as described below.

[0307] In some exemplary aspects, compiler 160 may be configured to provide technical solutions to support reduced (e.g., minimal and / or optimal) use of registers allocated for execution of a program, e.g., for a given instruction schedule based on source code 112, e.g., as described below.

[0308] In some exemplary aspects, compiler 160 may be configured to combine a pair of instructions (eg, a pair of similar but independent instructions) into a SIMD instruction.

[0309] In some exemplary aspects, compiler 160 may be configured to generate efficient instruction schedules based on SIMD instructions, eg, as described below.

[0310] In some exemplary aspects, compiler 160 may be configured to determine whether efficient instruction scheduling may result in invalid register allocations, such as due to insufficient registers and / or any other reason, eg, as described below.

[0311] In some exemplary aspects, compiler 160 may be configured to identify valid instruction schedules that may produce invalid register allocations, such as due to insufficient registers and / or any other reason.

[0312] In some exemplary aspects, compiler 160 may be configured to attempt to find one or more identified SIMD instructions (e.g., SIMD output instructions) with asymmetric usage in an efficient instruction schedule, e.g., based on a determination that the efficient instruction schedule may produce an invalid register allocation, e.g., as described below.

[0313] In some exemplary aspects, compiler 160 may be configured to split identified SIMD instructions (eg, each identified SIMD instruction or only some identified SIMD instructions) into two or more instructions, eg, if possible, eg, as described below.

[0314] In some exemplary aspects, compiler 160 may be configured to generate a new instruction set including two or more instructions, and perform instruction scheduling and register allocation for the new instruction set, for example, as described below.

[0315] In some exemplary aspects, selectively splitting an identified SIMD instruction into multiple instructions may provide a technical solution to avoid asymmetry and / or remove asymmetry from the identified SIMD instruction, eg, as described below.

[0316] For example, removing asymmetry from identified SIMD instructions may provide a technical solution to release one or more constraints in instruction scheduling, eg, as described below.

[0317] In some demonstrative aspects, compiler 160 may be configured to identify, for example, based on source code 112 , first data operations and second data operations that can be performed in parallel, for example, as described below.

[0318] In some exemplary aspects, compiler 160 may be configured to identify a first data operation and a second data operation that can be executed in parallel, such as according to SIMD instructions, eg, as described below.

[0319] In some exemplary aspects, SIMD instructions may be configured to be executed, for example, by a single data processing unit (ALU) of a target processor (eg, processor 180), for example, as described below.

[0320] In one example, compiler 200 ( Figure 2 ) may be configured to identify a first data operation and a second data operation, the first data operation and the second data operation may be, for example, based on a vector processor 300 ( Figure 3 ) of a single data processing unit 316 ( Figure 3 ) to execute SIMD instructions in parallel.

[0321] In some exemplary aspects, the first data operation and the second data operation may comprise data operations of the same instruction, eg, as described below.

[0322] In some demonstrative aspects, the first data operation may include a data operation of a first instruction, eg, as described below.

[0323] In some demonstrative aspects, the second data operation may include, for example, a data operation of the second instruction separate from and / or independent of the first instruction, eg, as described below.

[0324] In some demonstrative aspects, compiler 160 may be configured to determine a selected coding scheme to be applied to compile the first data operation and the second data operation, eg, as described below.

[0325] In some demonstrative aspects, compiler 160 may be configured to determine a selected coding scheme, eg, from a first coding scheme and a second coding scheme, eg, as described below.

[0326] In some demonstrative aspects, compiler 160 may be configured to determine the selected coding scheme, eg, by selecting the selected coding scheme from the first coding scheme and the second coding scheme, eg, based on a predefined selection criterion, eg, as described below.

[0327] In some demonstrative aspects, the first compilation scheme may include compiling the first data operation and the second data operation into SIMD instructions, eg, as described below.

[0328] In some exemplary aspects, the second compilation scheme may include compiling the first data operation into a first ALU instruction, and compiling the second data operation into a second ALU instruction, eg, as described below.

[0329] In some exemplary aspects, the second ALU instruction can be configured to be executed separately from the first ALU instruction, eg, as described below.

[0330] In some exemplary aspects, compiler 160 may be configured to generate object code 115, such as based on compiling the first data operation and the second data operation according to a selected compilation scheme, such as described below.

[0331] In some exemplary aspects, compiler 160 may be configured to generate target code 115 configured for execution by, for example, a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor (eg, processor 180).

[0332] In other aspects, compiler 160 may be configured to generate object code 115 configured for execution by, for example, any other suitable type of processor.

[0333] In some exemplary aspects, compiler 160 may be configured to generate target code 115 based on source code 112 including Open Computing Language (OpenCL) code, for example.

[0334] In other aspects, compiler 160 may be configured to generate object code 115 based on source code 112 , including any other suitable type of code, for example.

[0335] In some exemplary aspects, compiler 160 may be configured to compile source code 112 into target code 115 , for example, according to a low-level virtual machine (LLVM)-based compilation scheme.

[0336] In other aspects, compiler 160 may be configured to compile source code 112 into target code 115 according to any other suitable compilation scheme.

[0337] In some exemplary aspects, the first ALU instruction may be configured to execute in a first execution cycle, eg, as described below.

[0338] In some exemplary aspects, the second ALU instruction can be configured to execute in a second execution cycle that is different from the first execution cycle, for example, as described below.

[0339] In some exemplary aspects, the second compilation scheme may include compiling the first data operation into a first ALU instruction to be executed, e.g., during a first cycle, and compiling the second data operation into a second ALU instruction to be executed, e.g., during a second cycle that may be different from the first cycle, e.g., as described below.

[0340] In some exemplary aspects, compiler 160 may be configured to identify an ALU to be assigned to execute a second ALU instruction, for example, based on a determination that the ALU is available during the second cycle, eg, as described below.

[0341] In some demonstrative aspects, the first ALU instruction may be configured to be executed by a first ALU of target processor 180, eg, as described below.

[0342] In some demonstrative aspects, the second ALU instructions may be configured to be executed by a second ALU of target processor 180, eg, as described below.

[0343] In some exemplary aspects, the first ALU instruction and the second ALU instruction may be configured to be executed by the same ALU of target processor 180, eg, as described below.

[0344] In some demonstrative aspects, compiler 160 may be configured to generate object code (eg, object code 115) that may be configured for execution by a target processor (eg, processor 180), eg, as described below.

[0345] In some exemplary aspects, compiler 160 may be configured to determine a selected compilation scheme for compilation of the first and second data operations, e.g., based on predefined selection criteria, which may include, for example, a register utilization criterion corresponding to, for example, register utilization of one or more registers of target processor 180, e.g., as described below.

[0346] In some exemplary aspects, the register utilization criterion may be based on register utilization of one or more registers of the target processor 180, e.g., according to the first compilation scheme, e.g., based on utilizing SIMD instructions to perform the first and second data operations, e.g., as described below.

[0347] In some exemplary aspects, the register utilization criteria may be configured, for example, such that, for example, when a first register utilization of one or more registers of the target processor 180 according to a first compilation scheme is greater than a second register utilization of one or more registers of the target processor 180 according to a second compilation scheme, the selected compilation scheme may include the second compilation scheme, e.g., as described below.

[0348] In some exemplary aspects, the compiler 160 may be configured to determine that the selected coding scheme is to include the second coding scheme, e.g., based on a determination that a first register utilization of one or more registers of the target processor 180 according to the first coding scheme is greater than a second register utilization of one or more registers of the target processor 180 according to the second coding scheme, e.g., as described below.

[0349] In some exemplary aspects, the compiler 160 may be configured to determine that the selected coding scheme will include the second coding scheme, e.g., based on a determination that a first register utilization of one or more registers of the target processor 180 according to the first coding scheme is greater than a predefined register utilization threshold, e.g., as described below.

[0350] In other aspects, the register utilization criteria may include any other additional or alternative criteria related to register utilization of the target processor 180 .

[0351] In some exemplary aspects, compiler 160 may be configured to determine the selected coding scheme, e.g., according to a predefined selection criterion, which may be based on, for example, an active scope of at least one variable corresponding to at least one of the first data operation and / or the second data operation, e.g., as described below.

[0352] In some demonstrative aspects, at least one variable corresponding to the first data operation and / or the second data operation may include, for example, at least one input variable of the first data operation and / or the second data operation, eg, as described below.

[0353] In some demonstrative aspects, at least one variable corresponding to the first data operation and / or the second data operation may include, for example, at least one output variable produced by the first data operation and / or the second data operation, eg, as described below.

[0354] In some exemplary aspects, the predefined selection criteria can be based on, for example, active ranges of input variables of the SIMD instruction, eg, according to a first compilation scheme, eg, as described below.

[0355] In some exemplary aspects, the predefined selection criteria can be based on, for example, an active range of an output variable produced by a SIMD instruction, eg, according to a first compilation scheme, eg, as described below.

[0356] In some exemplary aspects, the predefined selection criteria may be configured, for example, such that when, for example, a first active range of a variable according to a first coding scheme is greater than a second active range of a variable according to a second coding scheme, the selected coding scheme will include the second coding scheme, e.g., as described below.

[0357] In some demonstrative aspects, compiler 160 may be configured to determine that the selected coding scheme is to include the second coding scheme, e.g., based on a determination that a first active range of a variable according to the first coding scheme is greater than a second active range of the variable according to the second coding scheme, e.g., as described below.

[0358] In some demonstrative aspects, compiler 160 may be configured to determine that the selected coding scheme is to include a second coding scheme, e.g., based on a determination that an active range of a variable according to a first coding scheme is greater than a predefined active range threshold, e.g., as described below.

[0359] In some exemplary aspects, compiler 160 can be configured to determine the selected coding scheme, e.g., according to a predefined selection criterion, which can be based on, for example, a difference between an active range of a first variable and an active range of a second variable, e.g., as described below.

[0360] In some exemplary aspects, the first variable may be generated by executing a first data operation via a SIMD instruction, eg, according to a first compilation scheme, eg, as described below.

[0361] In some exemplary aspects, the second variable may be generated by executing a second data operation via a SIMD instruction, eg, according to a first compilation scheme, eg, as described below.

[0362] In some exemplary aspects, the first variable may include input for performing a first data operation via a SIMD instruction, eg, according to a first compilation scheme, eg, as described below.

[0363] In some exemplary aspects, the second variable may include input for performing a second data operation via a SIMD instruction, eg, according to a first compilation scheme, eg, as described below.

[0364] In some exemplary aspects, the predefined selection criteria may be based on register utilization corresponding to the first coding scheme, eg, as described below.

[0365] In some exemplary aspects, register utilization corresponding to the first compilation scheme can be based on, for example, a count of cycles in which registers will be occupied by results of SIMD instructions, eg, as described below.

[0366] In some exemplary aspects, register utilization corresponding to the first compilation scheme can be based on a cycle count between, for example, a first cycle (e.g., an initial cycle) in which a result of a SIMD instruction is to be available in a register and a second cycle, for example, after the first cycle, in which a result of the SIMD instruction is to be retrieved from the register, e.g., as described below.

[0367] In some exemplary aspects, register utilization corresponding to the first compilation scheme can be based on, for example, a count of cycles that registers will be occupied by input variables of a SIMD instruction, eg, as described below.

[0368] In some exemplary aspects, register utilization corresponding to the first compilation scheme can be based on a cycle count between, for example, a first cycle (e.g., an initial cycle) in which input variables of the SIMD instruction are to be available in registers and a second cycle, for example, after the first cycle, in which the input variables of the SIMD instruction are to be retrieved from registers, e.g., for execution of the SIMD instruction, e.g., as described below.

[0369] In other aspects, the predefined selection criteria may include any other additional or alternative criteria related to the active scope of the variable resulting from the first data operation and / or the second data operation.

[0370] In some exemplary aspects, the predefined selection criteria may be based on, for example, latency of an ALU of a target processor used to execute SIMD instructions, eg, as described below.

[0371] In other aspects, the predefined selection criteria may include any other additional or alternative criteria, such as relating to any other additional or alternative parameters and / or attributes.

[0372] In some exemplary aspects, compiler 160 may be configured to provide compiled code (eg, object code 115), for example, based on compilation of first and second data operations, for example, according to a selected compilation scheme, for example, as described below.

[0373] In some exemplary aspects, compiler 160 may be configured to determine the selected compilation scheme, e.g., according to a predefined selection criterion, which may be based on, for example, an active range of at least one output variable produced by execution of the SIMD instruction, e.g., according to a first compilation scheme, e.g., as described below.

[0374] In some exemplary aspects, compiler 160 may be configured to process source code 112 of a program for execution.

[0375] In one example, source code 112 may include an instruction schedule that may be configured to, for example, calculate an expression based on a plurality of variables, for example, the plurality of variables including four variables, denoted as a, b, c, d, for example, as follows:

[0376] (a·b·3)+c·d (1)

[0378] For example, variables a, b, c, d may be stored in four registers denoted R0, R1, R2, and R3 respectively.

[0379] For example, compiler 160 may be configured to compile source code 112 for a processor that supports SIMD instructions (SIMD processor), such as processor 180. For example, a processor may include a first ALU and a second ALU that may be configured to perform addition and multiplication instructions, such as with a latency of 2 cycles.

[0380] For example, a SIMD processor may have a memory unit that has a latency of 1 cycle for memory access operations (eg, a store operation to store a value to the memory unit or a load operation to load a value from the memory unit).

[0381] For example, the execution of Expression 1 can be implemented according to the following instruction scheduling that utilizes SIMD instructions, for example:

[0382]

[0383] Table (1)

[0384] As shown in Table 1, the instructions of the executed program can be performed using only the first ALU (ALU1).

[0385] As shown in Table 1, it may not be necessary to use the second ALU (ALU2) for any operation of the executed program.

[0386] As shown in Table 1, the execution of Expression 1 can be performed during 5 cycles.

[0387] As shown in Table 1, the instruction scheduling can include a SIMD instruction to be executed in cycle 0. The SIMD instruction can include a combination of a first data operation (e.g., %0 = a * b) and a second data operation (e.g., %1 = c * d) that can be executed in parallel by ALU1.

[0388] In some exemplary aspects, the compiler 160 can be configured to identify the active ranges of variables in the instruction scheduling according to Table 1, for example, as follows:

[0389] value Active range in cycles [start, end] (inclusive) %0 [2,2] %1 [2,4] %2 [4,4] %3 [6,6]

[0390] Table (2)

[0391] As shown in Table 2, the output variable %1 that can be generated by the execution of the second data operation in the SIMD instruction can be active between cycle 2 and cycle 4. The active range of the variable %1 may require occupying a register to store the value of the variable %1, for example, between cycle 2 and cycle 4.

[0392] In one instance, the requirement to reserve a register for storing the variable %1 between cycle 2 and cycle 4 can result in a register shortage.

[0393] In some exemplary aspects, the compiler 160 can identify that using SIMD instructions can result in a register shortage.

[0394] In some exemplary aspects, the compiler 160 can identify that SIMD instructions can result in an asymmetric use of registers. For example, the compiler 160 can identify that SIMD instructions can generate an output variable %0 that can be active only in cycle 2, while the output variable %1 can be required to remain active between cycle 2 and 4.

[0395] In some exemplary aspects, compiler 160 may be configured to determine that it may be preferred to compile a first data operation (e.g., %0=a*b) into a first instruction (first ALU instruction) that can be executed, for example, by a first ALU (ALU1), and to compile a second data operation (e.g., %1=c*d) into a second instruction (second ALU instruction) that can be executed, for example, by a second ALU (ALU2).

[0396] In some exemplary aspects, compiler 160 may be configured to determine that it may be preferable to split a SIMD instruction into a first ALU instruction to be executed by a first ALU and a second ALU instruction to be executed by a second ALU, e.g., as described below.

[0397] In some exemplary aspects, compiler 160 may be configured to generate another instruction schedule including a first ALU instruction to be executed by a first ALU (ALU1) and a second ALU instruction to be executed by a second ALU (ALU2), for example, as follows:

[0398]

[0399]

[0400] Table (3)

[0401] As shown in Table 3, a first ALU instruction (eg, %0=a*b) may be executed, for example, by the first ALU during cycle 0.

[0402] As shown in Table 3, the second ALU instruction (eg, %1=c*d) may be executed, for example, by the second ALU during cycle 2.

[0403] In some exemplary aspects, one or more of the variables scheduled according to the instructions of Table 3 may have an active range that may be different than, for example, the active range of Table 2, for example, as follows.

[0404] value Active range in cycles [start, end] (inclusive) %0 [2,2] %1 [4,4] %2 [4,4] %3 [6,6]

[0405] Table (4)

[0406] As shown in Table 4, variable %1 may be active only during cycle 4, and / or may be idle during other cycles.

[0407] For example, variable %1 may occupy a register only during cycle 4.

[0408] Accordingly, the instruction schedule of Table 3 may provide a technical solution to reduce register utilization, reduce register starvation, and / or reduce the number of registers used.

[0409] In some exemplary aspects, the number of cycles of a program executed according to the instruction set of Table 1 may be equal to the number of cycles of a program executed according to the instruction set of Table 3, eg, 5 cycles.

[0410] In some exemplary aspects, a program executed according to the instruction set of Table 3 may utilize a reduced number of registers, such as compared to a program executed according to the instruction set of Table 1.

[0411] Accordingly, the instruction set of Table 3 may be implemented to provide a technical solution to improve the performance of an executed program.

[0412] In some exemplary aspects, compiler 160 may be configured to determine a selected compilation scheme, e.g., according to a predefined selection criterion, which may be based on, for example, an active range of at least one input variable to be input for execution of the SIMD instruction, e.g., according to a first compilation scheme, e.g., as described below.

[0413] In one example, source code 112 may include an instruction schedule that may be configured to compute a first expression (data operation) and a second expression (data operation) based on a plurality of variables, for example, including two variables (denoted as a, b), for example, as follows:

[0414] (3–a)·b(2)

[0415] 4·a (3)

[0416] In some exemplary aspects, compiler 160 may be configured to load variable a, for example, from a first memory location (a_add), and / or to load variable b, for example, from a second memory location (b_add).

[0417] In some exemplary aspects, compiler 160 may be configured to store variables a, b in two registers, denoted R0 and R1, respectively.

[0418] For example, compiler 160 may be configured to compile source code 112 for a processor (eg, a SIMD processor) that may include a single ALU (eg, ALU1).

[0419] For example, the ALU of a processor may be configured to perform SIMD instructions, such as add, multiply, and minus instructions, with a latency of 2 cycles, for example.

[0420] For example, a SIMD processor may have a memory unit that has a latency of 1 cycle for memory access operations (eg, a store operation to store a value to the memory unit or a load operation to load a value from the memory unit).

[0421] For example, the execution of Expressions 2 and 3 may be implemented according to the following instruction schedule using, for example, SIMD instructions:

[0422]

[0423] Table (5)

[0424] As shown in Table 5, a single ALU1 may be used to execute instructions of a program.

[0425] As shown in Table 5, the execution of Expression 2 and Expression 3 can be performed during 9 cycles.

[0426] As shown in Table 5, the instruction schedule may include a SIMD instruction to be executed in cycle 5.

[0427] For example, as shown in Table 5, a SIMD instruction may include a combination of a first instruction (eg, %2=%1*b=(3-a)*b) and a second instruction (eg, %3=4*a) that may be executed in parallel by ALU1.

[0428] In some exemplary aspects, compiler 160 may be configured to identify the active scope of a variable in instruction scheduling, and / or assign registers to the variable, for example, according to Table 5, for example, as follows:

[0429]

[0430] Table (6)

[0431] As shown in Table 6, register R0 may be occupied between cycle 1 and cycle 5, for example, to store a variable a, which may be used as an input for a second instruction in the SIMD instructions.

[0432] As shown in Table 6, register R1 may be occupied during cycle 5, for example, to store variable b.

[0433] As shown in Table 6, variable %1, which may be generated by data operation %1=add(%0,3), which may be used as input for the first instruction in the SIMD instruction, may be active during cycle 5. For example, this may require occupying a register (e.g., register R2) to store the value of variable %1, for example, during cycle 5. For example, variable %1 may occupy register R2, for example, because register R0 and register R1 may be occupied during cycle 5.

[0434] In one example, the requirement to reserve a register for storing variable %1 during cycle 5 may create a register shortage.

[0435] In some exemplary aspects, compiler 160 may recognize that using SIMD instructions may create a register shortage.

[0436] In some exemplary aspects, compiler 160 may recognize that a SIMD instruction generates asymmetric usage of registers, for example, input variable b is active only in cycle 5, while input variable a is active between cycles 1-5.

[0437] In some exemplary aspects, compiler 160 may be configured to determine that it may be preferred to compile a second data operation (e.g., %3=4*a) into a first multiplication instruction to be executed by ALU1, and to compile a first data operation (e.g., %2=(3-a)*b) into a second multiplication instruction to be executed by ALU1, e.g., after the first multiplication instruction.

[0438] In some exemplary aspects, compiler 160 may be configured to determine that it may be preferable to split the SIMD instruction into a first multiplication instruction and a second multiplication instruction.

[0439] In some exemplary aspects, compiler 160 may be configured to generate another instruction schedule including a first multiplication instruction to be executed by ALU1 and a second multiplication instruction to be executed by ALU1, e.g., after the first multiplication instruction, e.g., as follows:

[0440] 0 a=load(a_add) 1 %0=neg(a) 2 %3,_=mul(4,a,_,_) "_" is an unused symbol 3 %1=add(%0,3) b=load(b_add) 4 store(%3,res2_add) 5 %2,_=mul(%1,b,_,_) %2=(3–a)·b 6 7 store(%2,res1_add)

[0441] Table (7)

[0442] As shown in Table 7, the first multiplication instruction (eg, %3=4*a) may be executed by ALU1 during cycle 2.

[0443] As shown in Table 7, the second multiplication instruction (eg, %2=%1*b=(3-a)*b) may be executed by ALU1 during cycle 5.

[0444] In some exemplary aspects, variables scheduled according to the instructions of Table 7 may have an active range that may be different from, for example, the active range of Table 6, and / or registers assigned to variables according to the instruction scheduling of Table 7 may be different from, for example, assigning registers according to Table 6, e.g., as follows.

[0445]

[0446] Table (8)

[0447] As shown in Table 8, variable a may be active between cycle 1 and cycle 2, and variable b may be active between cycle 4 and cycle 5. For example, separate, non-overlapping active ranges of variables a and b may be utilized, for example, to specify variable a and variable b to occupy the same register, for example, R0, for example, between cycle 1 and cycle 5, instead of occupying two registers, for example, registers R0 and R1. Accordingly, the instruction schedule of Table 7 may be implemented to provide a technical solution to utilize register R1, for example, instead of register R2, for example, to store variable %1 during cycle 5.

[0448] Accordingly, the instruction schedule of Table 7 may provide a technical solution to reduce register utilization, reduce register starvation, and / or reduce the number of registers used.

[0449] In some exemplary aspects, the number of cycles (eg, 8 cycles) of a program executed according to the instruction set of Table 7 may be less than the number of cycles (eg, 9 cycles) of a program executed according to the instruction set of Table 5.

[0450] In some exemplary aspects, a program executed according to the instruction set of Table 7 may utilize a reduced number of registers, eg, 2 registers, eg, as compared to a program executed according to the instruction set of Table 5 (eg, 3 registers).

[0451] Accordingly, the instruction set of Table 7 may be implemented to provide a technical solution to both improve the performance of an executed program and improve register starvation of an executed program.

[0452] In one example, attempting to schedule the execution of expressions 2 and 3 according to the instruction schedule of Table 5 may result in unsuccessful register allocation, for example, if only two registers are available. One option to address this situation may be to relax the instruction schedule, for example, to try to reduce the number of active variables that share the same execution cycle. However, this option may result in performance degradation.

[0453] In some exemplary aspects, execution of Expressions 2 and 3 according to the above-described register allocation scheme (e.g., using the instruction scheduling of Table 7) may provide a technical solution to support successful register allocation, e.g., even when only two registers are available, e.g., while avoiding the performance degradation caused by the instruction scheduling of Table 4.

[0454] refer to Figure 4 , which schematically illustrates a method of compiling code for a processor. For example, Figure 4 One or more operations of the method may be performed by: a system, for example, system 100 ( Figure 1 ); devices, for example, device 102 ( Figure 1 ); a server, for example, server 170 ( Figure 1); and / or a compiler, for example, compiler 160 ( Figure 1 ) and / or compiler 200( Figure 2 ).

[0455] In some exemplary aspects, as indicated at block 402, the method may include processing code to identify one or more pairs of data operations and / or instructions, e.g., similar and independent data operations and / or instructions, that may potentially be combined into SIMD instructions. For example, compiler 160( Figure 1 ) can identify a pair of similar independent instructions that can potentially be combined into a single SIMD, for example, as described above.

[0456] In some exemplary aspects, as indicated at block 404, the method may include performing instruction scheduling and register allocation, for example, using SIMD instructions. For example, compiler 160 ( Figure 1 ) may use SIMD instructions for instruction scheduling and register allocation, for example, as described above.

[0457] In some exemplary aspects, as indicated at block 406, the method may include determining whether the SIMD instruction can (eg, should) be replaced by two separate instructions, eg, as described below.

[0458] In some exemplary aspects, as indicated at block 406, the method may include observing instructions that are combined into SIMD instructions, e.g., each of the instructions, e.g., to determine whether a single SIMD has an asymmetric use definition chain, e.g., such that parallel execution of the pair of instructions adds constraints to the schedule, which may cause a large active range. For example, compiler 160( Figure 1 ) can determine whether instructions in SIMD instructions cause a large active range, for example, as described above.

[0459] In some exemplary aspects, as indicated at block 406, the method may include splitting the SIMD instruction into separate instructions, e.g., based on a determination that available resources (e.g., ALU resources) exist to perform the separate instructions. For example, if it is determined that ALU resources are available to perform the separate instructions, the compiler 160 ( Figure 1 ) can split SIMD instructions into separate instructions, for example, as described above.

[0460] In some exemplary aspects, as indicated at block 408, the method may include, for example, generating instruction scheduling and register allocation based on the individual instructions after splitting the SIMD instructions, for example, to reduce the range of the active interval. For example, the compiler 160 ( Figure 1 ) may generate target code 115 based on instruction scheduling and register allocation that may be configured to reduce the size of the active range of variable %1 ( Figure 1 ), for example, as described above.

[0461] refer to Figure 5 , which schematically illustrates a method of compiling code for a processor. For example, Figure 5 One or more operations of the method may be performed by: a system, for example, system 100 ( Figure 1 ); devices, for example, device 102 ( Figure 1 ); a server, for example, server 170 ( Figure 1 ); and / or a compiler, for example, compiler 160 ( Figure 1 ) and / or compiler 200( Figure 2 ).

[0462] In some exemplary aspects, as indicated at block 502, the method may include identifying, for example, based on source code, a first data operation and a second data operation that can be executed in parallel according to a SIMD instruction to be executed, for example, by a single ALU of a target processor. For example, compiler 160( Figure 1 ) may be configured, for example, based on source code 112 ( Figure 1 ) to identify the first and second data operations, for example, as described above.

[0463] In some exemplary aspects, as indicated at block 504, the method may include determining a selected compilation scheme from the first compilation scheme and the second compilation scheme based on a predefined selection criterion. For example, the first compilation scheme may include compiling the first data operation and the second data operation into SIMD instructions, and the second compilation scheme may include compiling the first data operation into a first ALU instruction and compiling the second data operation into a second ALU instruction to be executed separately from the first ALU instruction. For example, compiler 160( Figure 1 ) may be configured to identify a selected coding scheme to compile the first and second data operations, for example, as described above.

[0464] In some exemplary aspects, as indicated at block 506, the method may include generating target code that may be configured for execution by a target processor. For example, the target code may be based on compiling the first data operation and the second data operation according to the selected compilation scheme. For example, the compiler 160( Figure 1 ) may be configured to generate target code 115 (e.g., by compiling the first and second data operations according to a selected compilation scheme) Figure 1 ), for example, as described above.

[0465] refer to Figure 6, which schematically illustrates an article of manufacture 600 according to some exemplary aspects. Article 600 may include one or more tangible computer-readable ("machine-readable") non-transitory storage media 602, which may include, for example, computer-executable instructions implemented by logic 604, which are operable to enable, when executed by at least one computer processor, at least one computer processor to perform operations on a device 102 ( Figure 1 )、Server 170( Figure 1 ) and / or compiler 160( Figure 1 ) to implement one or more operations to enable device 102 ( Figure 1 )、Server 170( Figure 1 ) and / or compiler 160( Figure 1 ) performs, triggers and / or implements one or more operations and / or functionalities, and / or performs, triggers and / or implements reference Figures 1 to 5 One or more operations and / or functionalities described, and / or one or more operations described herein. The phrases "non-transitory machine-readable medium" and "computer-readable non-transitory storage medium" may be directed to include all computer-readable media, with the sole exception of transitory propagating signals.

[0466] In some exemplary aspects, the product 600 and / or the machine-readable storage medium 602 may include one or more types of computer-readable storage media capable of storing data, including volatile memory, non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. For example, the machine-readable storage medium 602 may include RAM, DRAM, double data rate DRAM (DDR-DRAM), SDRAM, static RAM (SRAM), ROM, programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory (e.g., NOR or NAND flash memory), content addressable memory (CAM), polymer memory, phase change memory, ferroelectric memory, silicon-oxide-nitride-oxide-silicon (SONOS) memory, disk, hard drive, etc. The computer-readable storage medium may include any suitable medium involved in downloading or transferring a computer program from a remote computer to a requesting computer via a communication link (e.g., a modem, radio, or network connection), the computer program being carried by a data signal embedded in a carrier wave or other propagation medium.

[0467] In some exemplary aspects, logic 604 may include instructions, data, and / or code that, if executed by a machine, may cause the machine to perform methods, processes, and / or operations as described herein. The machine may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, etc., and may be implemented using any suitable combination of hardware, software, firmware, etc.

[0468] In some exemplary aspects, logic 604 may include or may be implemented as software, a software module, an application, a program, a subroutine, an instruction, an instruction set, a computing code, a word, a value, a symbol, etc. Instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Instructions may be implemented according to a predefined computer language, manner, or syntax for instructing a processor to perform a specific function. Instructions may be implemented using any suitable high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language, machine code, etc.

[0469] Examples

[0470] The following examples relate to further aspects.

[0471] Example 1 includes a product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions operable to, when executed by at least one processor, enable the at least one processor to cause a computing device to: identify, based on source code, a first data operation and a second data operation that can be executed in parallel according to a single instruction / multiple data (SIMD) instruction to be executed by a single arithmetic logic unit (ALU) of a target processor; determine a selected compilation scheme from a first compilation scheme and a second compilation scheme based on a predefined selection criterion, wherein the first compilation scheme includes compiling the first data operation and the second data operation into SIMD instructions, wherein the second compilation scheme includes compiling the first data operation into a first ALU instruction, and compiling the second data operation into a second ALU instruction to be executed separately from the first ALU instruction; and generate a target code configured for execution by the target processor, the target code being based on compilation of the first data operation and the second data operation according to the selected compilation scheme.

[0472] Example 2 includes the subject matter of example 1, and optionally wherein the predefined selection criteria includes a register utilization criterion corresponding to register utilization of one or more registers of the target processor.

[0473] Example 3 includes the subject matter of example 2, and optionally wherein the register utilization criterion is based on register utilization of one or more registers of the target processor according to the first compilation scheme.

[0474] Example 4 includes the subject matter of example 2 or 3, and optionally wherein the register utilization criterion is configured such that when a first register utilization of one or more registers of the target processor according to the first compilation scheme is greater than a second register utilization of one or more registers of the target processor according to the second compilation scheme, the selected compilation scheme includes the second compilation scheme.

[0475] Example 5 includes the subject matter of any of Examples 1 to 4, and optionally wherein the predefined selection criteria is based on an active range of at least one variable corresponding to at least one of the first data operation or the second data operation.

[0476] Example 6 includes the subject matter of Example 5, and optionally wherein the at least one variable comprises at least one of: an input variable of the first data operation or the second data operation or an output variable produced by the first data operation or the second data operation.

[0477] Example 7 includes the subject matter of example 5 or 6, and optionally wherein the predefined selection criterion is based on an active range of a SIMD instruction or an output variable produced by the SIMD instruction.

[0478] Example 8 includes the subject matter of any of Examples 5 to 7, and optionally wherein the predefined selection criteria is configured such that when a first active range of the variable according to the first coding scheme is greater than a second active range of the variable according to the second coding scheme, the selected coding scheme is to include the second coding scheme.

[0479] Example 9 includes the subject matter of any of Examples 1 to 8, and optionally wherein the predefined selection criterion is based on a cycle count between a first cycle in which a result of the SIMD instruction is to be available in a register of the target processor and a second cycle after the first cycle in which the result of the SIMD instruction is to be retrieved from the register of the target processor.

[0480] Example 10 includes the subject matter of any of Examples 1 to 9, and optionally wherein the predefined selection criteria is based on a cycle count between a first cycle in which input variables of the SIMD instruction are to be available in registers of the target processor and a second cycle after the first cycle in which the input variables of the SIMD instruction are to be retrieved from registers of the target processor.

[0481] Example 11 includes the subject matter of any of Examples 1 to 10, and optionally wherein the predefined selection criteria is based on a difference between an active range of a first variable and an active range of a second variable, the first variable being generated by performing a first data operation through a SIMD instruction, and the second variable being generated by performing a second data operation through a SIMD instruction.

[0482] Example 12 includes the subject matter of any of Examples 1 to 11, and optionally wherein the predefined selection criteria is based on a difference between an active range of a first variable and an active range of a second variable, the first variable comprising an input for performing a first data operation via a SIMD instruction, and the second variable comprising an input for performing a second data operation via a SIMD instruction.

[0483] Example 13 includes the subject matter of any of Examples 1-12, and optionally wherein the predefined selection criteria is based on latency of an ALU of a target processor to execute the SIMD instruction.

[0484] Example 14 includes the subject matter of any of Examples 1-13, and optionally wherein the first ALU instruction is configured to execute in a first execution cycle, and the second ALU instruction is configured to execute in a second execution cycle different from the first execution cycle.

[0485] Example 15 includes the subject matter of any of Examples 1 to 14, and optionally wherein the second compilation scheme includes compiling the first data operation into a first ALU instruction to be executed during a first execution cycle, and compiling the second data operation into a second ALU instruction to be executed during a second cycle different from the first cycle.

[0486] Example 16 includes the subject matter of any of Examples 1-15, and optionally wherein the first ALU instruction is to be executed by the first ALU, and the second ALU instruction is to be executed by the second ALU.

[0487] Example 17 includes the subject matter of any of Examples 1-15, and optionally wherein the first ALU instruction and the second ALU instruction are to be executed by the same ALU.

[0488] Example 18 includes the subject matter of any of Examples 1 to 17, and optionally wherein the first data operation and the second data operation comprise data operations of the same instruction.

[0489] Example 19 includes the subject matter of any of Examples 1 to 17, and optionally wherein the first data operation comprises a data operation of a first instruction, and the second data operation comprises a data operation of a second instruction separate from the first instruction.

[0490] Example 20 includes the subject matter of any of Examples 1-19, and optionally wherein the source code includes Open Computing Language (OpenCL) code.

[0491] Example 21 includes the subject matter of any one of Examples 1 to 20, and optionally wherein the computer-executable instructions, when executed, cause a computing device to compile source code into target code according to a low-level virtual machine (LLVM)-based compilation scheme.

[0492] Example 22 includes the subject matter of any of Examples 1 to 21, and optionally wherein the target code is configured for execution by a very long instruction word (VLIW) single instruction / multiple data (SIMD) target processor.

[0493] Example 23 includes the subject matter of any one of Examples 1 to 22, and optionally wherein the object code is configured for execution by a target vector processor.

[0494] Example 24 includes a compiler configured to perform any of the operations described in any of Examples 1 to 23.

[0495] Example 25 includes a computing device configured to perform any of the operations described in any of Examples 1 to 23.

[0496] Example 26 includes a computing system comprising: at least one memory for storing instructions; and at least one processor for retrieving instructions from the memory and executing the instructions to cause the computing system to perform any of the operations described in any of Examples 1 to 23.

[0497] Example 27 includes a computing system comprising: a compiler configured to generate a target code according to any of the operations described in any of Examples 1 to 23; and a processor configured to execute the target code.

[0498] Example 28 includes a device comprising means for performing any of the operations described in any of Examples 1-23.

[0499] Example 29 includes an apparatus comprising: a memory interface; and a processing circuit system configured to: perform any of the operations described in any of Examples 1-23.

[0500] Example 30 includes a method comprising any of the operations described in any of Examples 1 to 23.

[0501] The functions, operations, components, and / or features described herein with reference to one or more aspects may be combined with or utilized in combination with one or more other functions, operations, components, and / or features described herein with reference to one or more other aspects, or vice versa.

[0502] While certain features have been illustrated and described herein, numerous modifications, substitutions, changes, and equivalents will occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the disclosure.

Claims

1. A product comprising one or more tangible computer-readable non-transitory storage media containing computer-executable instructions operable to, when executed by at least one processor, enable the at least one processor to cause a computing device to: identifying, based on the source code, a first data operation and a second data operation that are executable in parallel according to a single instruction / multiple data (SIMD) instruction to be executed by a single arithmetic logic unit (ALU) of a target processor; determining a selected compilation scheme from a first compilation scheme and a second compilation scheme based on a predefined selection criterion, wherein the first compilation scheme includes compiling the first data operation and the second data operation into the SIMD instruction, wherein the second compilation scheme includes compiling the first data operation into a first ALU instruction, and compiling the second data operation into a second ALU instruction to be executed separately from the first ALU instruction; as well as Object code configured for execution by the target processor is generated, the object code being based on compilation of the first data operation and the second data operation according to the selected compilation scheme.

2. The product of claim 1, wherein the predefined selection criteria comprises a register utilization criteria corresponding to a register utilization of one or more registers of the target processor.

3. The product of claim 2, wherein the register utilization criterion is based on the register utilization of the one or more registers of the target processor according to the first compilation scheme.

4. The product of claim 2, wherein the register utilization criterion is configured such that when a first register utilization of the one or more registers of the target processor according to the first compilation scheme is greater than a second register utilization of the one or more registers of the target processor according to the second compilation scheme, the selected compilation scheme will include the second compilation scheme. 5 . The product of claim 1 , wherein the predefined selection criteria is based on an active range of at least one variable corresponding to at least one of the first data operation or the second data operation.

6. The product of claim 5, wherein the at least one variable comprises at least one of: an input variable to the first data operation or the second data operation or an output variable produced by the first data operation or the second data operation.

7. The product of claim 5, wherein the predefined selection criteria is based on an active range of an input variable of the SIMD instruction or an output variable produced by the SIMD instruction.

8. The product of claim 5, wherein the predefined selection criterion is configured such that when a first active range of the variable according to the first coding scheme is greater than a second active range of the variable according to the second coding scheme, the selected coding scheme is to include the second coding scheme.

9. The product of any one of claims 1 to 8, wherein the predefined selection criterion is based on a cycle count between a first cycle in which a result of the SIMD instruction is to be available in a register of the target processor and a second cycle after the first cycle in which the result of the SIMD instruction is to be retrieved from the register of the target processor.

10. The product of any one of claims 1 to 8, wherein the predefined selection criterion is based on a cycle count between a first cycle in which input variables of the SIMD instruction are to be available in registers of the target processor and a second cycle after the first cycle in which the input variables of the SIMD instruction are to be retrieved from the registers of the target processor.

11. The product of any one of claims 1 to 8, wherein the predefined selection criterion is based on a difference between an active range of a first variable resulting from executing the first data operation by the SIMD instruction and an active range of a second variable resulting from executing the second data operation by the SIMD instruction.

12. The product of any one of claims 1 to 8, wherein the predefined selection criterion is based on a difference between an active range of a first variable and an active range of a second variable, the first variable comprising an input for performing the first data operation by the SIMD instruction, and the second variable comprising an input for performing the second data operation by the SIMD instruction.

13. The product of any one of claims 1 to 8, wherein the predefined selection criteria is based on latency of the ALU of the target processor to execute the SIMD instruction.

14. The product of any one of claims 1 to 8, wherein the first ALU instruction is configured to execute in a first execution cycle, and the second ALU instruction is configured to execute in a second execution cycle different from the first execution cycle.

15. The product of any one of claims 1 to 8, wherein the second compilation scheme comprises compiling the first data operation into the first ALU instruction to be executed during a first execution cycle, and compiling the second data operation into the second ALU instruction to be executed during a second cycle different from the first cycle.

16. The product of any one of claims 1 to 8, wherein the first ALU instruction is to be executed by a first ALU and the second ALU instruction is to be executed by a second ALU.

17. The product of any one of claims 1 to 8, wherein the first ALU instruction and the second ALU instruction are to be executed by the same ALU.

18. The product of any one of claims 1 to 8, wherein the first data operation and the second data operation comprise data operations of the same instruction.

19. The product of any one of claims 1 to 8, wherein the first data operation comprises a data operation of a first instruction, and the second data operation comprises a data operation of a second instruction separate from the first instruction.

20. The product of any one of claims 1 to 8, wherein the source code comprises Open Computing Language (OpenCL) code.

21. The product of any one of claims 1 to 8, wherein the computer-executable instructions, when executed, cause the computing device to compile the source code into the target code according to a Low Level Virtual Machine (LLVM)-based compilation scheme.

22. The product of any one of claims 1 to 8, wherein the object code is configured for execution by a Very Long Instruction Word (VLIW) Single Instruction / Multiple Data (SIMD) target processor.

23. A product according to any one of claims 1 to 8, wherein the object code is configured for execution by a target vector processor.

24. A computing system, include: at least one memory for storing instructions; as well as at least one processor to retrieve the instructions from the memory and to execute the instructions to cause the computing system to: identifying, based on the source code, a first data operation and a second data operation that are executable in parallel according to a single instruction / multiple data (SIMD) instruction to be executed by a single arithmetic logic unit (ALU) of a target processor; determining a selected compilation scheme from a first compilation scheme and a second compilation scheme based on a predefined selection criterion, wherein the first compilation scheme includes compiling the first data operation and the second data operation into the SIMD instruction, wherein the second compilation scheme includes compiling the first data operation into a first ALU instruction, and compiling the second data operation into a second ALU instruction to be executed separately from the first ALU instruction; as well as Object code configured for execution by the target processor is generated, the object code being based on compilation of the first data operation and the second data operation according to the selected compilation scheme.

25. The computing system of claim 24, wherein the predefined selection criteria comprises a register utilization criteria corresponding to register utilization of one or more registers of the target processor.

26. A computing system according to claim 24 or 25, comprising the target processor.

27. A method wherein include: identifying, based on the source code, a first data operation and a second data operation that are executable in parallel according to a single instruction / multiple data (SIMD) instruction to be executed by a single arithmetic logic unit (ALU) of a target processor; determining a selected compilation scheme from a first compilation scheme and a second compilation scheme based on a predefined selection criterion, wherein the first compilation scheme includes compiling the first data operation and the second data operation into the SIMD instruction, wherein the second compilation scheme includes compiling the first data operation into a first ALU instruction, and compiling the second data operation into a second ALU instruction to be executed separately from the first ALU instruction; as well as Object code configured for execution by the target processor is generated, the object code being based on compilation of the first data operation and the second data operation according to the selected compilation scheme.

28. The method of claim 27, wherein the predefined selection criteria is based on an active range of at least one variable corresponding to at least one of the first data operation or the second data operation.