Generation and use of memory access instruction order encodings

By using a block-based processor architecture and an explicit data graphics execution architecture, the limitations of existing ISA architectures in terms of performance and energy efficiency are addressed, achieving efficient intra-instruction block communication and energy savings, thereby improving processor performance and compiler efficiency.

CN115390926BActive Publication Date: 2026-03-17MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2016-09-13
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing processor ISA architectures have limited improvements in areas such as transistor expansion, integrated circuit cost, manufacturing capital, and energy efficiency, and out-of-order superscalar implementations have failed to continuously improve performance.

Method used

It adopts a block-based processor architecture (BB-ISA) and uses the explicit data graph execution (EDGE) architecture to generate and compile code by utilizing the relative ordering encoding of memory access instructions. This reduces the complexity of register renaming and data flow analysis, supports high instruction-level parallelism and out-of-order execution, and enables efficient intra-instruction block communication.

Benefits of technology

It improves processor performance and energy efficiency, reduces energy consumption, supports mainstream programming languages, and improves compiler and processor performance, avoiding complex hardware and software overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115390926B_ABST
    Figure CN115390926B_ABST
Patent Text Reader

Abstract

Apparatuses and methods are disclosed for controlling execution of memory access instructions in a block-based processor architecture using a hardware structure that indicates the relative ordering of memory access instructions in an instruction block. In one example of the disclosed technology, a method of executing an instruction block having a plurality of memory load and / or memory store instructions includes selecting a next memory load or memory store instruction to execute based on a dependency encoded within the block and a store vector to memory data indicating which memory load and memory store instructions in the instruction block have already executed. The store vector can be masked using a store mask. The store mask can be generated at the time the instruction block is decoded or copied from an instruction block header. Based on the encoded dependency and the masked store vector, the next instruction can be issued when its dependency is available.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese invention patent application filed on September 13, 2016, with application number 201680054500.0 and title "Generation and Use of Sequential Encoding for Memory Access Instructions". Technical Field

[0002] Microprocessors have benefited from the continuous increase in transistor count, integrated circuit costs, manufacturing capital, clock frequency, and energy efficiency, as predicted by Moore's Law, while the associated processor instruction set architecture (ISA) has seen minimal changes. However, the benefits from the lithographic expansion that has driven the semiconductor industry over the past 40 years are slowing or even reversing. Reduced Instruction Set Computing (RISC) architecture has been the dominant paradigm in processor design for many years. Out-of-order superscalar implementations have not yet shown sustained improvements in area or performance. Therefore, there are ample opportunities for improvements to the processor ISA that extend performance. Summary of the Invention

[0003] Methods, apparatus, and computer-readable storage devices for configuring, operating, and compiling code for block-based processor architectures (BB-ISA), including explicit data graphics execution (EDGE) architectures, are disclosed. The described techniques and tools for solutions such as improving processor performance and / or reducing power consumption can be implemented separately or in various combinations thereof. As will be described more fully below, the described techniques and tools can be implemented in digital signal processors, microprocessors, application-specific integrated circuits (ASICs), soft processors (e.g., multiprocessor cores implemented in field-programmable gate arrays (FPGAs) using reconfigurable logic), programmable logic, or other suitable logic circuitry. As will be readily apparent to those skilled in the art, the disclosed techniques can be implemented in a variety of computing platforms, including but not limited to servers, mainframes, cellular phones, smartphones, PDAs, handheld devices, handheld computers, PDAs, touchscreen tablet devices, tablet computers, wearable computers, and laptop computers.

[0004] In one example of the disclosed technology, the block-based processor is configured to control the order of memory access instructions (e.g., memory load and memory store instructions) based on a hardware structure of stored data indicating the relative ordering of the memory access instructions. In some examples, a memory mask is generated during instruction block decoding or read directly from the instruction block header and stored in the hardware structure. In some examples, memory access instructions are encoded using identifiers indicating their relative ordering. In some examples, a compiler or interpreter transforms source code and / or object code into executable code for the block-based processor, including memory access instructions encoded using ordering identifiers and / or memory mask information.

[0005] The present invention is provided to introduce a selection of concepts in simplified forms, which are further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. The foregoing and other objects, features, and advantages of the disclosed subject matter will become more apparent from the following detailed description with reference to the accompanying drawings. Attached Figure Description

[0006] Figure 1 The illustration shows a block-based processor core that can be used in some examples of the disclosed technologies.

[0007] Figure 2 The illustration shows a block-based processor core that can be used in some examples of the disclosed technologies.

[0008] Figure 3 The illustration shows several instruction blocks based on certain examples of the disclosed technology.

[0009] Figure 4 The illustrations show instruction blocks and portions of source code that can be used in some examples of the disclosed technologies.

[0010] Figure 5 The diagram illustrates a block-based processor header and instructions that may be used in some examples of the disclosed technologies.

[0011] Figure 6 It is a state diagram illustrating the multiple states assigned to an instruction block as it is mapped, executed, and retired.

[0012] Figure 7 The illustration shows multiple instruction blocks and processor cores that can be used in some examples of the disclosed technologies.

[0013] Figure 8 This is a flowchart outlining an example method for comparing and loading a storage identifier with a storage vector, as may be performed in some examples of publicly available technologies.

[0014] Figure 9 The illustrations show example sources and assembly code that can be used in some examples of the disclosed technologies.

[0015] Figure 10 The diagram illustrates an example control flow graph and load-store identifier, which may be used in some examples of the disclosed techniques.

[0016] Figure 11A and Figure 11B An example of generating a masked storage vector, as can be used in some examples of the disclosed techniques, is illustrated.

[0017] Figure 12 This is a flowchart outlining another example method for comparing a loaded storage identifier with a counter, which may be executed in some examples of the disclosed technology.

[0018] Figure 13 It is a control flow graph that includes multiple memory access instructions and load memory identifiers, as may be used in some examples of the disclosed technology.

[0019] Figure 14 This is a flowchart outlining an example method for transforming source code and / or object code into block-based processor executable code, which may be executed in some examples of the disclosed technology, including an indication of the relative ordering of memory access instructions.

[0020] Figure 15 This is a block diagram illustrating a suitable computing environment for implementing some embodiments of the disclosed techniques. Detailed Implementation

[0021] I. Overall considerations

[0022] This disclosure is set forth in the context of representative embodiments that are not intended to be limited in any way.

[0023] As used herein, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” include the plural forms. Furthermore, the term “comprising” means “including.” Moreover, the term “coupled” covers mechanical, electrical, magnetic, optical, and other practical ways of coupling or linking multiple items together, and does not exclude the presence of intermediate elements between the coupled items. Additionally, as used herein, the term “and / or” means any one or more of the phrases.

[0024] The systems, methods, and apparatuses described herein should not be construed as limiting in any way. Rather, this disclosure relates to all novel and non-obvious features and aspects of the various disclosed embodiments, individually and in various combinations and sub-combinations. The disclosed systems, methods, and apparatuses are not limited to any particular aspect or feature or combination thereof, and the disclosed content and methods do not require the presence of any one or more particular advantages or the solution of any problem. Furthermore, any feature or aspect of the disclosed embodiments may be used in various combinations and sub-combinations with each other.

[0025] While some of the methods disclosed are described in a specific order for ease of presentation, it should be understood that this manner of description encompasses rearrangement unless required by the specific language described below. For example, operations described sequentially may be rearranged or performed in parallel in some cases. Furthermore, for simplicity, the accompanying drawings may not show the various ways in which the disclosed content and methods can be combined with other content and methods. Additionally, the description sometimes uses terms such as “generate,” “produce,” “display,” “receive,” “emit,” “verify,” “execute,” and “initiate” to describe the disclosed methods. These terms are high-level descriptions of the actual operations performed. The actual operations corresponding to these terms will vary depending on the specific implementation and are readily discernible to those skilled in the art.

[0026] The operational theories, scientific principles, or other theoretical descriptions presented herein regarding the apparatus or methods of this disclosure have been provided for the purpose of better understanding and are not intended to be limiting in scope. The apparatuses and methods in the appended claims are not limited to those implemented in a manner described by such operational theories.

[0027] Any of the disclosed methods can be implemented as computer-executable instructions stored on one or more computer-readable media (e.g., computer-readable media such as one or more optical media optical discs, volatile memory components such as DRAM or SRAM) or non-volatile memory components such as hard disk drives) and executed on a computer (e.g., any commercially available computer, including smartphones or other mobile devices including computing hardware). Any of the computer-executable instructions used to implement the disclosed techniques, and any data created and used during the implementation of the disclosed embodiments, can be stored on one or more computer-readable media (e.g., computer-readable storage media). The computer-executable instructions can be, for example, part of a dedicated software application or a software application accessed or downloaded via a web browser or other software application (such as a remote computing application). Such software can be executed, for example, on a single local computer (e.g., having a general-purpose and / or block-based processor that executes on any suitable commercially available computer), or in a network environment using one or more networked computers (e.g., via the Internet, wide area network, local area network, client-server network (such as a cloud computing network), or other such networks).

[0028] For clarity, only certain selected aspects of the software-based implementation are described. Other details well-known in the art are omitted. For example, it should be understood that the disclosed techniques are not limited to any particular computer language or program. For example, the disclosed techniques can be implemented in C, C++, JAVA, or any other suitable programming language. Similarly, the disclosed techniques are not limited to any particular computer or hardware type. Certain details of suitable computers and hardware are well-known and do not need to be elaborated in this disclosure.

[0029] Furthermore, any of the software-based embodiments (including, for example, computer-executable instructions for causing a computer to perform any of the disclosed methods) can be uploaded, downloaded, or remotely accessed via suitable communication means. Such suitable communication means include, for example, the Internet, the World Wide Web, intranets, software applications, cable (including fiber optic cables), magnetic communication, electromagnetic communication (including RF, microwave, and infrared communication), electronic communication, or other such communication means.

[0030] II. Introduction to the disclosed technology

[0031] Superscalar out-of-order microarchitectures employ significant circuit resources to rename registers, schedule instructions in data stream order, clean up after misjudgments, and retire results for precise anomalies. This includes expensive circuitry such as deep, multi-port register files, multi-port Content-Accessible Memory (CAM) for data stream instruction scheduling wake-ups, and numerous wide-bus multiplexers and bypass networks—all of which are resource-intensive. For example, FPGA-based implementations of read-many, write-many RAM typically require a mix of copying, multi-loop operations, clock doubling, group interleaving, real-value tables, and other expensive techniques.

[0032] The disclosed techniques can achieve performance enhancements by applying techniques including high instruction-level parallelism (ILP), out-of-order execution (OoO), and superscalar execution, while avoiding significant complexity and overhead in both processor hardware and associated software. In some examples of the disclosed techniques, block-based processors utilize the EDGE ISA, designed for area and energy efficiency with high ILP execution. In other examples, the use of the EDGE architecture and associated compilers cleverly avoids register renaming, CAM, and other complexities.

[0033] In some examples of the disclosed techniques, the EDGE ISA can eliminate the need for one or more complex architectural features, including register renaming, dataflow analysis, false positive recovery, and in-order retirement, while supporting mainstream programming languages ​​such as C and C++. In some examples of the disclosed techniques, block-based processors execute multiple (two or more) instructions as atomic blocks. Block-based instructions can be used to express the semantics of program dataflow and / or instruction flow in a more explicit manner, allowing for improved compiler and processor performance. In some examples of the disclosed techniques, the Explicit Data Graphics Execution Instruction Set Architecture (EDGE ISA) includes information about program control flow that can be used to improve the detection of inappropriate control flow instructions, thereby increasing performance, saving memory resources, and / or saving energy.

[0034] In some examples of the disclosed techniques, instructions organized within an instruction block are atomically fetched, executed, and committed. Instructions within the block are executed in data-flow order, using register renaming to reduce or eliminate and provide power-efficient OoO execution. The compiler can be used to explicitly encode data dependencies via ISA, which reduces or eliminates the burden on processor core control logic to rediscover dependencies at runtime. Using asserted execution, intra-block branches can be translated into data-flow instructions, and dependencies other than memory dependencies can be limited to direct data dependencies. The disclosed object-form encoding technique allows instructions within a block to pass their operands directly via operand buffers, reducing access to power-constrained multi-port physical register files.

[0035] Instructions can communicate between instruction blocks using memory and registers. Therefore, by leveraging a hybrid dataflow execution model, the EDGE architecture can still support imperative programming languages ​​and sequential memory semantics, but also expect to enjoy the benefits of out-of-order execution with near-in-order power efficiency and complexity.

[0036] Apparatus, methods, and computer-readable storage media for generating and using memory access instruction sequence encoding for block-based processors are disclosed. In some examples of the disclosed technology, the instruction block includes an instruction block header and multiple instructions. In other words, the instructions executing the instruction block may or may not affect the state as a whole.

[0037] In some examples of the disclosed technology, the hardware structure stores data indicating the execution order to be adhered to for multiple memory access instructions, including memory load and memory store instructions. A control unit coupled to the processor core controls the issuance of memory access instructions based at least in part on the data stored in the hardware structure. Thus, memory read / write hazards can be avoided, allowing instructions in an instruction block to execute as soon as their dependencies become available. In some examples, the control unit includes wake-up and selection logic used to determine when memory instructions are issued to the load / store queue.

[0038] As will be readily understood by those skilled in the art, multiple implementations of the disclosed technology are possible with various trade-offs in area and performance.

[0039] III. Example of a block-based processor

[0040] Figure 1 Such as the block-based processor 100, which can be implemented in some examples of the disclosed technology. Figure 10 Processor 100 is configured to execute atomic instruction blocks according to an instruction set architecture (ISA). The ISA describes several aspects of processor operation, including a register model, several defined operations executed by block-based instructions, a memory model, interrupts, and other architectural features. The block-based processor includes multiple processor cores 110, including processor core 111.

[0041] As in Figure 1As shown, processor cores are interconnected via core interconnect 120. Core interconnect 120 carries data and controls signals between individual cores in core 110, memory interface 140, and input / output (I / O) interface 145. Core interconnect 120 can send and receive signals using electrical, optical, magnetic, or other suitable communication technologies, and can provide communication connections according to several different topologies depending on a specific desired configuration. For example, core interconnect 120 can have a crossbar switch, bus, point-to-point bus, or other suitable topology. In some examples, any core in core 110 can be connected to any core in other cores, while in other examples, some cores are connected only to a subset of other cores. For example, each core can be connected only to the nearest 4, 8, or 20 neighboring cores. Core interconnect 120 can be used to transfer input / output data to and from cores, and to transfer control signals and other information signals to and from cores. For example, each core 110 in core 110 can receive and transmit a semaphore indicating the execution status of an instruction currently being executed by each core in the corresponding core. In some examples, core interconnect 120 is implemented as a wiring connecting core 110 and the memory system, while in other examples, the core interconnect may include circuitry, switching and / or routing components for multiplexing data signals on one or more interconnect lines, including active signal drivers and repeaters or other suitable circuitry. In some examples of the disclosed technology, signals within and / or to / from processor 100 are not limited to full-swing electrical digital signals, but the processor may be configured to include differential signals, pulse signals, or other suitable signals for transmitting data and control signals.

[0042] exist Figure 1 In some examples, the processor's memory interface 140 includes interface logic for connecting to additional memory (e.g., memory located on another integrated circuit besides processor 100). The external memory system 150 includes an L2 cache 152 and main memory 155. In some examples, the L2 cache may be implemented using static RAM (SRAM), and the main memory 155 may be implemented using dynamic RAM (DRAM). In some examples, the memory system 150 is included on the same integrated circuit as other components of processor 100. In some examples, the memory interface 140 includes blocks that allow the transfer of data in memory without using register files and / or the processor 100's direct memory access (DMA) controller. In some examples, the memory interface manages the allocation of virtual memory, thereby expanding the available main memory 155.

[0043] I / O interface 145 includes circuitry for receiving and sending input and output signals to other components, such as hardware interrupts, system control signals, peripheral interfaces, coprocessor control and / or data signals (e.g., signals for a graphics processing unit, floating-point coprocessor, physical processing unit, digital signal processor, or other coprocessor), clock signals, semaphores, or other suitable I / O signals. I / O signals can be synchronous or asynchronous. In some examples, all or part of the I / O interface is implemented using memory-mapped I / O technology in conjunction with memory interface 140.

[0044] The block-based processor 100 may also include a control unit 160. The control unit 160 oversees the operation of the processor 100. Operations that can be performed by the control unit 160 may include allocating and deallocating code for instruction processing, controlling input and output data between any components in the core, register file, memory interface 140, and / or I / O interface 145, modifying the execution flow, and verifying the target locations of branch instructions, instruction headers, and other changes in the control flow. The control unit 160 can generate and control the processor based on control flow and metadata information representing exit points and control flow probabilities for instruction blocks.

[0045] Control unit 160 can also handle hardware interrupts and control the reading and writing of special system registers (e.g., a program counter stored in one or more register files). In some examples of the disclosed technology, control unit 160 is implemented using at least partially one or more of the processing cores 110, while in other examples, non-block-based processing cores (e.g., general-purpose RISC processing cores coupled to memory) are used to implement control unit 160. In some examples, control unit 160 is implemented using at least partially one or more of the following: hardwired finite state machines, programmable microcode, programmable gate arrays, or other suitable control circuitry. In alternative examples, the control unit functions may be performed by one or more cores 110.

[0046] Control unit 160 includes a scheduler 165 used to allocate instruction blocks to processor core 110. As used herein, scheduler allocation refers to the operation of directing instruction blocks, including initiating instruction block mapping, fetching, decoding, executing, committing, discarding, idling, and refreshing instruction blocks. During instruction block mapping, processor core 110 is assigned to the instruction block. The described stages of instruction operations are for illustrative purposes, and in some examples of the disclosed art, certain operations may be combined, certain operations may be omitted, certain operations may be separated into multiple operations, or additional operations may be added. Scheduler 165 schedules the flow of instructions, including the allocation and deallocation of cores for executing instruction processing, and the control of input and output data between any components in the core, register file, memory interface 140, and / or I / O interface 145. Control unit 160 may also be used to store memory access instruction hardware structure 167, including data in memory mask and memory vector registers, as discussed further in detail below.

[0047] The block-based processor 100 also includes a clock generator 170 that distributes one or more clock signals to various components within the processor (e.g., core 110, interconnect 120, memory interface 140, and I / O interface 145). In some examples of the disclosed technology, all components share a common clock, while in other examples, different components use different clocks (e.g., clock signals with different clock frequencies). In some examples, a portion of the clock is gated to allow power savings when some components in the processor are not in use. In some examples, the clock signal is generated using a phase-locked loop (PLL) to generate a signal with a fixed constant frequency and duty cycle. Circuitry receiving the clock signal may be triggered on a single edge (e.g., a rising edge), while in other examples, at least some of the receiving circuitry is triggered by both rising and falling clock edges. In some examples, the clock signal may be transmitted optically or wirelessly.

[0048] IV. Example of a block-based processor core

[0049] Figure 2 This is a block diagram further describing an example microarchitecture for a block-based processor 100, as may be used in some examples of the disclosed technology, and specifically an instance of a core in a block-based processor core. For ease of illustration, five stages are used to illustrate an exemplary block-based processor core: instruction fetch (IF), decode (DC), operand fetch, execution (EX), and memory / data access (LS). However, those skilled in the art will readily understand that the illustrated microarchitecture (e.g., adding / removing stages, adding / removing units for execution operations, and other implementation details) can be modified to suit specific applications for block-based processors.

[0050] like Figure 2 As shown, processor core 111 includes a control unit 205, which generates control signals for regulating core operations and uses an instruction scheduler 206 to schedule the instruction flow within the core. Operations that can be performed by the control unit 205 and / or the instruction scheduler 206 may include generating and using memory access instruction codes, allocating and deallocating cores for instruction processing, and controlling input and output data between any components in the core, register file, memory interface 140, and / or I / O interface 145. The control unit may also control load-store queues, schedulers, global control units, other units, or combinations thereof used to determine the rate and order of instruction issuance.

[0051] In some examples, the instruction scheduler 206 is implemented using a general-purpose processor coupled to memory, which is configured to store data for scheduling instruction blocks. In some examples, the instruction scheduler 206 is implemented using a dedicated processor or a block-based processor core coupled to memory. In some examples, the instruction scheduler 206 is implemented as a finite state machine coupled to memory. In some examples, an operating system executing on a processor (e.g., a general-purpose processor or a block-based processor core) generates priority, prediction, and other data that can be at least partially used to schedule instruction blocks using the instruction scheduler 206. As will be readily apparent to those skilled in the art, other circuit structures implemented in integrated circuits, programmable logic, or other suitable logic can be used to implement the hardware for the instruction scheduler 206.

[0052] Control unit 205 also includes memory (e.g., in SRAM or registers) for storing control flow information and metadata. For example, data for the memory access instruction sequence may be stored in a hardware structure (e.g., memory instruction data repository 207). Memory instruction data repository 207 may store data for a memory mask (e.g., by copying data encoded in an instruction block or generated by an instruction decoder during instruction decoding), memory vector registers (e.g., storing data indicating which and what types of memory access instructions have been executed), and masked memory vector register data (e.g., data generated by applying a memory mask to the memory vector registers). In some examples, memory instruction data repository 200 includes a counter that tracks the number and type of memory access instructions that have been executed.

[0053] Control unit 205 can also handle hardware interrupts and control the reading and writing of a program counter in special system registers (e.g., stored in one or more register files). In other examples of the disclosed technology, control unit 205 and / or instruction scheduler 206 are implemented using a non-block-based processing core (e.g., a general-purpose RISC processing core coupled to memory). In some examples, control unit 205 and / or instruction scheduler 206 are implemented using at least one or more of the following: hardwired finite state machine, programmable microcode, programmable gate array, or other suitable control circuitry.

[0054] An exemplary processor core 111 includes two instruction windows 210 and 211, each of which can be configured to execute an instruction block. In some examples of the disclosed technology, an instruction block is a block-based atomic collection of processor instructions, which includes an instruction block header and one or more instructions. As will be discussed further below, the instruction block header includes information that can be used to further define the semantics of one or more instructions within the instruction block. Depending on the specific ISA and processor hardware used, the instruction block header can be used during instruction execution to improve the performance of executing the instruction block, for example, by allowing early instruction and / or data fetching, improved branch prediction, speculative execution, improved energy efficiency, and improved code compactness. In other examples, different numbers of instruction windows are possible, such as one, four, eight, or other numbers.

[0055] Each instruction window in instruction windows 210 and 211 may receive instructions and data from one or more input ports 220, 221, and 222 connected to the interconnect bus and instruction cache 227, which in turn are connected to instruction decoders 228 and 229. Additional control signals may also be received on additional input port 225. Each instruction decoder in instruction decoders 228 and 229 decodes the instruction header and / or instructions for the instruction block and stores the decoded instructions in memory repositories 215 and 216 located in each corresponding instruction window 210 and 211. Additionally, each decoder in decoders 228 and 229 may send data to control unit 205 to configure the operation of processor core 111, for example, according to execution flags specified in the instruction block header or in the instructions.

[0056] Processor core 111 also includes a register file 230 coupled to L1 (Level 1) cache 235. Register file 230 stores data for registers defined in a block-based processor architecture and may have one or more read ports and one or more write ports. For example, a register file may include two or more write ports for storing data in the register file and multiple read ports for reading data from individual registers within the register file. In some examples, a single instruction window (e.g., instruction window 210) may access only one port of the register file at a time, while in other examples, instruction window 210 may access one read port and one write port simultaneously, or two or more read ports and / or write ports simultaneously. In some examples, register file 230 may include 64 registers, each holding a word of 32 bits of data. (For ease of illustration, 32 bits of data will be referred to as a word unless otherwise specified. Suitable processors according to the disclosed technology can operate with words of 8, 16, 64, 128, 256 bits, or another number of bits.) In some examples, some registers within register file 230 may be allocated for special purposes. For example, some registers in the register file may be dedicated as system registers. Examples of system registers include registers storing constant values ​​(e.g., all-zero words), a program counter (PC) indicating the current address of the executing program thread, the number of physical cores, the number of logical cores, core assignment topology, core control flags, execution flags, processor topology, or other suitable dedicated purposes. In some examples, there are multiple program counter registers, one or each program counter, to allow multiple execution threads to be executed in parallel across one or more processor cores and / or processors. In some examples, the program counter is implemented as a specified memory location rather than as a register in a register file. In some examples, the use of system registers may be restricted by the operating system or other supervisory counter instructions. In some examples, register file 230 is implemented as an array of flip-flops, while in other examples, the register file may be implemented using latches, SRAM, or other memory storage. The ISA specification for a given processor (e.g., processor 100) specifies how registers within register file 230 are defined and used.

[0057] In some examples, processor 100 includes a global register file shared by multiple processor cores. In some examples, individual register files associated with processor cores may be combined, either statically or dynamically, to form a larger file, depending on the processor ISA and configuration.

[0058] like Figure 2As shown, the memory repository 215 of the instruction window 210 includes multiple decoded instructions 241, a left operand (LOP) buffer 242, a right operand (ROP) buffer 243, an assertion buffer 244, three broadcast channels 245, and an instruction scoring board 247. In some examples of the disclosed technology, such as... Figure 2 The diagram shows the breakdown of each instruction in an instruction block into a line of decoded instructions, left and right operands, and scoreboard data. The decoded instruction 241 may include a partially or fully decoded version of an instruction stored as bit-level control signals. Operand buffers 242 and 243 store operands (e.g., register values ​​received from register file 230, data received from memory, immediate operands encoded within instructions, operands calculated from earlier issued instructions, or other operand values) until their corresponding decoded instructions are ready for execution. Instruction operands and assertions are read from operand buffers 242 and 243 and assertion buffer 244, respectively, instead of from the register file. The instruction scoreboard 245 may include buffers for assertions related to instructions, including wired OR logic for combining assertions sent to instructions from multiple instructions.

[0059] The memory repository 216 of the second instruction window 211 stores similar instruction information (decoded instructions, operands, and scoreboards) in memory repository 215, but for simplicity, it is not included in memory repository 215. Figure 2 The instruction blocks are shown in the diagram. The instruction blocks can be executed in parallel or sequentially with respect to the first instruction window, subject to the ISA, and by the second instruction window 211 as directed by the control unit 205.

[0060] In some examples of the disclosed technology, the front-end pipeline stages IF and DC can operate decoupled from the back-end pipeline stages (IS, EX, LS). The control unit can fetch two instructions per clock cycle and decode them into each instruction window in instruction windows 210 and 211. Control unit 205 provides instruction window dataflow scheduling logic for monitoring the readiness status of the inputs (e.g., assertions and operands for each corresponding instruction) of each decoded instruction using scoring board 245. An instruction is ready to be issued when all input operands and assertions for a particular decoded instruction are ready. Control unit 205 then initiates control signals per cycle to execute (issue) one or more subsequent instructions (e.g., the lowest-numbered ready instruction) and based on the decoded instruction, and sends the instruction's input operands to one or more functional units 260 for execution. The decoded instruction can also be encoded with multiple readiness events. The scheduler in control unit 205 receives these and / or other events from other sources and updates the readiness status of other instructions in the window. Therefore, execution continues from the ready zero-input instruction of processor core 111, the instruction that is the target of the zero-input instruction, and so on.

[0061] The decoded instructions 241 do not need to be executed in the same order they are arranged in the memory repository 215 of the instruction window 210. Instead, the instruction scoring board 245 is used to track the dependencies of the decoded instructions and, when dependencies are satisfied, schedules the associated individual decoded instructions for execution. For example, a reference to a corresponding instruction can be pushed onto the ready queue when the dependency for that instruction is satisfied, and ready instructions can be scheduled from the ready queue in a first-in, first-out (FIFO) order. For instructions encoded using load memory identifiers (LSIDs), the instruction order will also follow the priority enumerated in the instruction LSIDs or executed in an order that behaves as if the instructions were executed in a specified order.

[0062] The information stored in the scoring board 245 may include, but is not limited to, execution assertions of associated instructions (e.g., whether an instruction is waiting to compute an assertion bit and whether the instruction will execute if the assertion bit is true or false), operand availability for instructions, or other prerequisites required before issuing and executing associated individual instructions. The number of instructions stored in each instruction window generally corresponds to the number of instructions within an instruction block. In some examples, operands and / or assertions are received on one or more broadcast channels that allow the same operands or assertions to be sent to a larger number of instructions. In some examples, the number of instructions within an instruction block may be 32, 64, 128, 1024, or another number of instructions. In some examples of the disclosed technology, instruction blocks are allocated across multiple instruction windows within the processor core. Out-of-order operations and memory accesses can be controlled according to data specifying one or more operating modes.

[0063] In some examples, restrictions are imposed on the processor (e.g., based on architecture definitions or the processor's programmable configuration) to prevent the execution of instructions from proceeding in a non-ordered manner within the instruction block. In some examples, the lowest-numbered available instruction is configured as the next instruction to be executed. In some examples, control logic traverses the instructions within the instruction block and executes the next instruction ready for execution. In some examples, only one instruction can be issued and / or executed at a time. In some examples, instructions within an instruction block are issued and executed in a deterministic order (e.g., the order in which instructions are arranged within the block). In some examples, restrictions on instruction ordering can be configured when using a software debugger or by the user debugging a program executing on a block-based processor.

[0064] Instructions can be allocated and scheduled using a control unit 205 located within processor core 111. Control unit 205 orchestrates instruction fetching from memory, instruction decoding, execution of instructions once they have been loaded into their respective instruction windows, data inflow / outflow from processor core 111, and control signals input and output by the processor core. For example, control unit 205 may include a ready queue as described above for use in scheduling instructions. Instructions stored in memory repositories 215 and 216 located in each corresponding instruction window 210 and 211 can be executed atomically. Therefore, updates to the visible architectural state (e.g., register file 230 and memory) affected by the executed instructions can be locally buffered within core 200 until the instruction is committed. Control unit 205 can determine when instructions are ready to be committed, order the commit logic, and issue a commit signal. For example, the commit phase for an instruction block can begin when all register writes are buffered, all writes to memory are buffered, and the compute branch target is computed. An instruction block can be committed when updates to the visible architectural state are complete. For example, instruction blocks can be submitted when writing registers to the register file, sending repositories to the load / store unit or memory controller, and generating commit signals. Control unit 205 also at least partially controls the allocation of functional unit 260 to each instruction window within the corresponding instruction window.

[0065] like Figure 2As shown, a first router 250, having multiple execution pipeline registers 255, is used to send data from either instruction window in instruction windows 210 and 211 to one or more functional units in functional units 260. These functional units may include, but are not limited to, integer ALUs (arithmetic logic units) (e.g., integer ALUs 264 and 265), floating-point units (e.g., floating-point ALU 267), shift / rotation logic (e.g., barrel shifter 268), or other suitable execution units that may include graphics functions, physics functions, and other mathematical operations. The first router 250 also includes wake-up / selection logic 258 used to determine when to send a memory instruction to the load / store queue 275. For example, the wake-up / selection logic 258 may determine whether all source operands and assertion conditional statements are available for a memory access instruction and, based on this determination, send an address (and, if applicable, data) to the load / store queue 275.

[0066] Data from functional unit 260 may then be routed via second router 270 to outputs 290, 291, and 292, depending on the requirements of the specific instruction being executed, or routed back to operand buffers (e.g., LOP buffer 242 and / or ROP buffer 243) or fed back to another functional unit. Second router 270 includes a load / store queue 275 that can be used to issue memory instructions, a data cache 277 that stores data input to or output from the core to memory, and a load / store pipeline register 278.

[0067] Load / store queue 275 receives and temporarily stores information for executing memory access instructions. An instruction block can execute all memory access instructions as a single atomic transaction block. In other words, all or none of the memory access instructions may be executed. The relative order of memory access instructions is determined based on the LSID associated with each memory access instruction (e.g., an LSID encoded using the corresponding instruction) and, in some cases, a memory mask. In some examples, additional performance can be gained by executing memory accesses out of relative order specified by LSID, but the memory state must still behave as if the instructions were executed sequentially. Load / store queue 275 also receives addresses for loading instructions and addresses and data for storing instructions. In some examples, the load / store queue waits to execute queued memory access instructions until it is determined that the contained instruction block will actually be committed. In other examples, load / store queue 275 may speculatively issue at least some memory access instructions, but will need to flush memory operations if the block is not committed. In other examples, control unit 205 determines the order in which memory access instructions are executed by providing functions described as being performed by wake-up / selection logic and / or load / store queue 275. In some examples, processor 100 includes a debug mode that allows memory access instructions to be issued incrementally using a debugger. Load / store queue 275 can be implemented using control logic (e.g., utilizing a finite state machine) and memory (e.g., registers or SRAM) to perform memory transactions and store memory instruction operands, respectively.

[0068] The core also includes a control output 295, which indicates when the execution of all instructions in one or more instruction windows, such as instruction windows 210 or 211, has been completed. When the execution of an instruction block is complete, the instruction block is designated as "committed," and the signal from the control output 295 can then be used by other cores within the block-based processor 100 and / or by the control unit 160 to initiate the scheduling, fetching, and execution of other instruction blocks. Both the first router 250 and the second router 270 can send data back to the instructions (e.g., as operands for other instructions within the instruction block).

[0069] As will be readily understood by those skilled in the art, the components within the individual core 200 are not limited to... Figure 2 The components shown are not necessarily those of a specific application; rather, they can vary depending on the requirements of that application. For example, a core may have fewer or more instruction windows, a single instruction decoder may be shared by two or more instruction windows, and the number and type of functional units used may vary depending on the specific target application for the block-based processor. Other considerations for applications when selecting and allocating resources using instruction cores include performance requirements, energy usage requirements, integrated circuit chips, processing technology, and / or cost.

[0070] It will be readily apparent to those skilled in the art that trade-offs in processor performance can be made through the design and allocation of resources within the instruction window (e.g., instruction window 210) of processor core 110 and control unit 205. Area, clock cycles, capabilities, and limitations substantially determine the implementation performance of individual core 110 and the throughput of block-based processor 100.

[0071] Instruction scheduler 206 can have different functionalities. In some high-performance examples, the instruction scheduler is highly concurrent. For example, each cycle(s) of the decoder(s) writes the decode readiness status and decoded instruction to one or more instruction windows, selects the next instruction to issue, and sends a second readiness event as a response backend—either a target readiness event targeting the input slot (assertion, left operand, right operand, etc.) of a specific instruction or a broadcast readiness event targeting all instructions. Each instruction readiness status bit, along with the decode readiness status, can be used to determine if an instruction is ready to be issued.

[0072] In some cases, scheduler 206 accepts events for target instructions that have not yet been decoded and must also prevent the re-issuance of issued ready instructions. In some examples, the instruction can be unassertified or asserted (based on true or false conditions). An asserted instruction becomes ready only when the assertion result of another instruction targets it and that result matches the assertion condition. If the associated assertion condition does not match, the instruction is never issued. In some examples, asserted instructions can be speculatively issued and executed. In some examples, the processor can subsequently check that speculatively issued and executed instructions were correctly speculatively ...

[0073] When branching to a new instruction block, the corresponding instruction window's ready state is cleared (block reset). However, when the instruction block branch returns to itself (block refresh), only the active ready state is cleared. Therefore, the ready state for decoding the instruction block can be preserved, eliminating the need to re-fetch and re-decode the block's instructions. Thus, block refresh can be used to save time and energy within loops.

[0074] V. Example stream of instruction blocks

[0075] Turn now Figure 3Figure 300 illustrates portion 310 of a block-based instruction stream, comprising multiple variable-length instruction blocks 313-314. The instruction stream can be used to implement user applications, system services, or any other suitable use. The instruction stream can be stored in memory, received from another process in memory, received via a network connection, or stored or received in any other suitable manner. Figure 3 In the examples shown, each instruction block begins with an instruction header followed by a variable number of instructions. For example, instruction block 311 includes a header 320 and twenty instructions 321. The specific instruction header 320 shown includes multiple data fields that partially control the execution of the instructions within the instruction block and also allow for performance enhancement techniques, such as branch prediction, speculative execution, lazy evaluation, and / or other techniques. The instruction header 320 also includes an indication of the instruction block size. The instruction block size can be in a block larger than one instruction, for example, an instruction block containing a block of four instructions. In other words, the block size is shifted by 4 bits to compress the allocated header space to the specified instruction block size. Thus, a size value of 0 indicates the smallest instruction block size, which is a block header followed by four instructions. In some examples, the instruction block size is expressed as bytes, words, n-word blocks, addresses, address offsets, or other appropriate expressions used to describe the instruction block size. In some examples, the instruction block size is indicated by the termination bit pattern in the instruction block header and / or footer.

[0076] The instruction block header 320 may also include one or more execution flags indicating one or more operating modes for executing the instruction block. For example, operating modes may include kernel fusion operation, vector mode operation, memory dependency prediction, and / or sequential or deterministic instruction execution.

[0077] In some examples of the disclosed technology, instruction header 320 includes one or more identifier bits indicating that the encoded data is an instruction header. For example, in some block-based processor ISAs, a single ID bit in the least significant bit space is always set to a binary value of 1 to indicate the start of a valid instruction block. In other examples, different bit codes may be used for the identifier bits. In some examples, instruction header 320 includes information indicating that the associated instruction block is specific to a particular version of the ISA it is encoded in.

[0078] The block instruction header may also include multiple block exit types for use in, for example, branch prediction, control flow determination, and / or branch processing. Exit types can indicate what type of branch instruction it is, such as: a sequential branch instruction pointing to the next consecutive instruction block in memory; an offset instruction branching to another instruction block at a memory address calculated relative to the offset; a subroutine call; or a subroutine return. By encoding the branch exit types in the instruction header, the branch predictor can begin operation at least partially before the branch instructions within the same instruction block have been fetched and / or decoded.

[0079] The instruction block header 320 also includes a memory mask that indicates which load-store queue identifiers (LSIDs) encoded in the block instructions are assigned to memory operations. For example, for a block with eight memory access instructions, the memory mask 01011011 would indicate the presence of three memory store instructions (bits 0 corresponding to LSIDs 0, 2, and 5) and five memory load instructions (bits 1 corresponding to LSIDs 1, 3, 4, 6, and 7). The instruction block header may also include a write mask that identifies which global register(s) the associated instruction block will be written to. In some examples, the memory mask is stored in a memory vector register, for example, by an instruction decoder (e.g., decoder 228 or 229). In other examples, the instruction block header 320 does not include a memory mask, but the instruction decoder dynamically generates the memory mask by analyzing instruction dependencies as the instruction block is decoded. For example, the decoder may analyze the load-store identifiers of the instruction block instructions to determine the memory mask and store it in the memory vector register. Similarly, in other examples, write masks are not encoded in the instruction block header but are dynamically generated (e.g., by analyzing registers referenced by instructions within the block using an instruction decoder) and stored in a write mask register. The storage mask and write mask can be used to determine when the execution of the instruction block has completed and thus initiate a commit of the instruction block. The associated register file must receive writes to each entry before the instruction block can complete. In some examples, block-based processor architectures may include not only scalar instructions but also Single Instruction Multiple Data (SIMD) instructions that allow operations with a larger number of data operands within a single instruction.

[0080] Examples of suitable block-based instructions that can be used for instruction 321 may include instructions for performing integer and floating-point arithmetic, logical operations, type conversions, register reads and writes, memory loads and stores, branch and jump execution, and other suitable processor instructions. In some examples, the instructions include instructions for configuring the processor to operate, for example, by speculative execution based on control flow and data regarding memory access instructions stored in a hardware structure (such as the memory instruction data repository 207), according to one or more operations. In some examples, the memory instruction data repository 207 is not architecturally visible. In some examples, access to the memory instruction data repository 207 is configured to be limited to processor operations in the processor's supervised mode or other protected mode.

[0081] VI. Example block instruction target encoding

[0082] Figure 4 Figure 400 is an example of two sections 410 and 415 of C language source code and their corresponding instruction blocks 420 and 425, illustrating how block-based instructions can be explicitly encoded to their targets. In this example, the first two READ instructions 430 and 431 target the right (T[2R]) operand and left (T[2L]) operand of the ADD instruction 432, respectively (2R indicates targeting the right operand of instruction number 2; 2L targets the left operand of instruction number 2). In the illustrated ISA, a read instruction is an instruction that reads from a global register file (e.g., register file 230); however, any instruction can target a global register file. When the ADD instruction 432 receives the result of two register reads, it becomes ready and executes. Note that in this disclosure, the right operand is sometimes referred to as OP0 and the left operand as OP1.

[0083] When the TLEI (less than equal immediate test) instruction 433 receives its single input operand from ADD, it becomes ready to be issued and executed. The test then produces assertion operands broadcast on channel one (B[1P]) to all instructions listening for assertions on the broadcast channel; in this example, these instructions are branch instructions of two assertions (BRO_T 434 and BRO_F 435). Receiving a branch instruction that matches the assertion will trigger (execute), but the other instruction encoded using the assertion's complement will not trigger / execute.

[0084] The dependency graph 440 for instruction block 420 is also illustrated as instruction node array 450 and its corresponding operand targets 455 and 456. This illustrates the correspondence between block instruction 420, the corresponding instruction window entry, and the lower-level data flow graph represented by the instruction. Here, the decoded instructions READ 430 and READ 431 are ready to be issued because they have no input dependencies. As they are issued and executed, values ​​read from registers R0 and R7 are written to the right and left operand buffers of ADD 432, thereby marking the left and right operands of ADD 432 as "ready". As a result, the ADD 432 instruction becomes ready, is issued to the ALU, is executed, and a sum is written to the left operand of the TLEI instruction 433.

[0085] VII. Example of a block-based instruction format

[0086] Figure 5 This is a diagram illustrating a generalized example of the instruction format used for instruction header 510, general instructions 520, branch instructions 530, and memory access instructions 540 (e.g., memory load or store instructions). The instruction format can be used for a block of instructions executed according to multiple execution flags specified in the instruction header in a specified operating mode. The instruction header or each instruction within an instruction is labeled according to the number of bits. For example, instruction header 510 comprises four 32-bit words and is labeled from its least significant bit (lsb) (bit 0) to its most significant bit (msb) (bit 127). As shown, the instruction header includes a write mask field, a store mask field 515, multiple exit type fields, multiple execution flag fields, an instruction block size field, and an instruction header ID bit (the least significant bit of the instruction header). In some examples, the store mask field 515 is replaced or supplemented by an LSID count 517, which indicates the number of store instructions on each assertion path of the instruction block. For instruction blocks with different numbers of store instructions on different assertion paths, one or more instructions can be invalidated, and the count of executed store instructions can be incremented, so that each assertion path will indicate that the same number of store instructions have been executed at runtime. In some examples, header 510 does not indicate the LSID count or store mask, but the instruction decoder dynamically generates information based on the LSIDs encoded in individual store instructions.

[0087] exist Figure 5 The execution flag field depicted occupies bits 6 to 13 of the instruction block header 510 and indicates one or more operating modes for executing the instruction block. For example, operating modes may include kernel fusion operation, vector mode operation, branch predictor disabled, memory dependency predictor disabled, block synchronization, interrupt after block, interrupt before block, block failure, and / or sequential or deterministic instruction execution.

[0088] The exit type field includes data that can be used to indicate the type of control flow instruction encoded within an instruction block. For example, the exit type field may indicate that the instruction block includes one or more of the following: sequential branch instructions, offset branch instructions, indirect branch instructions, call instructions, and / or return instructions. In some examples, a branch instruction can be any control flow instruction used to transfer control flow (including relative and / or absolute addresses and the use of conditional or unconditional assertions) between instruction blocks. In addition to identifying implicit control flow instructions, the exit type field can also be used for branch prediction and speculative execution.

[0089] The illustrated general block instruction 520 is stored as a 32-bit word and includes an opcode field, an assertion field, a broadcast ID field (BID), a vector operation field (V), a single instruction multiple data (SIMD) field, a first destination field (T1), and a second destination field (T2). For instructions with more consumers than the destination field, the compiler can use move instructions to build a fan-out tree, or it can assign high fan-out instructions to the broadcast. Broadcasting supports sending operands to any number of consumer instructions in the core via a lightweight network.

[0090] Although the general instruction format outlined by General Instruction 520 can represent some or all instructions processed by a block-based processor, those skilled in the art will readily understand that one or more instruction fields may deviate from the general format used for a particular instruction, even for a specific example of the ISA. The opcode field specifies the operation performed by instruction 520, such as memory read / write, register load / store, add, subtract, multiply, divide, shift, rotate, system operation, or other appropriate instruction. The assertion field specifies the condition under which the instruction will be executed. For example, the assertion field may specify the value "true," and the instruction will only be executed if the corresponding condition flag matches the specified assertion value. In some embodiments, the assertion field at least partially specifies which is used for comparison assertion, while in other examples, assertion is performed against flags set by previous instructions (e.g., the preceding instruction in an instruction block). In some examples, the assertion field may specify that the instruction will always or never be executed. Therefore, the use of assertion fields can allow for denser object code, improved energy efficiency, and improved processor performance by reducing the number of branch instructions that are decoded and executed.

[0091] The target fields T1 and T2 specify the instruction to which the result of a block-based instruction is sent. For example, the ADD instruction at instruction slot 5 can specify that its computation result will be sent to instructions at slots 3 and 10, including the specification of operand slots (e.g., left operand, right operand, or assertion operand). Depending on the specific instruction and ISA, one or both of the illustrated target fields can be replaced by other information; for example, the first target field T1 can be replaced by immediate operand, additional opcode, specifying two targets, etc.

[0092] Branch instruction 530 includes an opcode field, an assertion field, a broadcast ID (BID) field, and an offset field. The opcode and assertion fields are similar in format and function to those described for general instructions. Offsets can be expressed in groups of four instructions, thus extending the range of memory addresses on which the branch can be executed. Assertions illustrated using general instruction 520 and branch instruction 530 can be used to avoid additional branch jumps within an instruction block. For example, the execution of a specific instruction can be asserted on the result of a previous instruction (e.g., comparing two operands). If the assertion is false, the instruction will not commit the value calculated by the specific instruction. If the assertion value does not match the required assertion, the instruction is not issued. For example, the BRO_F (false assertion) instruction will be issued if it receives a false assertion value.

[0093] It should be readily understood that, as used herein, the term "branch instruction" is not limited to changing program execution to a relative memory location, but includes jumping to absolute or symbolic memory locations, subroutine calls and returns, and other instructions that can modify the flow of execution. In some examples, the flow of execution is modified by changing the value of a system register (e.g., the program counter PC or instruction pointer), while in other examples, it is modified by changing a value stored at a specified location in memory. In some examples, jump register branch instructions are used to jump to a memory location stored in a register. In some examples, jump and link instructions, as well as jump register instructions, are used to implement subroutine calls and returns.

[0094] The memory access instruction 540 format includes an opcode field, an assertion field, a broadcast ID field (BID), a load-store ID field (LSID), an immediate field (IMM), an offset field, and a destination field. The opcode, broadcast, and assertion fields are similar in format and function to those described for general instructions. For example, the execution of a specific instruction can be asserted on the result of a previous instruction (e.g., comparing two operands). If the assertion is false, the instruction will not commit the value calculated by the specific instruction. If the assertion value does not match the required assertion, the instruction is not issued. The immediate field (e.g., and the number of bits shifted) can be used as an offset for the operand sent to the load or store instruction. The operand plus (shifted) immediate offset is used as the memory address for the load / store instruction (e.g., the address for reading data from or storing data in memory). The LSID field specifies the relative order of load and store instructions within a block. In other words, a higher LSID indicates that the instruction should be executed after a lower LSID. In some examples, the processor can determine that two load / store instructions do not conflict (e.g., based on the read / write address used for the instruction) and can execute the instructions in a different order, although the resulting state of the machine should not differ from if the instructions had been executed in the specified LSID order. In some examples, load / store instructions with mutual exclusion assertion values ​​can use the same LSID value. For example, if the first load / store instruction is asserted as true for the value p and the second load / store instruction is asserted as false for the value p, then each instruction can have the same LSID value.

[0095] VIII. Example processor state diagram

[0096] Figure 6 This is a state diagram 600 illustrating the multiple states assigned to an instruction block as it is mapped, executed, and retired. For example, one or more states can be assigned during the execution of an instruction based on one or more execution flags. It should be readily understood that... Figure 6 The states shown are for one example of the disclosed technology, but in other examples, the instruction block may have additional or fewer states, and may have states different from those depicted in state diagram 600. At state 605, the instruction block is unmapped. The instruction block may reside in memory coupled to the block-based processor, be stored on a computer-readable storage device (such as a hard drive or flash drive), and be accessible locally on the processor or on a remote server using a computer network. The unmapped instructions may also reside at least partially in cache memory coupled to the block-based processor.

[0097] At instruction block mapping state 610, control logic for block-based processors (such as an instruction scheduler) can be used to monitor the processing core resources of the block-based processor and map instruction blocks to one or more processing cores.

[0098] The control unit can map one or more instruction blocks to a processor core and / or the instruction window of a specific processor core. In some examples, the control unit monitors a processor core that has previously executed a specific instruction block and can reuse decoded instructions for instruction blocks still residing on a "warm-up" processor core. Once one or more instruction blocks have been mapped to a processor core, the instruction block can proceed to fetch state 620.

[0099] When an instruction block is in fetch state 620 (e.g., instruction fetch), the mapped processor core fetches computer-readable block instructions from the block-based processor's memory system and loads them into the memory associated with the specific processor core. For example, the instructions for fetching the instruction block can be fetched and stored in an instruction cache within the processor core. Instructions can be communicated to the processor core using core interconnects. Once at least one instruction of the instruction block has been fetched, the instruction block can enter instruction decode state 630.

[0100] During instruction decoding state 630, the various bits of the fetched instruction are decoded into signals that can be used by the processor core to control the execution of specific instructions. For example, this can be done in the above... Figure 2 The memory repository 215 or 216 shown stores the decoded instructions. Decoding includes the dependencies for generating the instructions for decoding, operand information for the instructions for decoding, and the target of the instructions for decoding. Once at least one instruction of the instruction block has been decoded, the instruction block can proceed to execution state 640.

[0101] During execution state 640, for example, using the above regarding... Figure 2The functional unit 260 under discussion performs operations associated with instructions. As discussed above, the functions performed may include arithmetic functions, logical functions, branch instructions, memory operations, and register operations. Control logic associated with the processor core monitors the execution of instruction blocks, and once it is determined that the instruction block can be committed or will be aborted, the instruction block state is set to Commit / Abort 650. In some examples, the control logic uses write masks and / or store masks for the instruction block to determine whether execution has been sufficiently performed to commit the instruction block. Executed memory access instructions send data and address information to a load / store queue for accessing memory. In some examples, some memory access instructions (e.g., memory load instructions) may be executed before block execution, while other instructions (e.g., memory store instructions) wait to be executed until the block is being committed. In some examples, all memory access instructions wait to access memory until the block is being committed. In some examples, memory load and store instructions access memory during the execution of the instruction block, but additional hardware catches dangerous memory conditions (e.g., read danger after write) to ensure that main memory behaves as if instructions were executed according to their relative order.

[0102] At commit / abandon state 650, the processor core control unit determines that the operations executed by the instruction block can be completed. For example, memory load operations, register read / write operations, branch instructions, and other instructions will be executed deterministically according to the control flow of the instruction block. Alternatively, if the instruction block is to be abandoned, for example because one or more dependencies of the instruction are not satisfied, or because an instruction is executed speculatively for an unsatisfied assertion used for the instruction block, the instruction block is abandoned so that it does not affect the state of the instruction sequence in memory or register files. Any incomplete memory access operations are also completed. Regardless of whether the instruction block has been committed or abandoned, the instruction block proceeds to state 660 to determine whether the instruction block should be refreshed. If the instruction block is refreshed, the processor core typically re-executes the instruction block using the new data values ​​(specifically, the registers and memory updated by the execution of the block just committed) and proceeds directly to execution state 640. Thus, the time and energy spent on mapping, fetching, and decoding the instruction block can be avoided. Alternatively, if the instruction block is not to be refreshed, the instruction block enters idle state 670.

[0103] In idle state 670, the processor core can be kept idle, for example, by reducing the power of the hardware within the processor core executing the instruction block, thereby maintaining at least a portion of the instructions used to decode the instruction block. At some point, the control unit determines 680 whether to flush the idle instruction block on the processor core. If the idle instruction block is to be flushed, the instruction block can resume execution at execution state 640. Alternatively, if the instruction block is not to be flushed, the instruction block is unmapped and can be flushed to the processor core, and the instruction block can subsequently be mapped to the flushed processor core.

[0104] Although state diagram 600 illustrates the state of an instruction block as being executed on a single processor core for ease of illustration, those skilled in the art will readily understand that in some examples, multiple processor cores may be used to execute multiple instances of a given instruction block in parallel.

[0105] IX. Example of block-based processor and memory configuration

[0106] Figure 7 Figure 700 illustrates an apparatus including a block-based processor 710, which includes a control unit 720 configured to execute instruction blocks based on data for one or more operating modes. The control unit 720 includes a core scheduler 725 and a memory access hardware structure 730. The core scheduler 725 schedules the instruction flow, including allocating and deallocating cores for instruction processing, and controlling input and output data between any components in the core, register file, memory interface, and / or I / O interface. The memory access hardware structure 730 stores data, including, for example, memory mask data, memory vector register data indicating which instructions have been executed, masked memory vector data, and / or control flow data. The memory access hardware structure 730 can be implemented using any suitable technology, including SRAM, registers (e.g., including flip-flops or latch arrays), or other suitable memory technologies. The memory mask can be generated during instruction decoding by the control unit 720. In some examples, the memory mask is read from memory 750 (memory mask 735), from the instruction block header (e.g., memory masks 737 and 738), or from a computer-readable storage medium (e.g., storage medium disk 736).

[0107] The block-based processor 710 also includes one or more processor cores 740-747 configured to fetch and execute instruction blocks. The illustrated block-based processor 710 has up to eight cores, but in other examples there may be 64, 512, 1024, or other numbers of block-based processor cores. The block-based processor 710 is coupled to a memory 750 comprising multiple instruction blocks (including instruction blocks A and B) and a memory mask 736 stored on a computer-readable storage medium disk 755.

[0108] X. Example method for using a storage mask to issue instructions

[0109] Figure 8 This is a flowchart 800 outlining an example method, such as using a storage mask to determine when instructions can be issued and executed using a block-based processor, which may be executed in some examples of the disclosed technology. For example, Figure 1 Block-based processor 100 (including the above) Figure 2 The block-based processor core 111 described herein can be used to execute the methods outlined. In some examples, the execution unit of the block-based processor is configured to execute memory access instructions in an instruction block, and the hardware structure stores data indicating the execution order of at least some of the memory access instructions, and the control unit of the block-based processor is configured to control the issuance of memory access instructions to the execution unit based at least in part on the hardware structure data.

[0110] At process block 810, a memory mask is generated for the block of instructions currently being executed. The memory mask includes data indicating which of a plurality of memory access instructions are store instructions. For example, zeros may be stored at the bits corresponding to store instructions having a specific load memory identifier (LSID) associated with a memory store instruction, and one may be stored for the LSID associated with a memory load instruction. As used herein, memory load and store instructions refer to processor instructions that operate on memory, while read and write instructions refer to register reads and writes, e.g., reads and writes to and from a register file. The memory mask may be stored in registers accessible by the control unit of the block-based processor core. In other examples, the memory mask may be stored in a small memory or stored using other suitable techniques.

[0111] The memory mask can be generated in any suitable manner. In some examples, the memory mask is generated by reading the memory mask encoded in the instruction block header by the compiler that generates the instruction block. In some examples, the memory mask is generated from a memory location storing previously generated memory masks. For example, a binary file for a block-based processor program may include a section that stores memory masks for any number of instruction blocks in the program. In some examples, previously generated memory masks are cached from previous executions of the instruction block and do not need to be regenerated for subsequent instances of the instruction block. In some examples, the memory mask is generated by generating a new memory mask during instruction decoding of the instruction block. For example, as each load and store instruction is decoded by the instruction decoder, the LSID field is extracted and appropriate bits in the memory mask can be set to indicate whether the LSID corresponds to a load or store instruction. In some examples, more than one bit is used to encode the LSID in the memory mask, for example, without considering or with an empty LSID. Once the memory mask has been generated, the method proceeds to process block 820. In some examples, the memory mask is generated from a previous instance of the instruction block that executed it. In other examples, it is generated by an instruction decoder that decodes memory access instructions.

[0112] At procedure block 820, one or more instructions in the currently executing instruction block are decoded, issued, and executed. If the instruction is a memory access instruction (such as a memory load or memory store), the memory vector can be updated. For example, the memory vector can have bits corresponding to each LSID within the instruction block. When a load or store instruction with an encoded LSID is executed, the corresponding bits in the memory vector are then set. Thus, the memory vector can indicate which memory access instructions in the instruction block have been executed. In other examples, other techniques can be used to update the memory vector; for example, a counter can be used instead of a memory vector, as discussed further in detail below. It should be noted that in some examples, the LSID is unique for each instruction in the block. In other words, each LSID value can only be used once within the instruction block. In other examples, for example, in the case of assertion instructions, the same LSID can be encoded for two or more instructions. Therefore, a set of instructions asserting true conditions can have some or all of their LSIDs overlap with the corresponding instructions asserting false conditions. Once the memory vector is updated, the method proceeds to procedure block 830.

[0113] At procedure block 830, the LSID used for the instruction is compared with the masked memory vector. In some examples, the block-based processor control unit is configured to compare memory vector register data with memory mask data from the hardware architecture to determine which memory store instructions in the memory store instruction set have been executed. The memory vector updated at procedure block 820 is combined with the memory mask generated at procedure block 810 to produce the value used in the comparison. For example, bitwise logical AND or OR operations can be used to mask the memory vector using the memory mask. The masked memory vector indicates which LSIDs are executable. For example, if the masked memory vector has all the bits set for instructions zero through five, then instruction number six is ​​acceptable to issue. Based on the comparison, the method continues as follows. If the comparison indicates that an instruction can be issued based on the LSID comparison, the method proceeds to procedure block 840. On the other hand, if the instruction in question is not acceptable to issue, the method proceeds to procedure block 820 to execute additional instructions in the instruction block and update the memory vector accordingly.

[0114] At process block 840, the load or store instruction associated with the LSID used in the comparison at process block 830 is issued to the execution phase of the processor pipeline. In some examples, the next memory load or memory store instruction to be executed is selected, at least in part, based on the LSID encoded within the instruction block and a memory vector register storing data indicating which memory store instructions have already been executed. Thus, an instruction can continue execution as its memory dependencies, as indicated by the comparison of its LSID with the masked memory vector, have been satisfied. In some examples, other dependencies may cause the issued instruction to be delayed due to factors unrelated to the masked memory vector comparison.

[0115] In some examples of the method outlined in flowchart 800, the block-based processor core includes an instruction unit configured to execute an instruction block encoded with multiple instructions, each of which may be issued based on a dependency specified for the corresponding instruction. The processor core also includes a control unit configured to control the issuance of memory load and / or memory store instructions within the instruction block to the execution unit based at least in part on data stored in a hardware structure indicating the relative ordering of loads and stores within the instruction block. In some examples, the hardware structure may be a memory mask, content-addressable memory (CAM), or a lookup table. In some examples, the data is stored in a hardware structure generated from a previous instance of the executed instruction block. In some examples, the data is stored in a hardware structure from data decoded from the instruction block header used for the instruction block. In some examples, the control unit includes a memory vector register for storing data indicating which memory access instructions (e.g., memory load and / or memory store instructions) have been executed. In some examples, the processor core control unit is configured to prevent the commit of the instruction block until the memory vector indicates that all memory access instructions have been executed. In some examples, the processor control unit instructs a counter to be updated (e.g., incremented) when a memory load or memory store instruction is executed, and indicates that an instruction block is completed when the counter reaches a predetermined value for the number of memory access instructions. In some examples, the processor core is configured to execute assertion instructions, including assertion memory access instructions.

[0116] XI. Example source code and target code

[0117] Figure 9 An example of source code 910 and corresponding object code 920 for a block-based processor, which may be used in some examples of the disclosed technology, is illustrated. Source code 910 includes if / else statements. Statements within each part of the if / else statements include reads and writes to multiple memories in arrays A and B. When source code 910 is transformed into object code, multiple load and store assembly instructions are generated.

[0118] The assembly code 920 used for source code section 910 includes 25 instructions numbered 0 to 24. The assembly instructions indicate multiple fields, such as the instruction opcode (pneumonic), metadata specified by the instruction (e.g., broadcast identifier or immediate argument), load memory ID identifier, and target identifier. The assembly code includes register read instructions (0-2), register write instructions (instruction 24), arithmetic instructions (e.g., instructions 3 and 4), and move instructions for sending data to multiple targets (e.g., move instructions 5 and 6). Assembly code 920 also includes test instruction 11, which is a test to determine if an assertion value generated on broadcast channel 2 is greater than or equal to a given value. Additionally, the assembly code includes two unasserted memory load instructions 7 and 8, and one asserted load instruction 16. Load instruction 23 is also unasserted. Assembly code 920 also includes multiple memory store instructions for storing data to memory addresses, such as asserted store instructions 12, 13, and 18, and an unasserted store instruction 21. As shown in assembly code 920, each load and store instruction in the load and store instruction set has been assigned a unique LSID. For example, load instruction 7 is assigned to LSID 0, load instruction 8 to LSID 1, and assertion store instruction 12 to LSID 2. LSIDs indicate the relative order in which instructions will be executed. For example, instructions 12 and 13 depend on load instructions 7 and 8 being executed first. This order is mandatory because load instructions 7 and 8 are used to generate the values ​​that will be stored by store instructions 12 and 13. In some examples, two or more load-store instructions may share an LSID. In some examples, the instruction set architecture requires LSIDs to be consecutive, while in other examples, LSIDs can be sparse (e.g., skipping intermediate LSID values). It should also be noted that in some examples, speculative or out-of-order execution of instructions within a block is possible, but the processor must still maintain semantics as if the memory dependencies specified by the LSIDs have not been violated. In some examples, whether memory access instructions can be issued out of order may depend on the memory addresses computed at runtime.

[0119] The assembly code section 920 can be converted into machine code for actual execution by a block-based processor.

[0120] XII. Example control flow graph

[0121] Figure 10 The diagram illustrates the above regarding... Figure 9The control flow graph 1000 is generated from the assembly code 920 described. For ease of illustration, it is depicted in the form of nodes and edges, but the control flow graph 1000 can be represented in other forms (e.g., according to a suitable graphical data structure, the arrangement of data in memory), as will be readily apparent to those skilled in the art. For ease of illustration, only load and store instructions from the assembly code 920 are shown in the control flow graph; however, it should be understood that the nodes of the control flow graph will be placed or referenced to other instructions according to the dependencies and assertions of each corresponding instruction.

[0122] As shown, the first node 1010 includes load instructions 7 and 8 associated with LSID 0 and 1, respectively. Instructions 7 and 8 are unassertified and can be issued and executed as soon as their operands become available. For example, the assembly code move instruction 5 sends the memory address associated with a[i] to move instruction 5, which in turn sends the address to load instruction 7. Load instruction 7 can be executed as soon as the address becomes available. Other instructions (such as fetch instructions 0 through 2) can also be executed without reference assertions.

[0123] If node 1020 is generated by conditional instruction 11, which generates a Boolean value by comparing two values ​​(e.g., for a test where one operand is greater than the other), the condition is asserted to be true if the left operand of the test instruction is greater, and only the instructions in code section 1030 will be executed. Conversely, if the condition is false, code section 1035 will be executed. In the disclosed block-based processor architecture, this can be executed without using branches or jumps because the associated instructions are asserted. For example, instruction 12 is a store instruction asserted on broadcast channel 2, generated by test instruction 11. Similarly, instruction 16 will be executed if the broadcast assertion is false. The store instructions in code section 1030 are associated with LSIDs 2 and 3, while the load and store instructions in code section 1035 are associated with LSIDs 4 and 5. As each instruction in the instruction set is executed, the memory vector is updated to indicate that an instruction has been executed. The control flow graph 1000 also includes a junction node 1040, which represents a transition back to statements outside the if / else statements contained in source code 910. For example, instructions 21 and 23 of code section 1050 are placed after the if / else statements. Instructions 21 and 23, as shown, have LSIDs 6 and 7. It should be noted that the compiler generating the assembly code 920 did not place memory access instructions 21 and 23 with code section 1010 because they may depend on values ​​generated within code sections 1030 or 1035. For example, load instruction 23 reads from array b at index 2, which may or may not be written to by store instruction 18 of code section 1035 depending on the value of i. It should be noted that although memory access instructions are executed according to the relative order encoded by LSIDs, the instructions will also wait for other dependencies before being issued.

[0124] XIII. Example storage mask / Vector comparison

[0125] Figure 11A and Figure 11B The illustration shows examples of comparing a storage mask with a storage vector, which can be performed at certain examples of the disclosed techniques. For example, as... Figure 11AAs shown, memory mask 1100 stores either 1 or 0 values ​​(bits ordered starting from LSID 0 on the left) for the LSIDs associated with assembly code 920. Therefore, load memory IDs 0, 1, 4, and 7 are associated with memory load instructions, while LSIDs 2, 3, 5, and 6 are associated with memory store instructions. The memory mask bits are set to 1 for load instructions and to 0 for store instructions. Memory store vector 1110 is shown in the state after instructions 7, 8, and 12 have been executed. Therefore, the bits associated with LSIDs 0, 1, and 2 are set to 1, while the corresponding bits for unexecuted instructions are set to 0. A bitwise OR gate 1120 is used to compare memory mask 1100 and memory mask vector 1110 to produce a masked memory vector 1130. The masked memory vector indicates which instructions are allowed to proceed. As shown, since the LSID order is maintained, the next instruction to be executed is the one associated with LSID 3 (instruction ID 13 in this example). In some examples, the LSID associated with the unacquired instruction can be marked as acquired (in other words, invalidated) by placing a 1 in the instruction of the unacquired assertion.

[0126] Figure 11B The illustration shows an example of generating a masked memory vector comparison after an additional number of instructions have been executed. As shown, the memory vector has been updated to indicate that instructions 0 through 5 have been executed. If the assertion result is true, the LSID associated with the unacquired assertion can be marked as 1. The masked memory result 1140 shows that the instruction associated with the first 0 (i.e., LSID6) has satisfied its previous instruction dependency and is ready to be issued.

[0127] XIV. Example method for issuing instructions according to a specified order

[0128] Figure 12 This is a flowchart 1200 outlining an example method for issuing instructions based on a relative ordering and counter specified in the instructions, which can be executed at certain examples of the disclosed technology. For example, the above regarding... Figure 1 and Figure 2 The block-based processor 100 and processor core 111 discussed can be used to implement the method shown.

[0129] In process block 1210, the asserted load or store instruction is executed. After the instruction is executed, a counter is updated (e.g., incremented by 1). Therefore, the counter can be used to indicate the execution status of memory access instructions instead of the memory storage vector discussed above. In some examples, a block-based processor control unit is configured to compare memory vector register data with memory mask data from the hardware architecture to determine if all memory store instructions ordered before the current memory access instruction have been executed, and based on this determination, issue the current memory access instruction to the execution unit. In some examples, the control unit includes a counter that is incremented during the execution of one of the memory load and / or memory store instructions, and wherein the control unit indicates that the instruction block has been completed when the counter reaches a predetermined value for the number of memory access instructions encoded in the instruction block. Once the counter has been updated, the method proceeds to process block 1220.

[0130] At process block 1220, the load memory identifier associated with the unacquired assertion path is invalidated. For example, if an assertion associated with assertion node 1320 discussed below is acquired, the memory access instruction associated with the unacquired portion can be invalidated. After invalidating the identifier depending on the acquired assertion path, the method proceeds to process block 1030.

[0131] At procedure block 1030, the load memory identifier for the next instruction is compared with a counter. If the counter is set to a value indicating that issuing the next instruction is acceptable, the method proceeds to procedure block 1240. If the comparison indicates that issuing a memory access instruction is unacceptable, the method proceeds to procedure block 1210 to execute additional instructions. For example, if the counter indicates that five instructions have been executed and the LSID of the next memory access instruction is 6, then issuing a memory access instruction is acceptable. Conversely, if the counter value is less than five, issuing a memory access instruction is inappropriate. In some examples, the number of store instructions or the number of memory access instructions is stored in the instruction block header as an LSID counter 517.

[0132] XV. Example control flow image

[0133] Figure 13 The illustration shows an example alternative control flow diagram that can be used to represent a slightly modified version of the control flow for assembly code 920 in other examples of the disclosed technology. For example... Figure 13As shown, code section 1335 contains a null instruction 18. Aside from being used to adjust the memory vector or counter to indicate that the next memory access instruction is ready to be issued, the null instruction 18 does not change the processor's state. For example, if an assertion is obtained at node 1320, both memory access instructions will be issued. Conversely, if an assertion is not obtained at node 1320, only one memory access instruction will be executed without a null instruction. Therefore, the null instruction is a way to balance the load memory IDs indicated in the memory vector. This simplifies the control flow for instruction blocks because the number of repositories for an instruction block can be the same, or the LSIDs at the junction nodes of the control flow graph can be the same, regardless of the assertion path obtained during the execution of the instruction block. In other examples, unbalanced conditions can be identified by the processor control unit and can automatically invalidate LSIDs without including null instructions in the instruction block code. It should also be noted that source code section 1320 has overlapping LSIDs (i.e., LSID2 and 3) as shown in source code section 1335. Since only one side of condition node 1320 will be acquired, it is possible to overlap and share the same LSID by allowing fewer bits to be used to encode the LSID, which allows for better allocation of LSID values. For example, source code section 1350 has instructions to assign LSIDs 4 and 5 instead of 6 and 7, as... Figure 10 The situation is the same as the control flow graph.

[0134] XVI. Example methods for transforming code

[0135] Figure 14 This is a flowchart 1400 outlining examples of transforming code into computer-executable code for a block-based processor, as can be executed at certain examples of the disclosed technologies. For example, general-purpose processors and / or block-based processors can be used to implement... Figure 14 The methods outlined herein. In some examples, the code is transformed by a compiler and stored as target code that can be executed by a block-based processor (e.g., block-based processor 100). In some examples, a just-in-time compiler or interpreter generates computer-executable code at runtime.

[0136] At procedure block 1410, memory references encoded in the source and / or object code are analyzed to determine memory dependencies. For example, memory dependencies can simply be the order in which memory access instructions are arranged in the program. In other examples, memory addresses likely to be written to by memory access instructions can be analyzed to determine if there is overlap between load memory instructions in the instruction block. In some examples, determining memory dependencies involves identifying two memory access instructions in the instruction block, asserting a first memory access instruction on a complementary condition of a second memory access instruction, and assigning the same identifier to both the first and second memory access instructions based on the identifier. After analyzing the memory references, the method proceeds to procedure block 1420.

[0137] At process block 1420, the source code and / or object code is transformed into block-based computer-executable code, which includes an indication of the relative ordering of memory access instructions within an instruction block. For example, LSID values ​​may be encoded in the instructions. In other examples, the relative ordering is indicated by instruction positioning within the block. In some examples, a memory mask is generated and stored as an instruction block header for the instruction block. In some examples, the memory mask indicates which load / store identifier in the load / store identifier corresponds to a memory access instruction. In some examples, special instructions are provided to load the memory mask into the control unit's memory for use when masking memory vectors. Once the code has been transformed into block-based processor code, it can be stored on a computer-readable storage medium or transferred via a computer network to another location for execution by a block-based processor.

[0138] XVII. Exemplary computing environment

[0139] Figure 15 The illustration depicts a generalized example of a suitable computing environment 1500 in which the described embodiments, skills, and techniques (including configuring a block-based processor) can be implemented. For example, computing environment 1500 can implement the disclosed skills, as described herein, for configuring a processor to generate and use memory access instruction sequence encoding or encoding code into computer-executable instructions for performing such operations.

[0140] The Computing Environment 1500 is not intended to impose any limitations on the scope of the technology's use or functionality, as the technology can be implemented in various general-purpose or special-purpose computing environments. For example, the disclosed technology can be implemented using other computer system configurations, including handheld devices, multiprocessor systems, programmable consumer electronics, network PCs, microcomputers, mainframes, and so on. The disclosed technology can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules (including executable instructions for block-based instruction blocks) can be located in both local and remote memory storage devices.

[0141] Reference Figure 15 The computing environment 1500 includes at least one block-based processing unit 1510 and a memory 1520. Figure 15 The most basic configuration 1530 is included within the dashed lines. A block-based processing unit 1510 executes computer-executable instructions and can be a physical or virtual processor. In a multiprocessor system, multiple processing units execute computer-executable instructions to increase processing power, and thus, multiple processors can operate simultaneously. Memory 1520 can be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or a combination of both. Memory 1520 stores, for example, software 1580, images, and video that can implement the techniques described herein. Memory 1520 can be accessed by a block-based processor using memory access instructions discussed herein, including load and store instructions with LSIDs. The computing environment can have additional features. For example, computing environment 1500 includes a storage device 1540, one or more input devices 1550, one or more output devices 1560, and one or more communication connections 1570. Interconnect mechanisms (not shown) (e.g., buses, controllers, or networks) interconnect the components of computing environment 1500. Typically, operating system software (not shown) provides an operating environment for other software running in computing environment 1500 and coordinates the activities of components of computing environment 1500.

[0142] Storage device 1540 may be removable or non-removable and includes a disk, magnetic tape or tape cartridge, CD-ROM, CD-RW, DVD, or any other medium that can be used to store information and is accessible within computing environment 1500. Storage device 1540 stores instructions, inserted data, and messages for software 1580, which can be used to implement the techniques described herein.

[0143] One or more input devices 1550 may be touch input devices, such as a keyboard, keypad, mouse, touchscreen display, pen or trackball, voice input device, scanning device, or another device that provides input to computing environment 1500. For audio, one or more input devices 1550 may be a sound card or similar device that accepts audio input in analog or digital form, or a CD-ROM reader that provides audio samples to computing environment 1500. One or more output devices 1560 may be a display, printer, speaker, CD burner, or another device that provides output from computing environment 1500.

[0144] Communication connection 1570 enables communication with another computing entity via a communication medium (e.g., a network connection). The communication medium transmits information, such as computer-executable instructions, compressed graphic information, video, or other data in modulated data signals. Communication connection 1570 is not limited to wired connections (e.g., megabit or gigabit Ethernet, wireless bandwidth, fiber optic channels over electrical or fiber optic connections), but also includes wireless technologies (e.g., via Bluetooth, WiFi (IEEE 802.11a / b / n), WiMax, cellular, satellite, laser, infrared RF connections) and other suitable communication connections for providing network connectivity for the disclosed methods. In a virtual hosting environment, the communication connection can be a virtualized network connection provided by the virtual host.

[0145] Some embodiments of the disclosed methods can be executed using computer-executable instructions that implement all or part of the techniques disclosed in Computing Cloud 1590. For example, the disclosed compiler and / or a server with a block-based processor is located in a computing environment, or the disclosed compiler can be executed on a server located in Computing Cloud 1590. In some examples, the disclosed compiler executes on a conventional central processing unit (e.g., a RISC or CISC processor).

[0146] Computer-readable media is any available medium that can be accessed within computing environment 1500. By way of example, and not limitation, using computing environment 1500, computer-readable media includes memory 1520 and / or storage device 1540. As should be readily understood, the term computer-readable storage medium includes media for data storage (such as memory 1520 and storage device 1540) but not transmission media (such as modulated data signals).

[0147] X. Additional examples of publicly available technologies

[0148] Additional examples of the disclosed topics are discussed here, based on the examples discussed above.

[0149] In some examples of the disclosed technology, an apparatus includes a memory and one or more block-based processor cores, at least one of the cores comprising: an execution unit configured to execute memory access instructions contained in an instruction block, including a plurality of memory load and / or memory store instructions; a hardware structure storing data indicating the execution order of at least some of the memory access instructions; and a control unit configured to control the issuance of memory access instructions to the execution unit based at least in part on the hardware structure data.

[0150] In some examples of the device, the hardware structure is a memory mask, content-addressable memory (CAM), or lookup table. In some examples, data stored in the hardware structure is generated from a previous instance of an instruction block being executed. In some examples, data stored in the hardware structure is generated from the instruction block header used for the instruction block. In some examples, the data stored in the hardware structure is generated by an instruction decoder that decodes memory access instructions.

[0151] In some examples of the device, the control unit includes a storage vector register storing data indicating which memory access instructions in a memory access instruction sequence have been executed. In some examples, the control unit is also configured to compare the storage vector register data with memory mask data from the hardware architecture to determine which memory access instructions in a memory access instruction sequence have been executed. In some examples, the control unit is also configured to compare the storage vector register data with memory mask data from the hardware architecture to determine that all memory access instructions prior to the current memory access instruction in the memory access instruction sequence have been executed, and based on the determination, issue the current memory access instruction to a load / store queue coupled to the control unit. In some examples, the control unit includes a counter that is incremented as one of the memory access instructions is executed, and the control unit indicates that the instruction block has been completed when the counter reaches a predetermined value for the number of memory access instructions encoded in the instruction block. In some examples, the control unit generates signals used to control wake-up / selection logic and / or one or more memory load / store queues coupled to the control unit. In some examples, the control unit generates signals used to directly control components within the processor core and / or memory cells. In some examples, the wake-up / selection logic and / or memory load / store queue perform some or all of the operations related to the generation and use of memory access instruction sequence codes.

[0152] In some examples of the apparatus, the data instructing the execution of the sorting is based at least in part on load / store identifiers encoded for each memory access instruction within the instruction block. In some examples, the apparatus is block-based, using the processor itself. In some examples, the apparatus includes a computer-readable storage medium storing data for the instruction block header and for memory access instructions within the instruction block.

[0153] In some examples of the disclosed technology, a method of operating a processor to execute a block of instructions comprising multiple memory load and / or memory store instructions includes selecting the next memory load or memory store instruction to execute based at least in part on dependencies encoded within the block of instructions and at least in part on a storage vector register of stored data indicating which memory store instructions among the memory store instructions have been executed, and executing the next instruction.

[0154] In some examples of the method, the choice includes comparing the memory mask encoded in the header of the instruction block with the memory vector register data. In some examples, dependencies are encoded using identifiers in each instruction within the memory load and / or memory store instructions. In some examples, execution of at least one instruction in the memory load and memory store instructions is asserted on a condition value generated by another instruction in the instruction block.

[0155] Some examples of the method also include comparing storage vector register data with a storage mask for a block of instructions that indicates which of the encoded dependencies corresponds to a memory-store instruction, and based on the comparison, performing one of the following actions: stopping the execution of the next instruction, stopping the execution of the next instruction block, generating a memory dependency prediction, initiating the execution of the next instruction block, or initiating an exception handle to indicate a memory access error.

[0156] Some examples of the method also include memory store instructions for executing blocks of instructions, which are encoded with identifiers indicating the location of the instructions in a relative order, an indication of whether the memory store instructions are executed is stored in a storage vector register, and memory load instructions with a later relative order are executed based on the indication.

[0157] In some examples of the method, selection is also based on comparing a counter value with an identifier encoded in the next instruction of selection, the counter value indicating the number of memory load and / or memory store instructions that have been executed. In some examples, the LSID values ​​are consecutive (e.g., 1, 2, 3, ..., n), while in other examples, the LSID values ​​are not all consecutive (e.g., 1, 2, 3, 5, 7, 9, ..., n). In some examples, the relative order is specified by one or more of the following: data encoded in the instruction block header, data encoded in one or more instructions of the instruction block, data encoded in a table dynamically generated during the decoding of the instruction block, and / or cached data encoded during previous execution of the instruction block.

[0158] In some examples of the disclosed techniques, one or more computer-readable storage media store computer-readable instructions that, when executed by a block-based processor, cause the processor to perform any or more of the disclosed methods.

[0159] In some examples of the disclosed technology, one or more computer-readable storage media store computer-readable instructions for instruction blocks. When executed by a block-based processor, these computer-readable instructions cause the processor to perform methods. These computer-readable instructions include instructions for analyzing memory accesses encoded in source code and / or object code to determine memory dependencies for the instruction block, and instructions for transforming the source code and / or object code into computer-executable code for the instruction block, the computer-executable code including indications of the ordering of memory access instructions within the instruction block. In some examples, the computer-readable instructions include: instructions for identifying two memory access instructions within the instruction block, wherein a first memory access instruction is asserted on a complementary condition of a second memory access instruction, and instructions for assigning the same identifier to the first and second memory access instructions based on the identifier. In some examples, the ordering indication is indicated by load / store identifiers encoded in the memory access instructions, and the instructions include instructions for generating a memory mask in the header of the instruction block indicating which load / store identifiers correspond to stored memory access instructions.

[0160] Given the many possible embodiments to which the principles of the disclosed subject matter can be applied, it should be recognized that the illustrated embodiments are merely preferred examples and should not be construed as limiting the scope of the claims to those preferred examples. Rather, the scope of the claimed subject matter is defined by the appended claims. We therefore claim protection under our invention for all that falls within the scope of these claims.

Claims

1. An apparatus comprising a memory and at least one processor core, the at least one processor core comprising: an execution unit configured to execute memory access instructions, the memory access instructions comprising memory load instructions, memory store instructions, or both memory load and memory store instructions, the memory access instructions being contained in an instruction block, each of the memory access instructions being associated with a respective one of a plurality of load / store identifiers; and a control unit comprising: a load / store counter, the load / store counter being updated upon execution of one of the memory access instructions, the load / store counter indicating a number of the memory access instructions that have been executed by the execution unit, the control unit being configured to control issuance of the memory access instructions to the execution unit in an execution order by comparing one of the load / store identifiers associated with the memory access instructions to a value stored in the load / store counter to determine whether an additional one of the memory access instructions is suitable for issuance.

2. The apparatus of claim 1, wherein the load / store identifiers are encoded within the memory access instructions.

3. The apparatus of claim 1, wherein the control unit indicates that the instruction block has completed when the load / store counter reaches a predetermined value for a number of memory access instructions encoded in the instruction block.

4. The apparatus of claim 1, wherein at least one of the memory access instructions is asserted on a condition value generated by another instruction of the instruction block.

5. The apparatus of claim 1, wherein at least one of the memory access instructions is asserted on a complemented condition of a different one of the memory access instructions.

6. The apparatus of claim 1, wherein the load / store counter is incremented when the memory store instruction is executed.

7. A method of operating a processor to execute a plurality of memory load and memory store instructions, the method comprising: comparing data stored in a storage vector register to a storage mask of the plurality of memory load and memory store instructions, the storage mask indicating which of a plurality of load / store identifiers corresponds to a respective memory load or memory store instruction, the load / store identifiers specifying a relative execution order of the plurality of memory load and memory store instructions; selecting the memory instructions of the plurality of memory load and memory store instructions to execute based on the respective load / store identifiers associated with the memory instructions and the comparison to the data; and executing the selected memory instructions.

8. The method of claim 7, further comprising: Based on the comparison, one of the following is performed: stopping execution of the selected memory instruction, generating a memory dependency prediction, or initiating an exception handler to indicate a memory access error.

9. The method of claim 7, wherein the respective load / store identifiers are encoded within the memory instructions to which the respective load / store identifiers are associated.

10. The method of claim 7, further comprising: invalidating a load / store identifier associated with an unfulfilled one of the plurality of memory load and store instructions.

11. The method of claim 7, further comprising: performing a no-op instruction to balance a load / store identifier of an unfulfilled one of the plurality of memory load and store instructions, wherein the inclusion of the no-op instruction balances the number of memory access instructions between assertion paths.

12. The method of claim 7, wherein the executing the selected memory instruction is asserted on a condition value generated by another instruction.

13. The method of claim 7, wherein the executing the selected memory instruction is asserted on a complement condition of another instruction.

14. An apparatus comprising a memory and at least one processor core, the at least one processor core comprising: an execution unit configured to execute memory access instructions, the memory access instructions comprising a plurality of memory load instructions, memory store instructions, or memory load and store instructions included in an instruction block; and a control unit comprising: a store vector register storing data indicating which of the memory access instructions have executed, and the control unit configured to operate control of issuance of the memory access instructions to the execution unit, the operation comprising comparing load / store identifiers associated with the memory access instructions to the store vector register data indicating which of the memory access instructions have executed.

15. The apparatus of claim 14, wherein the memory access instructions comprise memory store instructions, and wherein the control unit is further configured to compare the store vector register data indicating which of the memory access instructions have executed to store mask data to determine which of the memory store instructions have executed.

16. The apparatus of claim 14, wherein the memory access instructions comprise memory store instructions, and wherein the control unit is further configured to compare the store vector register data indicating which of the memory access instructions have executed to store mask data to determine that a memory store instruction ordered before a current memory access instruction of the memory access instructions has executed, and based on the determination, to issue the current memory access instruction to a load / store queue coupled to the control unit.

17. The apparatus of claim 14, wherein the load / store identifiers are encoded within the memory access instructions in the instruction block.

18. The apparatus of claim 14, wherein the control unit is further configured to generate a signal to control when the memory access instructions are sent to a load / store queue or one or more memory load / store queues coupled to the control unit.

19. The apparatus of claim 14, wherein execution of at least one of the plurality of memory load instructions, memory store instructions, or memory load and memory store instructions is predicated on a condition value generated by another instruction of the instruction block.

20. The apparatus of claim 14, wherein a first one of the plurality of memory load instructions, memory store instructions, or memory load and memory store instructions is predicated on a complementary condition of a second one of the plurality of memory load instructions, memory store instructions, or memory load and memory store instructions.