Low latency synchronization for operation cache and instruction cache fetch and decode instructions
By introducing shared extraction logic and synchronization mechanisms into the processor, the problems of high power consumption and performance bottlenecks during instruction extraction and decoding are solved, and more efficient instruction processing and lower latency are achieved.
Patent Information
- Application Number
- CN201980041084.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-21
- Filing Date
- 2019-05-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-05-15
AI Technical Summary
The prior art has problems with high power consumption and performance bottlenecks in the instruction extraction and decoding process, especially when switching from the operation cache and the instruction cache can result in delays.
By introducing shared extraction logic and synchronization mechanisms, it is determined that the predicted address block should be extracted from the operation cache path or the instruction cache path, and detect multiple operation cache hits in a single cycle, reducing switching delays.
The minimum delay of switching between different paths is achieved, the efficiency of instruction extraction and decoding is improved, and the power consumption is reduced.
Smart Images

Figure CN112313619B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Non-Provisional Patent Application No. 16 / 014,715, filed on June 21, 2018, the contents of which are incorporated herein by reference. Background Art
[0003] The microprocessor instruction execution pipeline fetches instructions and decodes them into micro-operations for execution. Instruction fetching and decoding consumes a lot of power and can also be a performance bottleneck. Improvements to instruction fetching and decoding are continually being made. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] A more detailed understanding may be obtained from the following description given by way of example in conjunction with the accompanying drawings, in which:
[0005] Figure 1 is a block diagram of an example apparatus in which one or more disclosed features may be implemented;
[0006] Figure 2 Based on the example Figure 1 A block diagram of the instruction execution pipeline of a processor;
[0007] Figure 3 According to the example shown Figure 2 A block diagram of components of an instruction fetch and decode unit of an instruction execution pipeline;
[0008] Figure 4A is a flow chart of a method of operation of shared fetch logic of an instruction fetch and decode unit according to an example;
[0009] Figure 4B is a flow chart of a method for processing entries of an operation cache queue to extract cached decoded micro-ops from the operation cache according to an example; and
[0010] Figure 4C is a flow chart of a method for fetching and decoding instruction bytes stored in an instruction byte buffer according to an example. DETAILED DESCRIPTION
[0011] The instruction fetch and decode unit includes an operation cache that stores previously decoded instructions and an instruction cache that stores undecoded instruction bytes. The instruction fetch and decode unit fetches instructions corresponding to a predicted address block predicted by a branch predictor. The instruction fetch and decode unit includes a fetch control block that determines whether the predicted address block should be fetched from an operation cache path or an instruction cache path, and which entries in these caches hold the associated instructions. When instructions are available in the operation cache, the operation cache path is used. The operation cache path retrieves decoded micro-ops from the operation cache. The instruction cache path retrieves instruction bytes from the instruction cache and decodes those instruction bytes into micro-ops.
[0012] The fetch control logic checks the tag array of the operational cache to detect hits for the predicted address block. Due to special features of the fetch control logic described elsewhere herein, multiple hits can be detected in a single cycle in the event that more than one operational cache entry is required to fetch the predicted address block. For instructions that do not find a hit in the operational cache, the instruction cache path fetches instruction bytes from the instruction cache or a higher level cache and decodes these instructions.
[0013] The fetch control logic generates operation cache queue entries. These entries indicate whether the instruction preceding the instruction address of the entry will be serviced by the instruction cache path, and therefore whether the operation cache path must wait for the instruction cache path to output the decoded operation of the previous instruction before outputting the decoded micro-operation of the operation cache queue entry. The fetch control logic also generates instruction cache queue entries for the instruction cache paths, which indicate whether the decoded micro-operation of the instruction corresponding to the instruction byte buffer entry must wait for the operation cache path to output the micro-operation of the previous instruction before outputting themselves. Therefore, for any particular decoded operation, both paths know whether such operation must wait for the decoded operation from the opposite path to be output. The synchronization mechanism allows work to proceed in either path until the work needs to be stopped because it must wait for another path. For predicted address blocks, the combination of the ability to detect multiple operation cache hits in a single cycle and the synchronization mechanism allows switching between different paths with minimal delay.
[0014] Figure 11 is a block diagram of an example device 100 in which aspects of the present disclosure are implemented. Device 100 includes, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, or a tablet computer. Device 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. Device 100 includes an input driver 112 and an output driver 114. It should be understood that device 100 may include Figure 1 Additional components not shown.
[0015] As is well known, processor 102 is a computing device capable of executing software, such as a microprocessor, microcontroller, or other device. Memory 104 stores instructions and data for use by processor 102. In one example, memory 104 is located on the same die as processor 102. In another example, memory 104 is located on a different die from processor 102. Memory 104 includes volatile memory, such as random access memory (RAM), dynamic RAM, or cache. In some examples, memory 104 includes non-volatile memory.
[0016] Storage device 106 includes fixed or removable memory, such as a hard drive, solid state drive, optical disk, or flash drive. Input device 108 includes a keyboard, a keypad, a touch screen, a touch pad, a detector, a microphone, an accelerometer, a gyroscope, a biometric scanner, a network connection (e.g., a wireless LAN card for transmitting and / or receiving wireless IEEE 802 signals), and / or other input devices. Output device 110 includes a display, a speaker, a printer, a tactile feedback device, one or more lights, an antenna, a network connection (e.g., a wireless LAN card for transmitting and / or receiving wireless IEEE 802 signals), and / or other output devices.
[0017] The input driver 112 communicates with the processor 102 and the input device 108 and allows the processor 102 to receive input from the input device 108. In various examples, the device 100 includes one or more input drivers 112 (although only one is shown). The input driver 112 is embodied as custom fixed-function hardware, programmable hardware, software executed on a processor (such as the processor 102), or any combination thereof. In various examples, the input driver 112 includes an expansion card inserted into a port (such as a peripheral component interconnect express (PCIe) port) that is coupled to both the processor 102 and the input device 108. The output driver 114 communicates with the processor 102 and the output device 110 and allows the processor 102 to send output to the output device 110. In various examples, the device 100 includes one or more output drivers 114 (although only one is shown). The output driver 114 is embodied as custom fixed-function hardware, programmable hardware, software executed on a processor (such as the processor 102), or any combination thereof. In various examples, the output driver 114 comprises an expansion card plugged into a port, such as a Peripheral Component Interconnect Express (PCIe) port, that is coupled to both the processor 102 and the input device 108 .
[0018] Figure 2 is included in the example Figure 1 1 is a block diagram of an instruction execution pipeline 200 within the processor 102 of FIG. The instruction execution pipeline 200 retrieves instructions from memory and executes the instructions, outputs data to memory, and modifies the state of elements within the instruction execution pipeline 200 (such as registers within the register file 218).
[0019] The instruction execution pipeline 200 includes an instruction fetch and decode unit 202, which fetches instructions from a system memory (such as memory 104) via an instruction cache and decodes the fetched instructions. Decoding the fetched instructions converts the fetched instructions into micro-operations (also just "operations") for execution by the instruction execution pipeline 200. The term "instruction" refers to a task specified in the instruction set architecture for the processor 102. Instructions can be specified to be executed by software. Micro-operations are sub-tasks that are generally not directly usable by software. Instead, micro-operations are single tasks that the processor 102 actually performs in order to execute the instructions requested by the software. Therefore, decoding instructions includes identifying control signals to be applied to the functional units 216, the load / store unit 214, and other parts of the instruction execution pipeline 200. Decoding some instructions results in multiple micro-operations per instruction, while decoding other instructions results in one micro-operation per instruction.
[0020] Execution pipeline 200 also includes functional units 216 that perform calculations to process micro-operations, load / store units 214 that load or store data from or to system memory via data cache 220 as specified by the micro-operations, and register files 218 that include registers that store working data for the micro-operations.
[0021] The reorder buffer 210 tracks the instructions currently being transferred and ensures orderly retirement of the instructions, although out-of-order execution is allowed when transferred. A "transferred" instruction is an instruction that has been received by the reorder buffer 210 but has not yet "retired" - that is, the result has not yet been submitted to the processor's architectural state (e.g., the result of writing an architectural register). When all micro-operations of an instruction are executed, the instruction is considered to have retired. The reservation station 212 maintains the transferred micro-operations and tracks the operands for the micro-operations. When all operands are ready to execute a particular micro-operation, the reservation station 212 sends the micro-operation to the functional unit 216 or the load / store unit 214 for execution.
[0022] The various elements of the instruction execution pipeline 200 communicate via a common data bus 222. For example, the functional units 216 and the load / store unit 214 write results to the common data bus 222, which may be written to the reservation station 212 to execute the associated instructions / micro-ops, and to the re-order buffer 210 as final processing results for in-progress instructions that have completed execution.
[0023] In some techniques, the instruction fetch and decode unit 202 includes an operation cache that stores micro-operations corresponding to previously decoded instructions. During operation, the instruction fetch and decode unit 202 checks the operation cache to determine whether decoded micro-operations are stored in the operation cache, and outputs such found micro-operations instead of performing the decoded operations on those instructions. Since decoding instructions is relatively power-hungry and has high latency, using an operation cache helps reduce power consumption and improve processing throughput.
[0024] One problem with the operation cache is due to the need to switch between using the operation cache when decoded micro-ops are present in the operation cache and fetching from the instruction cache and decoding those instructions in the decode unit when the decoded micro-ops are not present in the operation cache. For a variety of reasons (some of which are described elsewhere herein), in some implementations, switching between fetching from the operation cache and decoding from the instruction cache may result in the loss of several cycles. For this reason, techniques are provided herein to reduce the delay associated with switching between fetching decoded micro-ops from the operation cache and decoding instructions to generate micro-ops for instructions that do not have decoded micro-ops in the operation cache. Although Figure 2 One example of a processor pipeline 200 in which these techniques are applied is shown, but those skilled in the art will appreciate that the teachings of the present disclosure are also applicable to other pipeline architectures.
[0025] Figure 3 3 is a block diagram illustrating components of the instruction fetch and decode unit 202 according to an example. The instruction fetch and decode unit 202 includes a branch predictor 302, a prediction queue 304, a shared fetch logic 305, an operation cache path 301, an instruction cache path 303, and an operation queue 328. The shared fetch logic 305 includes an operation cache tag lookup circuit 310, an operation cache tag array 306, an instruction cache tag lookup circuit 312, an instruction cache tag array 308, and a fetch control circuit 313. The operation cache path 301 includes an operation cache queue 314, an operation cache data read circuit 320, and an operation cache data array 322. The instruction cache path 303 includes an instruction cache queue 315, an instruction cache data read circuit 316, an instruction cache data array 318 communicatively coupled to a higher level cache and system memory, an instruction byte buffer 324, and a decoder 326. As used herein, the term “operation cache” refers to the combination of the operation cache tag lookup circuit 310, the operation cache tag array 306, the operation cache data read circuit 320, and the operation cache data array 322. As used herein, the term “instruction cache” refers to the combination of the instruction cache tag lookup circuit 312, the instruction cache tag array 308, the instruction cache data read circuit 316, and the instruction cache data array 318.
[0026] Any cache, such as an operational cache and an instruction cache, includes a tag array that allows determination of whether a particular tag and index hits the cache and a data array that stores the cached data. Thus, the operational cache includes an operational cache tag array 306 and an operational cache data array 322, and the instruction cache includes an instruction cache tag array 308 and an instruction cache data array 318. In various examples, Figure 3 Any or all of the elements shown are implemented as fixed-function hardware (ie, as fixed-function circuits).
[0027] The branch predictor 302 generates predicted addresses for use by the rest of the instruction fetch and decode unit 202. By known techniques, the branch predictor 302 attempts to identify a sequence of instructions to be executed by the software executed in the instruction execution pipeline 200, the instruction sequence being specified as a sequence of predicted instruction addresses. The instruction sequence identification includes branch prediction, which uses various execution state information (such as the current instruction pointer address and branch prediction history, in various examples, the branch prediction history includes data indicating the history of whether a particular branch is taken and / or other data). The branch predictor 302 may predict incorrectly, in which case the branch predictor 302 and other parts of the instruction execution pipeline 200 perform actions to remedy the prediction error, such as undoing the operation of instructions from the wrongly predicted branch path, and the branch predictor 302 changes the instruction address to the correct branch path. There are a variety of branch prediction techniques, and in various examples, the branch predictor 302 uses any technically feasible branch prediction technique to identify a sequence of predicted instruction addresses.
[0028] The prediction queue 304 stores predicted addresses from the branch predictor 302. The prediction queue 304 acts as a decoupled queue between the branch predictor 302 and the rest of the instruction fetch and decode unit 202, allowing the timing of operation of the branch predictor 302 to be independent of the timing of operation of the rest of the instruction fetch and decode unit 202. The format of the predicted addresses output by the branch predictor 302 and stored in the prediction queue 304 is in the form of a predicted address block, which is a group of instruction addresses defined by a range.
[0029] The shared fetch logic 305 operates on predicted address blocks, sending addresses from such predicted address blocks to the operation cache way 301, the instruction cache way 303, or both, based on whether the addresses corresponding to the predicted address blocks are determined to be serviceable by the operation cache way 301 due to the corresponding translations being stored in the operation cache. More specifically, the shared fetch logic 305 includes an operation cache tag lookup 310 that probes the operation cache tag array 306 to determine whether any addresses of a particular predicted address block are stored in the operation cache. For a particular predicted address block, based on the lookup, the fetch control circuit 313 generates one or both of an operation cache queue entry or an instruction cache queue entry for servicing by the operation cache way 301 and / or the instruction cache way 303. If at least one operation cache queue entry is generated for a particular predicted address block, the operation cache way 301 services the operation cache queue entry, retrieves cached decoded operations from the operation cache, and forwards such cached decoded operations to the operation queue 328 for processing by the rest of the instruction execution pipeline 200. For predicted address blocks that the operation cache way 301 cannot fully service (i.e., some but not all or no addresses of the predicted address block have corresponding decoded instructions stored in the operation cache) the instruction cache way 303 partially or fully services the predicted address block by extracting instruction bytes from the instruction cache and then decoding those instruction bytes. The operation cache queue entries and instruction cache queue entries reflect the work to be performed by the operation cache way 301 and / or the instruction cache way 303 for a particular predicted address block.
[0030] The cache queue entries output by the shared fetch logic 305 include an indication that allows the operation cache path 301 and the instruction cache path 303 to coordinate with respect to the relative order of the addresses to be served. More specifically, for the cache queue entries flowing to each of the operation cache path 301 and the instruction cache path 303, such coordination to follow the relative program order is facilitated by including an indication of whether the reverse path (i.e., the instruction cache path 303 for the operation cache path 301 or the operation cache path 301 for the instruction cache path 303) obtains the decoded micro-operation of the instruction immediately preceding the instruction address of the particular cache queue entry. If such an indication exists, the operation cache path 301 or the instruction cache path 303 waits for the reverse path to serve such immediately preceding instruction address before serving the instruction address for which there is a "wait" indication. By waiting in this manner, the decoded micro-operations are output to the operation queue 328 in program order. The operation queue 328 acts as an endpoint or final stage of the instruction fetch and decode unit 202. More specifically, the operation queue 328 stores the program order output of the operation cache path 301 and the instruction cache path 303 for servicing by the rest of the instruction execution pipeline 200. The operation queue 328 also acts as a decoupling buffer that decouples the operation timing of the instruction fetch and decode unit 202 from the operation timing of subsequent stages of the instruction execution pipeline 200.
[0031] Now, with respect to the various components of those paths shown and including additional details, and also with reference to FIG. 4A to FIG. 4C , the operation of the shared fetch logic 305, the operation cache path 301 and the instruction cache path 303 are described in more detail. Specifically, Figure 4A shows the operation of the shared extraction logic 305, Figure 4B The operation of operating cache path 301 is shown, and Figure 4C The operation of instruction cache path 303 is shown.
[0032] Figure 4A is a flow chart of a method 400 of operation of the shared extraction logic 305 according to an example. Figures 1 to 3 However, those skilled in the art will appreciate that the steps may be performed in any technically feasible order. Figure 4A Any system that implements the steps of FIG. 1 is within the scope of the present disclosure.
[0033] The method 400 begins at step 402, where the operation cache tag lookup 310 and the instruction cache tag lookup 312 retrieve a predicted address block from the prediction queue 304. At step 404, the operation cache tag lookup circuit 310 and the instruction cache tag lookup circuit 312 consume the predicted address block to determine whether the operation cache stores cached decoded operations for all, part, or none of the predicted address block. While any technically feasible technique for doing so is possible, an example of a detailed lookup operation for finding a predicted address block is now provided.
[0034] According to the described example, the operational cache tag lookup circuit applies the address representing the predicted address block to the operational cache tag array 306 to determine whether there are any hits for the predicted address block in the operational cache. The operational cache is a set associative cache in which each index is associated with a particular set and each tag is associated with one or more ways. The tag corresponds to the high-order bits of the address (such as bits [47:12] of the address), and the index corresponds to the bits of the next low-order bits of the address (such as bits [11:6] of the address). The tag may be a partial tag, which is a version of the full tag shortened by some technique (such as hashing). Throughout this disclosure, it should be understood that the term "tag" may be replaced by "partial tag" where appropriate.
[0035] The operational cache tag lookup 310 applies the address representing the predicted address block to the operational cache tag array 306 as follows. The operational cache tag lookup 310 applies the index derived from the predicted address block to the operational cache tag array 306, which outputs the tags in the set identified by the index. If the tag derived from the predicted address block matches one or more of the read tags, a hit occurs.
[0036] Each hit of the combination of index and tag indicates that there is an entry in the operation cache that can store translated micro-operations for one or more instructions in the predicted address block. Since the operation cache stores entries corresponding to one or more individual instructions that may not completely cover the instruction cache line, information other than just the tag and index is required to indicate which specific instructions correspond to a specific hit among all instructions in the address range corresponding to the predicted address block. The identification information is stored as a starting offset value in each entry of the operation cache tag array 306. More specifically, a typical instruction cache stores instructions at the granularity of a cache line. In other words, a cache hit indicates that a specific amount of data aligned with the address portion represented by the index and tag is stored in the cache. However, the operation cache stores entries (decoded micro-operations) for one or more individual instructions that cover an address range smaller than a cache line. Therefore, a hit for a specific instruction in the operation cache requires a match of the index and tag, which represents the high-order bits of the address (e.g., bits [47:6]), and a match of the starting offset with the low-order bits (e.g., bits [5:0]) representing the address, which is aligned down to the byte level (for some instruction set architectures, instructions can only exist at a granularity of 2 or 4 bytes, so bits [0] or bits [1:0] of the offset can be excluded).
[0037] As described above, the starting offset information is stored in all entries in the operation cache tag array 306. In order to identify the instruction address range covered by the cached decoded micro-operations stored in the operation cache, each entry in the operation cache tag array 306 stores information indicating the ending offset of the instruction corresponding to the entry, because the size of the operation cache entry can be variable. The ending offset is stored as the address offset of the next instruction after the last instruction corresponding to the entry. The purpose of the information indicating the next instruction offset is to allow the hit status of the next operation cache entry to be found in the same cycle if the next instruction has decoded the micro-operations stored in the operation cache. By comparing the ending offset of the first entry with the starting offsets of other tag array entries with matching tags at the same index, if the second entry exists, the second entry is identified. Similarly, the ending offset of the second entry is compared with the starting offsets of other tag array entries with matching tags at the same index to find the third entry (if any).
[0038] In summary, the offset information allows identification of a specific address range in the predicted address block that is covered by an operation cache entry hit (via a match with a tag when an index is applied), and allows the operation cache tag lookup 310 to identify the address of the next instruction that is not covered by the entry. When the index is applied to the operation cache tag array 306, multiple entries are read out. Each entry includes a tag, a starting offset, and an ending offset (also referred to as the next instruction offset). A match between both the tag and starting offset of the operation cache entry and the tag and starting offset of the predicted address block signals an operation cache hit for the first entry. The ending offset of the first entry is compared to the starting offsets of other entries that also match the tag of the predicted address block to identify the second entry. As described elsewhere herein, the ending offset of the second entry is used to link together sequence hit detection for multiple entries in a single cycle. For each of the identified entries, the ending offset is compared to the ending offset of the predicted address block. When the end offset of a predicted address block is matched or exceeded, the block has been completely overwritten and further link entries are ignored for processing of the predicted address block.
[0039] At step 406, the operational cache tag lookup circuit 310 determines whether there is at least one hit in the operational cache tag array 306. If there is at least one hit, and the at least one hit covers all instructions for all predicted address blocks, the method 400 proceeds to step 408. If there is at least one hit, and the at least one hit covers some but not all instructions for the predicted address blocks, the method 400 proceeds to step 410. If there is no hit in the operational cache tag array 306, the method proceeds to step 414.
[0040] At step 408, the fetch control circuit 313 generates and writes an operation cache queue entry having information about hits for the predicted address block in the operation cache. The operation cache queue entry includes information indicating the index and entry in the operation cache for each hit, and an indication of whether a path change is required to serve the address of the operation cache queue entry. "Entry" is stored because the set (corresponding to the index) and entry identify a unique entry in the operation cache. The way is unique for each combination of tag and offset, and uniquely identifies a single entry in the set identified by the index. A path change occurs when an instruction prior to the instruction served by the operation cache queue entry is served by the instruction cache path 303. Therefore, the indication is set if the previously predicted address block is served by the instruction cache path 303 (or more specifically, if the last instruction of the previously predicted address block is served by the instruction cache path 303). As described in reference Figure 4B As depicted, the operation cache queue entries are serviced by operation cache way 301 .
[0041] Returning to step 406, if the hits in the operation cache indicate that some but not all of the predicted address blocks are covered by the operation cache, the method 400 proceeds to step 410. At step 410, the fetch control circuit 313 generates operation cache queue entries for the operation cache way 301 and writes those operation cache queue entries to the operation cache queue 314. The operation cache queue 314 acts as a decoupling buffer, isolating the timing in which the shared fetch logic 305 processes the predicted address blocks and creates the operation cache queue entries from the timing in which the operation cache data read circuit 320 consumes the operation cache queue entries, reads the operation cache data array 322, and outputs the decoded operations to the operation queue 328 in program order. These operation cache queue entries include the index and entry in the operation cache for each hit, as well as a path change indication indicating whether there is a path change from the immediately previous instruction. More specifically, the path change indication indicates whether the instruction immediately preceding the instruction address served by the particular cache queue entry will be served by the instruction cache way 303. In addition, at step 412, the extraction control circuit 313 generates an instruction cache queue entry for storage in the instruction cache queue 315. The instruction cache queue entry includes the index and entry of the instruction cache in which the hit occurred, the starting offset and the ending offset extracted from the instruction cache, and an indication that a path change occurred for the first such entry. The starting offset written to the instruction cache queue is the ending offset of the last operation cache entry that hit for the predicted address block. The ending offset written to the instruction cache queue is the ending offset of the predicted address block. The path change indication is set for the instruction cache queue entry because the previous instruction with the same predicted address block was served by the operation cache path 301.
[0042] Returning again to step 406, if there is no hit for the predicted address block in the operational cache, the method 400 proceeds to step 414. At step 414, the extraction control circuit 313 generates an instruction cache queue entry for storage in the instruction cache queue 315. The instruction cache queue entry includes the index and entry of the instruction cache in which the hit occurred, the starting offset and ending offset extracted from the instruction cache, and an indication of whether a path change occurred for the first such entry. The starting offset written to the instruction cache queue is the starting offset of the predicted address block. The ending offset written to the instruction cache queue is the ending offset of the predicted address block. If the instruction immediately preceding the instruction of the first entry will be served by the operational cache path 301, the indication is set.
[0043] After step 412 or 414, method 400 proceeds to step 416. At step 416, instruction cache data read circuit 316 reads instruction cache data array 318 based on the instruction cache queue entry of the instruction cache queue. Specifically, instruction cache data read circuit 316 accesses the entry specified by the index and entry for the particular instruction cache queue entry and obtains the instruction byte from the instruction cache. Some instructions may not be stored in the instruction cache or the operation cache, in which case the instruction byte must be extracted from a higher level cache or system memory. Instruction cache path 303 will also serve such cases. At step 418, instruction cache data read circuit 316 writes the instruction byte read from the instruction cache into instruction byte buffer 324 for decoding into decoded micro-operations by decoder 326.
[0044] In an implementation, the determining and finding steps (steps 404 and 406) occur in a pipelined manner (and thus at least partially overlapped in time) over several cycles relative to the steps of generating an operation cache queue entry and an instruction cache queue entry (steps 408, 410, 412, and 414). An example of such operation is now provided, wherein the checking of the tag array and the generation of the cache queue entry are performed in a pipelined manner, and at least partially overlapped in time is now provided.
[0045] According to the example, for any particular predicted address block in which a hit occurs in the operating cache, the operating cache tag lookup circuit 310 identifies a first offset to use as the lowest instruction address offset in which a hit occurs. In any particular cycle, identifying the first offset depends on several factors, including whether the current cycle is the first cycle in which the operating cache tag lookup 310 is checking the current predicted address block. More specifically, the operating cache tag lookup 310 is able to identify multiple hits in the operating cache tag array 306 in a single cycle. However, there is a limit to the number of hits that the operating cache tag lookup 310 can check in a cycle. If the operating cache tag lookup 310 checks a particular predicted address block in the second or higher cycle, the first offset for the cycle is specified by the end offset specified in the last hit of the previous cycle (in other words, the last hit of the previous cycle provides the first offset for the cycle). Otherwise (i.e., if the current cycle is the first cycle in which the operating cache tag lookup 310 is checking the predicted address block), a different technique is used to identify the first offset to be used from the operating cache tag array 306.
[0046] More specifically, if the current cycle is the first cycle in which the cache tag lookup 310 is checking the currently predicted address block, then one of the following is true:
[0047] there is a hit in the operation cache of the previously predicted address block, which indicates that the next instruction for which there is a hit in the operation cache belongs to the next predicted address block relative to the previously predicted address block (i.e., the entry in the operation cache tag array 306 includes an indication that the entry spans into the sequence predicted address block. When the "span" indication is set, the end offset refers to the offset into the sequence predicted address block),
[0048] the last instruction of the previously predicted address block is serviced by the instruction cache way 303 (ie, no decoded operation is stored in the operation cache for the last instruction); or
[0049] - The currently predicted address block represents the target of the taken branch.
[0050] In the event that there is an operation cache entry hit in the previously predicted address block, indicating that the operation cache entry spans into the sequence predicted address block, a first offset is provided for the currently predicted address block from the end offset of the entry.
[0051] In the case where the last instruction of the previously predicted address block is served by instruction cache way 303, the first offset used is the starting offset of the predicted address block. In an implementation, the op cache tag will not be looked up in this case, and instruction cache way 303 will be used. However, other implementations that do not allow "stepping" of op cache entries may choose to loop over op cache tag hits, and use op cache way 301 in the event of a hit.
[0052] In the case where the current predicted address block represents the target of a taken branch, the first offset used is the starting offset of the predicted address block.
[0053] When the index is applied to the operation cache tag array 306, multiple entries are read out and used in the operation cache tag lookup 310. Each entry includes a tag, a starting offset, and an ending offset (also called the next instruction offset). A match between both the tag and the starting offset of the operation cache entry and the tag of the predicted address block and the first offset just described signals an operation cache hit for the first entry. The ending offset of the first entry is compared to the starting offsets of other entries that also match the tags of the predicted address block to identify the second entry. As described elsewhere herein, the ending offset of the second entry is also used to link together sequence hit detections for multiple entries in a single cycle. The operation cache tag lookup 310 repeats the operation in the same cycle until one or more of the following occurs: the maximum number of sequential operation cache entries for a single cycle is reached; or the most recently checked operation cache tag array entry indicates that the next offset exceeds the ending offset of the predicted address block.
[0054] The operating cache has several properties that allow multiple hits to occur in a cycle. More specifically, the fact that the operating cache is set associative, combined with the fact that all hits in a single predicted address block fall into the same set, allows multiple hits to occur in a single cycle. Due to the described properties, when the index derived from the current predicted address block is applied to the operating cache tag array 306, all entries that may belong to the predicted address block are found in the same set. In response to the index being applied to the operating cache tag array 306, all entries of the operating cache tag array 306 are read out (due to the nature of the set associative cache). The operating cache tag lookup 310 is therefore able to obtain the first entry based on the first entry offset described above. The op-cache tag lookup 310 can then obtain the next entry by using the next instruction address of the first entry to match the offset of one of the entries already read, and continue to sort the entries of the op-cache tag array 306 that have been read in that cycle in this manner until the number of entries that can be sorted in one cycle is reached, or until there are no more sequence entries to read (e.g., the op-cache does not store decoded micro-ops for the next instruction following the instruction that hit in the op-cache).
[0055] In summary, the fact that all entries of any particular predicted address block fall into a single set means that when an index is applied to the operational cache tag array 306, the offsets and next instruction addresses of all entries that can be matched in the predicted address block are read out. This fact allows the use of simple sequential logic to link multiple entries of the operational cache tag array 306 during a clock cycle. If entries for a single predicted address block could be found in multiple sets, linking through multiple entries would be difficult or impossible because the index would need to be applied to the operational cache tag array 306 multiple times, which would take longer. Therefore, by reading multiple entries of the operational cache tag array 306, the operating speed of the operational cache path 301 is improved, and the compactness of the operational cache queue entries is also improved by allowing information of multiple hits in the operational cache to be written to the operational cache queue 314 in a single cycle.
[0056] Although any technically feasible means can be used to ensure that all entries of any particular predicted address block fall within the same set of the operational cache, a specific technique is now described. According to the technique, the branch predictor forms all predicted address blocks so that the addresses are within a specific aligned block size (such as a 64-byte aligned block). The predicted address blocks do not have to start or end on an aligned boundary, but they are not allowed to cross aligned boundaries. If an implementation of the branch predictor 302 that does not naturally comply with the restrictions is used, additional combinatorial logic can be inserted between the branch predictor 302 and the shared extraction logic 305 to decompose any predicted address blocks that cross the alignment boundary into multiple blocks for processing by the shared extraction logic 305. The alignment means that all legal starting offsets within the predicted address block differ only in the low-order bits. The index used to find the tag in the operational cache tag array 306 (which defines the set) does not have these lowest-order bits. Due to the absence of these lowest-order bits, the index cannot change for different addresses in a single predicted address block, and therefore the set must be the same for any specific operational cache entry required to satisfy the predicted address block.
[0057] The just described operation of reading multiple entries from the operational cache tag array 306 using the next instruction offset in sequence may be referred to herein as “reading sequence tags from the operational cache tag array 306 ” or via similar wording.
[0058] Once the operation cache tag lookup 310 has determined that the sequence tag read from the operation cache tag array 306 is complete for the current cycle, the operation cache tag lookup 310 generates an operation cache queue entry for storage in the operation cache queue 314. The operation cache queue entry includes information indicating the index and entry in the operation cache for each hit indicated by the operation cache tag array 306 for the current cycle, so that the micro-operation can be read out from the operation cache data array 322 later. In addition, each operation cache queue entry includes an indication of whether the instruction immediately preceding the first hit represented by the operation cache queue entry will be serviced by the instruction cache way 303 (i.e., because there is no corresponding set of cached micro-operations in the operation cache). The described indication facilitates storing the micro-operations in the operation queue 328 in program order (i.e., the order indicated by the sequence of instructions executed for the executing software), which will be described in more detail elsewhere herein.
[0059] The fetch control logic 313 determines whether a particular operation cache queue entry includes an indication of whether the instruction immediately preceding the first hit is to be serviced by the cache way 303 as follows. If the operation cache queue entry corresponds to the second or later cycle that the particular predicted address block is checked by the operation cache tag lookup 310, then the instruction immediately preceding the operation cache queue entry is not serviced by the instruction cache way 303 because the instruction is part of a multi-cycle read for a single predicted address block. Therefore, in the previous cycle, the operation cache tag lookup 310 determined that the operation cache stores decoded micro-ops for the immediately preceding instruction.
[0060] If the operation cache queue entry corresponds to the first cycle of checking a particular predicted address block by the operation cache tag lookup 310, a determination is made based on whether the previously predicted address block is completely covered by the operation cache. In the event that the previously predicted address block is completely covered by the operation cache (also shown in the "yes, full PAB covered" arc branching from decision block 406 of method 400), the instruction immediately preceding the operation cache queue entry is not serviced by the instruction cache way 303. If the previously predicted address block is not completely covered by the operation cache, the instruction immediately preceding the operation cache queue entry is not serviced by the instruction cache way 303.
[0061] Turning now to the instruction cache side of the shared fetch logic 305, the instruction cache tag lookup circuit 312 consumes and processes the prediction queue entries as follows. The instruction cache tag lookup 312 examines the predicted address block, obtains an index of the predicted address block, and applies the index to the instruction cache tag array 308 to identify a hit. The instruction cache tag lookup 312 provides information indicating the hit to the fetch control circuit 313. In some implementations, it is possible that the size of the predicted address block may be larger than the number of addresses that the instruction cache tag lookup 312 can look up in one cycle. In this case, the instruction cache tag lookup 312 sorts the different address range portions of the predicted address block in program order and identifies hits for each of those address range portions. Based on these hits, the fetch control circuit 313 generates instruction cache queue entries for storage in the instruction cache queue 315. These instruction cache queue entries include information indicating a hit in the instruction cache (such as index and entry information), as well as information indicating whether a path change has occurred for a particular instruction cache queue entry. Since the selection between employing the operational cache path 301 and the instruction cache path 303 is independent of the instruction cache hit status, those skilled in the art will appreciate that the instruction cache tag lookup may be accomplished in the shared fetch logic 305 or the instruction cache path 303 based on various tradeoffs not directly related to the techniques described herein.
[0062] Figure 4B is a flow chart of a method 430 for processing entries of the operation cache queue 314 to extract cached decoded micro-operations from the operation cache according to an example. Figures 1 to 3 However, those skilled in the art will appreciate that the steps may be performed in any technically feasible order. Figure 4B Any system that implements the steps of FIG. 1 is within the scope of the present disclosure.
[0063] At step 432, the operation cache data read circuitry 320 determines whether the operation cache queue 314 is empty. If it is empty, the method 430 performs step 432 again. If it is not empty, the method 430 proceeds to step 434. At step 434, the operation cache data read circuitry 320 determines whether the path change indication of the "head" (or "next") cache queue entry is set. If the path change indication is set (not cleared), the operation cache queue entry needs to wait for the instruction path before being processed, so the method 430 proceeds to step 436. If the path change indication is not set (cleared), no waiting occurs, and the method 430 proceeds to step 438. At step 436, the operation cache data read circuitry 320 determines whether all micro-ops preceding the micro-op represented by the operation cache queue entry have been written to the operation queue 328, or are currently being transmitted in the process of being decoded and written to the operation queue 328. If all micro-ops preceding those micro-ops represented by the operation cache queue entry have been written to the operation queue 328 or are currently being decoded and written to the operation queue 328, the method 430 proceeds to step 438, and if not all micro-ops preceding those micro-ops represented by the operation cache queue entry have been written to the operation queue 328 or are currently being decoded and written to the operation queue 328, the method returns to step 436.
[0064] At step 438, the operation cache data read circuit 320 reads the operation cache data array 322 to obtain the cached decoded micro-ops based on the contents of the operation cache queue entry. At step 440, the operation cache data read circuit 320 writes the read micro-ops to the operation queue 328 in program order.
[0065] A detailed example of steps 438 and 440 is now provided. According to the example, the operation cache data read circuit 320 obtains an index and a tag from an operation cache queue entry and applies the index and the tag to the operation cache data array 322 to obtain a decoded micro-operation for the operation cache queue entry. Since the operation cache queue entry can store multiple hit data, the operation cache data read circuit 320 performs the above-mentioned lookup one or more times depending on the number of hits represented for each operation cache queue entry. In some implementations, the operation cache data read circuit 320 performs multiple lookups from the same operation cache queue 314 entry in a single cycle, while in other implementations, the operation cache data read circuit 320 performs one lookup per cycle.
[0066] Figure 4Cis a flow chart of a method 460 for fetching and decoding instruction bytes stored in the instruction byte buffer 324 according to an example. Figures 1 to 3 However, those skilled in the art will appreciate that the steps may be performed in any technically feasible order. Figure 4C Any system that implements the steps of FIG. 1 is within the scope of the present disclosure.
[0067] At step 462, the instruction byte buffer 324 determines whether the instruction byte buffer 324 is empty. If the instruction byte buffer is empty, the method 460 returns to step 462, and if the instruction byte buffer is not empty, the method 460 proceeds to step 464. At step 464, the instruction byte buffer 324 determines whether the path change indication of the "head" (or "next") entry in the instruction byte buffer 324 is clear (indicating that a path change is not required for the entry). If the indication is not clear (i.e., the indication is set), the method 460 proceeds to step 466, and if the indication is clear (i.e., the indication is not set), the method 460 proceeds to step 468.
[0068] At step 466, the instruction byte buffer 324 checks whether all micro-operations of the instruction before the next entry in the program order served by the operation cache path 301 are written to the operation queue 328 or are written to the operation queue 328 when being transferred. If all micro-operations of the instruction before the next entry in the program order served by the operation cache path 301 are written to the operation queue 328 or are written to the operation queue 328 when being transferred, the method proceeds to step 468, and if not all micro-operations of the instruction before the next entry in the program order served by the operation cache path 301 are written to the operation queue 328 or are written to the operation queue 328 when being transferred, the method returns to step 466.
[0069] At step 468, the instruction byte buffer 324 reads the header entry and sends the instruction bytes of the entry to the decoder 326 for decoding. At step 470, the decoder 326 decodes the instruction bytes according to known techniques to generate decoded micro-operations. At step 472, the decoder 326 writes the decoded micro-operations to the operation queue 328 in program order.
[0070] The operation queue 328 provides the decoded micro-operations to the reorder buffer 210 and / or other subsequent stages of the instruction execution pipeline 200, as those stages are capable of consuming those decoded micro-operations. The rest of the instruction execution pipeline 200 consumes and executes those micro-operations according to known techniques to run the software represented by the instructions that obtained the micro-operations. As is well known, such software is capable of generating any technically feasible results and performing any technically feasible operations, such as writing the result data to memory, interfacing with input / output devices, and performing other operations as needed.
[0071] Once decoded, the instruction cache uses the decoded operations to update the operation cache based on any technically feasible replacement policy. The purpose of these updates is to keep the operation cache current with the program execution state (e.g., to ensure that recently decoded operations are available in the operation cache for use in decoding instructions in subsequent program flow).
[0072] It will be appreciated that instruction fetch and decode unit 202 is a pipelined unit, meaning that work in one stage (e.g., branch predictor) can be performed for certain instruction addresses in the same cycle as work in a different stage (e.g., op cache tag lookup 310). It will also be appreciated that op cache path 301 and instruction cache path 303 are independent parallel units that, although they are synchronized to output decoded micro-ops in program order, can perform work for the same or different predicted address blocks in the same cycle.
[0073] A variation of the above is that, instead of using the instruction cache queue 315, the instruction cache path 303 may directly read the prediction block to determine which address ranges to obtain the instruction bytes from. To facilitate this, the fetch control circuit 313 may write information indicating which addresses to fetch to the prediction queue 304 as the head entry of the queue, and the instruction cache data read circuit 316 may identify the next address to be used based on the head entry in the prediction queue 304.
[0074] The techniques described herein provide an instruction fetch and decode unit having an operation cache with low latency when switching between fetching decoded operations from the operation cache and fetching and decoding instructions using the decode unit. The low latency is achieved through a synchronization mechanism that allows work to flow through both the operation cache path and the instruction cache path until the work is stopped due to the need to wait for output from the opposite path. The presence of decoupling buffers in the operation cache path 301 (i.e., the operation cache queue 314) and the instruction cache path 303 (i.e., the instruction byte buffer 324) allows work to be maintained until it is cleared to continue due to synchronization between the two paths. Other improvements such as a specially configured operation cache tag array that allows multiple hits to be detected in a single cycle improve bandwidth by, for example, increasing the speed at which entries are consumed from the prediction queue 304, and allowing the ability to have an implementation that reads multiple entries from the operation cache data array 322 in a single cycle. More specifically, since prediction queue 304 advances to the next entry after both operation cache way 301 and instruction cache way 303 have read in the current entry (e.g., via operation cache tag lookup 310 and instruction cache tag lookup 312), allowing operation cache tag lookup 310 to service multiple instructions per cycle speeds up the rate at which prediction queue entries are consumed.
[0075] A method is provided for converting instruction addresses of a first predicted address block into decoded micro-operations for output to an operation queue storing the decoded micro-operations in program order and subsequently executed by the remainder of an instruction execution pipeline. The method comprises: identifying that the first predicted address block includes at least one instruction for which decoded micro-operations are stored in an operation cache of an operation cache path; storing a first operation cache queue entry in the operation cache queue, the first operation cache queue entry including an indication indicating whether to wait to receive a signal from the instruction cache path to proceed; obtaining the decoded micro-operations corresponding to the first operation cache queue entry from the operation cache; and outputting the decoded micro-operations corresponding to the first operation cache queue entry to the operation queue at a time based on the indication of the first operation cache queue entry indicating whether to wait to receive a signal from the instruction cache path to proceed.
[0076] An instruction fetch and decode unit is used to convert instruction addresses of a first predicted address block into decoded micro-operations for output to an operation queue storing the decoded micro-operations in program order and subsequently executed by the rest of the instruction execution pipeline. The instruction fetch and decode unit includes: shared fetch logic configured to recognize that the first predicted address block includes at least one instruction for which decoded micro-operations are stored in an operation cache of an operation cache path; an operation cache queue storing a first operation cache queue entry, the first operation cache queue entry including an indication indicating whether to wait to receive a signal from the instruction cache path to proceed; and operation cache data read logic, which obtains the decoded micro-operation corresponding to the first operation cache queue entry from the operation cache and outputs the decoded micro-operation corresponding to the first operation cache queue entry to the operation queue at a time based on the indication of the first operation cache queue entry indicating whether to wait to receive a signal from the instruction cache path to proceed.
[0077] The processor includes an instruction fetch and decode unit for converting instruction addresses of a first predicted address block into decoded micro-operations for output to an operation queue storing the decoded micro-operations in program order and subsequently executed by the remainder of the instruction execution pipeline and the remainder of the instruction execution pipeline. The instruction fetch and decode unit includes: shared fetch logic that identifies that the first predicted address block includes at least one instruction for which decoded micro-operations are stored in an operation cache of an operation cache path; an operation cache queue that stores a first operation cache queue entry, the first operation cache queue entry including an indication indicating whether to wait to receive a signal from the instruction cache path to proceed; and operation cache data read logic that obtains the decoded micro-operation corresponding to the first operation cache queue entry from the operation cache and outputs the decoded micro-operation corresponding to the first operation cache queue entry to the operation queue at a time based on the indication of the first operation cache queue entry indicating whether to wait to receive a signal from the instruction cache path to proceed.
[0078] It should be understood that many variations are possible based on the disclosure herein. Although features and elements are described above in particular combinations, each feature or element may be used alone without the other features and elements, or in various combinations with or without the other features and elements.
[0079] The provided methods may be implemented in a general purpose computer, a processor, or a processor core. Suitable processors include, by way of example, a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and / or a state machine. Such processors may be manufactured by configuring a manufacturing process using the result of processed hardware description language (HDL) instructions and other intermediate data including a netlist (such instructions can be stored on a computer readable medium). The result of such processing may be a mask work, which is then used in a semiconductor manufacturing process to manufacture a processor that implements the features disclosed above.
[0080] The methods or flow charts provided herein may be implemented in a computer program, software, or firmware that is incorporated into a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs).
Claims
1. A method for converting instruction addresses of a first predicted address block into decoded micro-operations for output to an operation queue storing decoded micro-operations in program order and subsequently executed by the remainder of an instruction execution pipeline, the method comprising: providing an index associated with the first predicted address block to an operating cache tag array to obtain a first tag associated with a first operating cache tag array entry and a second tag associated with a second operating cache tag array entry; determining, in a first computer clock cycle, that both the first tag and the second tag match tags derived from the predicted block of addresses and that an end address associated with the first operating cache tag array entry matches a start address associated with the second operating cache tag array entry; in the first computer clock cycle, identifying a set of instructions for which decoded micro-operations are stored in an operation cache data array of an operation cache way as instructions at addresses associated with both the first operation cache tag array entry and the second operation cache tag array entry; storing a first operation cache queue entry in an operation cache queue for the set of instructions, the first operation cache queue entry including an indication indicating whether to wait for receiving a signal from an instruction cache way to proceed; obtaining, from the operation cache, a decoded micro-operation corresponding to the first operation cache queue entry; as well as The decoded micro-op corresponding to the first operation cache queue entry is output to the operation queue at a time based on the indication of the first operation cache queue entry indicating whether to wait to receive the signal from the instruction cache way to proceed.
2. The method of claim 1, wherein: The indication indicates a need to wait for a previous micro-operation from the instruction cache path to be written or transferred to the operation queue; as well as Outputting the decoded micro-operations to the operation queue includes waiting until the previous micro-operations from the instruction cache way are written or transferred to the operation queue before outputting the decoded micro-operations corresponding to the first operation cache queue to the operation queue.
3. The method of claim 1, wherein: The indication does not indicate a need to wait for a previous micro-operation from the instruction cache path to be written or transferred to the operation queue; and Outputting the decoded micro-op to an operation queue without waiting for a previous micro-op from the instruction cache way includes outputting the decoded micro-op corresponding to the first operation cache queue entry.
4. The method of claim 1, further comprising: The first operating cache queue entry is generated by chaining multiple hits of the first predicted address block in an operating cache tag array in a single cycle and including information for each such hit in the first operating cache queue entry.
5. The method of claim 1, further comprising: The instruction addresses of the second predicted address block are converted into decoded micro-operations for output to the operation queue and subsequent execution by the rest of the instruction execution pipeline by the following operations: identifying that the second predicted address block includes at least one instruction for which decoded micro-operations are not stored in the operation cache; storing the instruction cache queue entry in the instruction cache queue; obtaining an instruction byte in an instruction byte buffer for the instruction cache queue entry along with an indication of whether to wait for a previous operation from the operation cache to be written or transferred to the operation cache; decoding the instruction byte to obtain a decoded micro-operation corresponding to the instruction byte buffer entry at a time based on the indication indicating whether to wait for the previous operation from the operation cache way; as well as The decoded micro-operation corresponding to the instruction byte buffer entry is output to the operation queue for storage.
6. The method of claim 5, wherein: The indication indicating whether to wait for the previous operation from the operation cache path indicates that waiting for the previous operation from the operation cache path is required; as well as Decoding the instruction byte includes decoding the instruction byte after the previous operation from the operation cache way is written or transferred to the operation queue.
7. The method of claim 5, wherein: The indication indicating whether to wait for the previous operation from the operation cache path indicates that it is not necessary to wait for the previous operation from the operation cache path; as well as Decoding the instruction byte includes decoding the instruction byte without waiting for a previous operation from the operation cache way to be written or transferred to the operation queue.
8. The method of claim 1, wherein: The first predicted address block includes at least one instruction whose decoded micro-operation is stored in an operation cache and at least one instruction whose decoded operation is not stored in the operation cache.
9. The method of claim 1, further comprising: The micro-operations stored in the operation queue are executed.
10. An instruction fetch and decode unit for converting instruction addresses of a first predicted address block into decoded micro-operations for output to an operation queue storing decoded micro-operations in program order and subsequently executed by the remainder of an instruction execution pipeline, the instruction fetch and decode unit comprising: Shared extraction logic, which is configured to: providing an index associated with the first predicted address block to an operating cache tag array to obtain a first tag associated with a first operating cache tag array entry and a second tag associated with a second operating cache tag array entry, determining, in a first computer clock cycle, that both the first tag and the second tag match tags derived from the predicted block of addresses and that an end address associated with the first operating cache tag array entry matches a start address associated with the second operating cache tag array entry, and in the first computer clock cycle, identifying a set of instructions for which decoded micro-operations are stored in an operation cache data array of an operation cache way as instructions at addresses associated with both the first operation cache tag array entry and the second operation cache tag array entry; an operation cache queue configured to store, for the set of instructions, a first operation cache queue entry, the first operation cache queue entry including an indication of whether to wait for receiving a signal from an instruction cache way to proceed; and an operation cache data read logic configured to obtain the decoded micro-operation corresponding to the first operation cache queue entry from the operation cache at a time based on the indication of the first operation cache queue entry indicating whether to wait to receive the signal from the instruction cache way to proceed, and output the decoded micro-operation corresponding to the first operation cache queue entry to the operation queue.
11. The instruction fetch and decode unit of claim 10, wherein: The indication indicates a need to wait for a previous micro-operation from the instruction cache path to be written or transferred to the operation queue; as well as Outputting the decoded micro-operations to the operation queue includes waiting until the previous micro-operations from the instruction cache way are written or transferred to the operation queue before outputting the decoded micro-operations corresponding to the first operation cache queue to the operation queue.
12. The instruction fetch and decode unit of claim 10, wherein: The indication does not indicate a need to wait for a previous micro-operation from the instruction cache path to be written or transferred to the operation queue; and Outputting the decoded micro-op to an operation queue without waiting for a previous micro-op from the instruction cache way includes outputting the decoded micro-op corresponding to the first operation cache queue entry.
13. The instruction fetch and decode unit of claim 10, further comprising: A fetch control unit is configured to generate the first operating cache queue entry by chaining multiple hits of the first predicted address block in an operating cache tag array in a single cycle and including information for each such hit in the first operating cache queue entry.
14. The instruction fetch and decode unit of claim 10, wherein the shared fetch logic and the instruction cache path are configured to convert the instruction addresses of the second predicted address block into decoded micro-operations for output to the operation queue and subsequent execution by the rest of the instruction execution pipeline by: identifying that the second predicted address block includes at least one instruction for which decoded micro-operations are not stored in the operation cache; storing the instruction cache queue entry in the instruction cache queue; obtaining an instruction byte in an instruction byte buffer for the instruction cache queue entry along with an indication of whether to wait for a previous operation from the operation cache to be written or transferred to the operation cache; decoding the instruction byte to obtain a decoded micro-operation corresponding to the instruction byte buffer entry at a time based on the indication indicating whether to wait for the previous operation from the operation cache way; as well as The decoded micro-operation corresponding to the instruction byte buffer entry is output to the operation queue for storage.
15. The instruction fetch and decode unit of claim 14, wherein: The indication indicating whether to wait for the previous operation from the operation cache path indicates that waiting for the previous operation from the operation cache path is required; as well as Decoding the instruction byte includes decoding the instruction byte after the previous operation from the operation cache way is written or transferred to the operation queue.
16. The instruction fetch and decode unit of claim 14, wherein: The indication indicating whether to wait for the previous operation from the operation cache way indicates that it is not necessary to wait for the previous operation from the operation cache way; and Decoding the instruction byte includes decoding the instruction byte without waiting for a previous operation from the operation cache way to be written or transferred to the operation queue.
17. The instruction fetch and decode unit of claim 10, wherein: The first predicted address block includes at least one instruction whose decoded micro-operation is stored in an operation cache and at least one instruction whose decoded operation is not stored in the operation cache.
18. A processor comprising: an instruction fetch and decode unit for converting instruction addresses of the first predicted address block into decoded micro-operations for output to an operation queue storing the decoded micro-operations in program order and for subsequent execution by the remainder of the instruction execution pipeline; as well as the remainder of the instruction execution pipeline, The instruction fetch and decode unit comprises: Shared extraction logic, which is configured to: providing an index associated with the first predicted address block to an operating cache tag array to obtain a first tag associated with a first operating cache tag array entry and a second tag associated with a second operating cache tag array entry, determining, in a first computer clock cycle, that both the first tag and the second tag match tags derived from the predicted block of addresses and that an end address associated with the first operating cache tag array entry matches a start address associated with the second operating cache tag array entry, and in the first computer clock cycle, identifying a set of instructions for which decoded micro-operations are stored in an operation cache data array of an operation cache way as instructions at addresses associated with both the first operation cache tag array entry and the second operation cache tag array entry; an operation cache queue configured to store, for the set of instructions, a first operation cache queue entry, the first operation cache queue entry including an indication of whether to wait for receiving a signal from an instruction cache way to proceed; and an operation cache data read logic configured to obtain the decoded micro-operation corresponding to the first operation cache queue entry from the operation cache at a time based on the indication of the first operation cache queue entry indicating whether to wait to receive the signal from the instruction cache way to proceed, and output the decoded micro-operation corresponding to the first operation cache queue entry to the operation queue.
19. The processor of claim 18, wherein: The indication indicates a need to wait for a previous micro-operation from the instruction cache path to be written or transferred to the operation queue; as well as Outputting the decoded micro-operations to the operation queue includes waiting until the previous micro-operations from the instruction cache way are written or transferred to the operation queue before outputting the decoded micro-operations corresponding to the first operation cache queue to the operation queue.
20. The processor of claim 18, wherein: The indication does not indicate a need to wait for a previous micro-operation from the instruction cache path to be written or transferred to the operation queue; and Outputting the decoded micro-op to an operation queue without waiting for a previous micro-op from the instruction cache way includes outputting the decoded micro-op corresponding to the first operation cache queue entry.
Citation Information
Patent Citations
Method and apparatus for pipeline inclusion and instruction restarts in a micro-op cache of a processor
US8433850B2