Fast mapper recovery for refresh in processor

By accelerating the recovery process of the mapper in the processor core, the recovery delay and power consumption in the prior art is solved, and faster processor performance is achieved.

CN120303645APending Publication Date: 2025-07-11INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380083341.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-05
Filing Date
2023-11-17
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In processor design, during the refresh recovery process of the mapper, the prior art requires additional recovery ports to cause power and congestion problems, and the recovery delay is long, affecting processing performance.

Method used

The temporary latch accelerates the recovery process of the mapper by using the temporary latch in the processor core and immediately restores the mapper state when a refresh command is received without waiting for a comparison of the save and restore buffers.

Benefits of technology

Reduces recovery latency, avoids additional recovery port overhead, and improves processor processing performance and speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303645A_ABST
    Figure CN120303645A_ABST
Patent Text Reader

Abstract

A method for restoring a mapper of a processor core includes storing first information in a temporary latch. The first information represents a newly dispatched first instruction of the processor core and is saved in an entry latch of the save and restore buffer. In response to receiving a refresh command from the processor core, resume of the mapper is initiated with first information from the temporary latch without waiting for a comparison of a refresh tag of the refresh command to an entry latch of the save and resume buffer. A processor core configured to perform the above method is also provided. A processor core is also provided that includes a dispatcher, a mapper, a save and restore buffer including an entry latch and connected to the mapper via at least one pipeline, and a register disposed in the at least one pipeline.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present invention generally relates to a processor used in a computer system and reading and executing software code input into the processor. Such a processor is used on a computer chip implemented in a computer system such as a personal computer, a server, and / or other computers. The processor executes software code by accessing instructions from the code and executing these instructions.

[0002] In processor design, a mapping (i.e., logical to physical mapping) is read out from a mapper according to a request of a dispatched instruction. The mapper also provides a new physical mapping for a logical destination requested by the instruction. The dispatched instruction is written to the mapper, and this writing to the mapper causes the previous mapper state to be written to a save and restore buffer (SRB) for safe storage. This storage of the previous mapper state is for use in the case where an instruction is flushed out before it is completed. An instruction may need to be flushed due to a branch misprediction or an exception event. When an instruction is flushed, the flushed destination in the mapper must be restored to the previous state that the mapper held before the flush. This restoration is a rollback to the correct state in order to allow the correct execution of the instruction. This restoration must occur before new instruction dispatching can be restarted. A race condition occurs between (1) newly fetched instructions after the flush being issued from the I-cache and (2) the save and restore buffer restoring the mapper to its previous state. If the restoration of the previous destination to the mapper takes too long, the new instructions issued from the I-cache will need to stall at the dispatch until the mapper is fully restored. This stall at the dispatch impairs processing performance.

[0003] In the article "Speculative Restore of History Buffer in a Microprocessor" (ip.com number: IPCOM000250357D) from ip.com, a register restoration pipeline is disclosed, which includes checking each entry in a history buffer when a branch flush is encountered to determine whether its evictor has been flushed. If the evictor has been flushed, then the previous producer instruction tag data must be restored. If the evictor is younger than the flush and the producer entry instruction tag is older than the flush, then the entry is restored. A set of entries can set status bits. Entries can be selected from the entries to be restored. The entries are broadcast to an issue queue and to the mapper to restore the data and instruction tags back to the register file.

[0004] The process of the ip.com article improves the latency of the save and restore buffer, but it is achieved by speculating on the complete flush process.

[0005] Accelerated save and restore buffers have also been implemented by adding additional restore ports to the mapper.

[0006] However, adding more restore ports is expensive in terms of power and congestion. SUMMARY OF THE INVENTION

[0007] A method for restoring a mapper of a processor core is provided. First information is saved in a staging latch. The first information represents a first instruction newly dispatched by the processor core. The first information is also saved in an entry latch of a save and restore buffer of the processor core. In response to receiving a flush command for the processor core, restoration of the mapper is started with the first information from the staging latch without waiting for comparison of a flush tag of the flush command with the entry latch of the save and restore buffer. A processor core configured to execute the above method is also provided.

[0008] With these embodiments, a new way is provided to accelerate restoration after flushing, which avoids having to add additional expensive restore ports. With these embodiments, the restore latency is improved by a low-overhead mechanism.

[0009] In another embodiment, a processor core is provided that includes a dispatcher, a mapper, a save and restore buffer including entry latches and connected to the mapper by at least one pipeline, and a first register arranged in at least one pipeline. The processor core is configured to save first information in the register. The first information represents a first instruction newly dispatched from the dispatcher. The processor core is also configured to save the first information in an entry latch of the save and restore buffer. The processor is further configured such that in response to receiving a flush command for the processor core, the first instruction in the mapper is restored using the first information saved in the first register arranged in at least one pipeline.

[0010] With this embodiment, restoration after flushing is accelerated in a way that avoids adding additional expensive wires for restore ports between design entities. With these embodiments, the restore latency is improved by a low-overhead mechanism.

[0011] In at least some further embodiments, in response to dispatch of a second instruction after dispatch of the first instruction, the first information is transferred from the staging latch to another staging latch, second information representing the second instruction is saved in the staging latch, and the second information is saved in another entry latch of the save and restore buffer. Further in response to receiving a flush command for the processor core, restoration of the mapper is continued with the second information from the staging latch without waiting for comparison of a flush tag of the flush command with the other entry latch of the save and restore buffer.

[0012] Using this embodiment, the recovery after a refresh can be accelerated by at least two processor cycles. This two-processor-cycle acceleration allows improvement of the recovery latency without adding additional expensive recovery ports (e.g., without adding additional running wires between design elements).

[0013] In at least some further embodiments, transferring the first information from a staging latch to another staging latch occurs in a direction opposite to the mapper recovery direction of the recovery pipeline.

[0014] Using these further embodiments, the maximum assumption of the correlation with mapper recovery can be placed on the most recently dispatched instruction rather than on the second-most recently dispatched instruction.

[0015] In at least some further embodiments, further in response to receiving a refresh command, the first information is marked as recovery complete in the entry latch of the save and recovery buffer.

[0016] Using these embodiments, redundant recovery actions can be avoided, which can help reduce or eliminate unnecessary power consumption of the processor core during the refresh recovery process.

[0017] In at least some further embodiments, in response to the completion of a refresh corresponding to the refresh command or in response to the completion of the first information in the mapper, the first information is cleared from the staging latch.

[0018] Using these embodiments, the accuracy of the accelerated refresh recovery can be improved when the mapper recovers to the pre-error state. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] These and other objects, features, and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. The various features in the drawings are not drawn to scale as the illustrations are for clarity purposes to enable those skilled in the art to understand the present invention in conjunction with the detailed description. In the drawings:

[0020] Figure 1 is a block diagram showing a processing system according to at least one embodiment;

[0021] Figure 2 is a block diagram showing a portion of an accelerated refresh recovery processor core according to at least one embodiment, which can be implemented in one or more of the processors of the processing system shown in Figure 1 the processing system shown;

[0022] Figure 3is a schematic diagram of a processor pipeline showing an accelerated refresh recovery process according to at least one embodiment, which can be performed by Figure 2 the accelerated refresh recovery processor core shown;

[0023] Figure 4 is a block diagram showing a portion of an alternative accelerated refresh recovery processor core according to at least one embodiment, which can be implemented in one or more of the processors of the processing system shown in Figure 1 ;

[0024] Figure 5 is a schematic diagram of a processor pipeline showing an alternative accelerated refresh recovery process according to at least one embodiment, which can be performed by Figure 4 the alternative accelerated refresh recovery processor core shown; and

[0025] Figure 6 is a block diagram of a computer environment with multiple computer systems showing a processor pipeline diagram in which the Figure 2 and / or Figure 4 shown accelerated refresh recovery processor core and Figure 3 and / or Figure 5 can be implemented; DETAILED DESCRIPTION

[0026] Detailed embodiments of the claimed structures and methods are disclosed herein; however, it is to be understood that the disclosed embodiments are merely examples of the claimed structures and methods, which may be embodied in various forms. The present invention may be embodied in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope of the invention to those skilled in the art. In the description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments.

[0027] The exemplary embodiments described below provide a processor, a computer system, and a method for operating a processor that improve the latency of the recovery process associated with refreshing, thus increasing the overall computer processing speed. Due to a design improvement based on the observation that the most recently dispatched instructions are most likely to be refreshed, the recovery using the save and restore buffer is accelerated. By first recovering the most recently dispatched instructions without having to compare the instruction tags with the entry latches of the save and restore buffer, the time required for refresh recovery is reduced. This speculative recovery process can accelerate the save and restore buffer by one or more cycles from the time of refresh, e.g., by two or more cycles. These cycles may refer to processor core clock cycles. The accelerated refresh recovery can allow avoiding one or more cycles of dispatch stalls.

[0028] This latency speedup can be achieved by saving newly dispatched instructions for two cycles in one or more scratch latches, such as in a series of scratch latches, while also writing these instructions to the entry latch of the save and restore buffer. These scratch latches are part of and / or adjacent to the restore pipeline between the entry latch of the save and restore buffer and the mapper. The scratch latches can be examples of corresponding register states. These scratch latches can be latches with a one-cycle refresh and a two-cycle refresh. When a refresh occurs, the instructions from these scratch latches are compared with a refresh mask to determine whether these instructions are being refreshed. The refresh instructions identified from these stages will be restored first, while the remaining destinations in the entry latch of the save and restore buffer are being processed for restoration. The instructions to be restored from these stages will be multiplexed to the buffer restore port without comparison with the instruction tags from the entry latch of the save and restore buffer. Thereby, the restore process of the mapper can start one or more cycles earlier. The use of the most recently dispatched instructions in these scratch latches prioritizes information representing one or more of the most recently dispatched instructions during the refresh restore process. The scratch latches represent secondary backup entries closer to and / or in the restore pipeline.

[0029] Accelerated refresh saves the time previously spent selecting the correct instructions from the save and restore buffer and pipelining the selected instructions to the correct stages of the restore pipeline. Before receiving a refresh command, information from recently dispatched instructions is used to populate the correct start-up phase of the restore pipeline. Traditional refresh comparison can be used to identify the later instructions for completing the refresh restore rather than the recently dispatched instructions. In the accelerated path, the first dispatch cycle can be achieved in the traditional refresh + 3 cycles of the restore pipeline. Entries with restore information can be read from the save and restore buffer in parallel with determining whether a restore is needed. Additionally, in the accelerated path, information from additional scratch latches can be queued in an accelerated manner to move to the position of the traditional refresh + 3 cycles and compared in parallel with the refresh information. Thus, the traditional later stages of the refresh restore process are reached and executed faster through embodiments of the present disclosure. By allowing early processing of the most recently dispatched instructions, embodiments of the present disclosure compensate for the constraint of processing one instruction per pipeline per cycle. The early processing occurs by providing a new storage location for the instruction information that is also stored in the entry latch of the save and restore buffer.

[0030] Instructions restored from these scratch latches in the accelerated path are marked as "restore complete" in the save and restore buffer in at least some embodiments to prevent these instructions from being restored again. Such an attempt at repeated restoration would be redundant and would slow down the restore process.

[0031] Upon dispatch, in at least some embodiments, the processor core of the present disclosure may write the entries evicted from the mapper into the entries of the save and restore buffer. If the entry does not simultaneously flush / restore another thread, the save and restore buffer address of this entry is written into the restore pipeline of the slice to which this entry is dispatched. The thread ID is also stored in the restore pipeline together with the entry data. The SRB address and the valid bit are written into the staging latch as part of the main path flush +2 cycle phase. This information in the staging latch can be accessed without waiting for the traditional initial confirmation lookup without adding too much new state. This information can be used to write into the restore latch associated with the entry location. In embodiments having two staging latches, upon each new dispatch, the dispatch information can be pipelined from the first staging latch to the second staging latch. After the flush, the staging latch entries can be cleared as usual.

[0032] Upon flush, in at least some embodiments, the processor core described in the present disclosure detects whether the same thread is being flushed. If the flush is for a different thread, then this particular staging latch will not be used for the mapper restoration. If the processor core is flushing the same thread and if the restore address is valid, then the processor core retrieves the correct information from the entry latch of the save and restore buffer using the data from the staging latch and then drives the retrieved information to the next stage of the save and restore buffer restore pipeline, such as the output latch. If the mapper restoration of this stage is active and looking for an entry, then in some embodiments, the new dispatch from the save latch is given priority. The priority is design specific. If the flush restore pipeline is currently active on a known restore, then the existing known restore will take precedence over the newly dispatched restore operation.

[0033] Upon completing an instruction whose information is held in one or more staging latches as described herein, in at least some embodiments, the processor core of the present disclosure clears the instruction information located in one or more of the involved staging latches so as to prevent these instructions from being restored during the flush. This clearing can occur by changing the valid bit to indicate invalid information.

[0034] This embodiment helps to achieve faster mapper restoration, thus improving the processing speed. Therefore, a computer system having a processor with one or more of the improvements described herein executes and performs flush restoration more quickly to re - establish the correct order and information of the instructions in the mapper. Utilizing this improved mapper restoration, the described embodiments can improve the processing performance and processing speed of a computer processor.

[0035] Now refer to Figure 1, shows a processing system 100 according to an embodiment of the present disclosure. The depicted processing system 100 includes a plurality of processors, including a first processor 10A, a second processor 10B, a third processor 10C, and a fourth processor 10D. Each of the first processor 10A, the second processor 10B, the third processor 10C, and the fourth processor 10D may be designed and have components consistent with one or more of these embodiments, for example, may be designed and have components consistent with Figure 2 the accelerated refresh recovery processor core 22 shown, and configured to execute Figure 3 the accelerated refresh recovery processing described and shown therein. The processing system 100 depicted with a plurality of processors is illustrative. Other processing systems according to other embodiments may include a single processor with symmetric multi-threading (SMT) cores. The first processor 10A includes a first processor core 11A, a second processor core 11B, and local storage 12, which may be at the cache level or the internal system memory level. The second processor 10B, the third processor 10C, and the fourth processor 10D may have similar internal components and / or the same internal component design as shown for the first processor 10A. The first processor 10A, the second processor 10B, the third processor 10C, and the fourth processor 10D are coupled to a main system memory 14 and a storage subsystem 16, which includes a non-removable drive and an optical disc drive for reading a first portable computer-readable tangible storage device 17. The processing system 100 also includes input / output (I / O) interfaces and devices 18, such as a mouse and a keyboard for receiving user input and a graphical display for displaying information. The various processors described may be microprocessors, integrated circuits, embedded systems, and / or equivalents.

[0036] As will be discussed with reference to Figure 6 , the processing system 100 may also be implemented in one or more of the computers described in the computing environment 600.

[0037] Although a Figure 1 system is used to provide an illustration of a system implementing the processor architecture of these embodiments, it is understood that the depicted system is not restrictive and is intended to provide an example of a suitable computer system for applying the techniques of these embodiments. It should be understood that Figure 1 no limitation on the environment in which different processor core embodiments may be implemented is implied. Many modifications may be made to the depicted environment according to design and implementation requirements.

[0038] Figure 2 shows a block diagram of a portion of an accelerated refresh recovery processor core 200 according to at least one embodiment, which may be implemented in Figure 1in one or more processors of the processing system 100 shown. The accelerated refresh recovery processor core 200 may include an instruction cache 201, a dispatcher 202, a mapper 204, and a plurality of save and restore buffers 206a, 206b, 206c, 206d that are separated into slices 0, 1, 2, 3 respectively. The accelerated refresh recovery processor core 200 may include Figure 2 additional components not shown in the figure, such as an issue queue, execution slices, etc.

[0039] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices to be executed via the accelerated refresh processor core 200 and cause a series of operation steps to be executed on the computer, other programmable apparatus, or other devices to generate a computer-implemented process. The instructions are executed on the computer, other programmable apparatus, or other devices via the accelerated refresh recovery processor core 200.

[0040] The instruction cache 201 may store, for example, in a temporary manner, one or more instruction streams fetched from another cache or system memory and transmit the instruction streams to the dispatcher 202. The instruction cache 201 itself may include cache memory and, in some cases, an instruction buffer. The instruction cache 201 may feed into the dispatcher 202. The dispatcher 202 may include a routing network and may send the received instructions to a queue to be prepared for scheduling for issue and execution. In the depicted embodiment, the dispatcher 202 feeds directly into the mapper 204.

[0041] The mapper 204 manages register tags that act as pointers to data in an array in the processor, such as slice target file (STF) tags. The mapper 204 is logically a table that maps logical registers, such as general-purpose registers (GPRs), to the instructions that produce results and the destination locations for executing the instructions. The mapper 204 may use tags to perform this mapping. For example, the mapper may use an instruction tag (ITAG) to refer to the instruction that produces a result. The mapper 204 may use an STF tag to refer to the destination location of the instruction. The destination location is the physical location where data is written as part of the execution of the instruction.

[0042] The mapper 204 can be a system logic mapper and can be divided into multiple partitions according to the logical register type. For example, the mapper can be divided into a general-purpose register (GPR) partition, a floating-point register / low vector scalar register (V0) partition, and a high vector scalar register (V1) partition. As an example, the logical register mapper can have ninety-six registers, thirty-two GPR registers, thirty-two V0 registers, and thirty-two V1 registers. The most significant bit of each of the mapper registers can be used to distinguish various logical register allocations, where registers 0 - 31 are GPRs, 32 - 63 are V0s, and 64 - 95 are V1 registers.

[0043] As the instructions are executed, the logical register mapper values can be copied to save and restore buffer locations in one or more save and restore buffers. For example, in Figure 2 this copying is indicated by a copy arrow 216 that extends from the mapper 204 to the first save and restore buffer 206a and specifically to the entry latch 208a of the first save and restore buffer 206a. The copying can also occur from the mapper 204 to other save and restore buffers 206b, 206c, 206d, and specifically to other entry latches 208b, 208c, 208d of the other save and restore buffers 206b, 206c, 206d.

[0044] A save and restore buffer is a structure that is used to save a previous mapper state so that the state can be restored if one or more instructions need to be flushed. The save and restore buffer includes a memory structure (such as a latch) for storing entries that hold information for restoring specific mapper registers to what the mapper requires in response to a flush. The latch can include a memory structure, such as a pulse-sensitive memory cell circuit that has the ability to change state in response to a specific input pulse level. Latching is the process of temporarily storing a signal to maintain a specific level state and record information. The save and restore buffer in an accelerated flush recovery processor core can have a plurality of entry latches that are more numerous than the number of scratch latches in the accelerated flush recovery processor core, where the scratch latches store recently dispatched instructions as described in this disclosure, because the save and restore buffer saves a complete set of instruction information for flush recovery, but the scratch latches contain a partial set that may include information representing the most recently dispatched instructions. For example, the save and restore buffer entry latches can store information for four threads, eight threads, one thread, etc. The save and restore buffer entry latches are also located at and / or near the beginning stage of the recovery pipelines (“RP”) 210a, 210b, 210c, 210d, while the scratch latches are arranged in the middle stage of the recovery pipelines 210a, 210b, 210c, 210d. The entry latches of the save and restore buffer also maintain their state until their contents are flushed and restored or until the instructions corresponding to their contents are completed; however, as part of the cycle of newly received information, the contents of the scratch latches can be replaced with information corresponding to the next dispatch.

[0045] Instruction information can be retrieved from the save and restore buffer entry latch according to the type of instruction information and sent to specific restore pipelines 210a, 212a, 214a. Instruction information from the first save and restore buffer 206a and directed towards the general - purpose processor register "GPR" can be restored through the restore pipeline 210a. Instruction information from the first save and restore buffer 206a and directed towards the floating - point register / low - vector scalar register ("V0") partition can be restored through the restore pipeline 212a. Instruction information from the first save and restore buffer 206a and directed towards the high - vector scalar register ("V1") partition can be restored through the restore pipeline 214a. Similar routing of instruction information types can also be performed in other save and restore buffers 206b, 206c, 206d. For example, the restore pipelines 210b, 210c, 210d of the other three save and restore buffers 206b, 206c, 206d can each be GPR restore pipelines and directed towards the GPR partition in the mapper 204. The restore pipelines 212b, 212c, 212d of the other three save and restore buffers 206b, 206c, 206d can each be V0 restore pipelines and directed towards the V0 partition in the mapper 204. The restore pipelines 214b, 214c, 214d of the other three save and restore buffers 206b, 206c, 206d can each be V1 restore pipelines and directed towards the V1 partition in the mapper 204.

[0046] Figure 2It shows that in addition to copying the instruction information into the entry latch of the save and restore buffer, the instruction information is also copied into the staging latches, such as the first staging latches 22a, 22c, 22e and / or the second staging latches 22b, 22d, 22f. Such additional copying to the first staging latch 22a is indicated by a staging arrow 218 extending from the mapper 204 to the first staging latch 22a. These staging latches 22a, 22c, 22e, 22b, 22d, 22f may also include a memory structure, such as a type of pulse-sensitive memory cell circuit capable of changing state in response to a particular input pulse level. These staging latches can temporarily store signals to maintain a particular level state and record information. The first staging latches may be arranged within and / or beside the corresponding restore pipelines. For example, the first staging latch 22a may be arranged within and / or beside the GPR restore pipeline 210a of the first save and restore buffer 206a, the first staging latch 22c may be arranged within and / or beside the V0 restore pipeline 212a of the first save and restore buffer 206a, and the first staging latch 22e may be arranged within and / or beside the V1 restore pipeline 210c of the first save and restore buffer 206a. The first staging latches 22a, 22c, 22e may have port connections to various restore pipelines, for example, to the GPR restore pipeline 210a, to the V0 restore pipeline 212a, and to the V1 restore pipeline 214a. The second staging latches 22b, 22d, 22f may be formed of the same type of memory structure as that of the first staging latches 22a, 22c, 22e. In some embodiments, the first staging latches and the second staging latches 22a, 22c, 22e, 22b, 22d, 22f are partially arranged in the conventional restore pipelines but have some additional information to allow for refresh calculations.

[0047] The instruction information saved in the first staging latch and / or the second staging latch may include the corresponding instruction tag and a logical register indicator. The logical register indicator indicates which part and / or address of the save and restore buffer stores the information of the corresponding instruction. For example, which particular entry latch or latches of the save and restore buffer store the information of the corresponding instruction. The restore pipeline 210a may include metal wiring 220 extending between different logical parts of the accelerated refresh restore processor core 200. Although these staging latches are depicted in the various restore pipelines of the first save and restore buffer 206a ("slice 0"), additional similar staging latches may also be implemented in the restore pipelines of other save and restore buffers, such as the second, third, and fourth save and restore buffers 206b, 206c, 206d.

[0048] The most recently dispatched instruction from dispatcher 202 can be stored in one of the first scratchpad latches 22a, 22c, 22e. In response to another new instruction (e.g., a second instruction) being dispatched from dispatcher 202 and the previous first instruction still being in first scratchpad latch 22a, 22c, or 22e, the first instruction can be transferred from first scratchpad latch 22a, 22c, or 22e to the corresponding second scratchpad latch 22b, 22d, or 22f to make room in the first scratchpad latch to hold the updated instruction (the second instruction). The second scratchpad latch will be stored partially in the normal recovery flow latch but with some additional information to allow for flush calculations. Figure 2 The latch transfer arrow 219 shown in Figure 2 illustrates an example of such a cyclic flow transfer between first scratchpad latch 22a and second scratchpad latch 22b. Thus, after this rolling transfer scenario is completed, the first scratchpad latch can hold the second instruction, i.e., the new first instruction or the most recently dispatched instruction, and the second scratchpad latch holds the original first instruction or the second most recently dispatched instruction. In at least some embodiments, the accelerated flush recovery processor core 200 can include three pipelines (GPR / V0 / V1) per slice, and for each of the three pipelines per slice, multiple dispatch cycles can occur, such as first and second dispatch cycles. The rolling cycle occurs per pipeline / register type, e.g., GPR / VR0 / VR1.

[0049] In response to a further instruction, e.g., a third instruction, being dispatched from dispatcher 202, the original first instruction can be pushed out or deleted from second scratchpad latch 22b, the second instruction can be transferred from first scratchpad latch 22a to second scratchpad latch 22b, and the third instruction can be stored in first scratchpad latch 22a as a further new first instruction.

[0050] Thus, the scratchpad latches can capture the instruction information of recently dispatched instructions in a rolling cyclic manner, temporarily holding the information of a specific number of instructions according to the number of scratchpad latches for a particular save and restore buffer slice.

[0051] It will be Figure 3 described in more detail in the accelerated flush recovery processor pipeline shown in Figure 3 the scratchpad latches and their role in accelerated flush recovery.

[0052] To perform a flush recovery operation on the basis of the partition type, recovery pipelines 210a, 212a, 214a, 210b, 212b, 214b, 210c, 212c, 214c, 210d, 212d, 214d can be provided between the entry latches 208a, 208b, 208c, 208d of the save and restore buffers 206a, 206b, 206c, 206d and the logical register mapper partitions of the mapper 204. The recovery pipelines can include one or more wired connections 220 in the architecture of the processor unit that link the save and restore buffers to the registers at the logical register mapper locations. As an example, each partition of the mapper, including thirty-two GPR registers, thirty-two V0 registers, and thirty-two V1 registers, can have a dedicated recovery pipeline from the SRB such that each mapper register receives specific type of instruction information via a recovery pipeline of the same allocated logical register type. Any mapper content can be written to any one of the SRB entry latches. In at least some embodiments, the save and restore buffers are not bound to the register type until the retrieved instruction information enters the recovery pipeline on its way back to the mapper 204. When an instruction is dispatched, the mapper values are copied to the SRB locations on a per-slice basis. In at least some embodiments, the save and restore buffers contain one or more of the mapper register values, the INSTRUCTION TAG of the instruction, the register file tag (such as the STF tag) associated with the value and pointing to the register file location, and the flush recovery field. All or part of this information can be referred to as the entry data of the entry latches of the save and restore buffers.

[0053] For the normal execution of an instruction, the mapper 204 will transfer the instruction to other components of the accelerated flush recovery processor core 200, such as the issue queue and the execution unit. The normal transfer is indicated by the exit arrow 230 shown in Figure 2 The accelerated flush recovery processor core 200 can be part of a single integrated circuit processor (such as a superscalar processor) and can include various execution units, registers, buffers, memories, and other functional units, such as an instruction decoding unit, an instruction issue unit, a load / store unit, an operand address generation unit, a fixed-point unit, etc. In at least one embodiment, the accelerated flush recovery processor core 200 is capable of issuing and executing instructions out of order.

[0054] After a refresh recovery event, such as a system exception, the accelerated refresh recovery processor core 200 sets the bits required for recovery to valid. The refresh recovery fields of the SRBs associated with the mapper values at the point prior to the refresh recovery event are set to indicate the need to restore these values to the mapper. Then, the SRB values are restored to the mapper registers of the mapper 204 using the recovery pipeline 210. Since the recovery pipeline 210 and the mapper registers are each divided by logical type, the accelerated refresh recovery processor core 200 can restore the mapper registers in parallel, thus allowing multiple registers to be restored per clock cycle. In one embodiment, two GPR, V0, and V1 register values can be restored from each of the four execution slices stored in the SRB location per clock cycle. Then, the restored values can be mapped to the mapper 204 through the logical register dedicated recovery pipeline 210. The ninety-six register values of the mapper 204 can be restored two per cycle from each of the four slices, and thus can be fully restored in only four cycles. In at least some embodiments, the recovery pipeline 210 is divided into each of the thirty mapper registers of the corresponding same logical register type by logical partial registers, rather than being directly connected to each of the entire ninety-six registers of the mapper 204. This division and / or partitioning of the recovery pipeline 210 and the mapper 204 enables the mapper 204 to be quickly and parallelly restored from the save and restore buffers 206a, 206b, 206c, 206d.

[0055] Figure 3 is a processor pipeline diagram 300 showing an accelerated refresh recovery process that can be performed by the Figure 2 accelerated refresh recovery processor core 200 shown in. The processor pipeline diagram 300 shows the repetition of some structures in some cycles but not in others to emphasize the particularly important elements of a specific cycle. Thus, for example, the SRB entry latch is depicted in the main path FL+1, FL+2, and FL+3 cycles but not in the main path FL+4, FL+5, and FL+6 cycles.

[0056] Figure 3 The left side of the processor pipeline diagram 300 shown in shows the accelerated refresh recovery and its corresponding cycle modifications, while Figure 3 the right side of the processor pipeline diagram 300 shown in shows the main path refresh recovery and its corresponding processor cycles. The accelerated refresh recovery processor core 200 can perform Figure 3 both the accelerated refresh recovery and the main path refresh recovery shown in and described herein.

[0057] Using the above regarding Figure 2A staging latch that stores information representing recently dispatched instructions is described to perform accelerated flush recovery. Accelerated flushing takes advantage of the concept that the most recently dispatched instructions are most likely to be flushed. Thus, placing the recently dispatched instructions in the staging latches of the save and restore buffer that are further along in the restore pipeline stage of the save and restore buffer constitutes a prioritization of these most recently dispatched instructions. Instruction information from these staging latches is restored first, without waiting for the instruction tag comparison that is normally performed with the entry latches of the save and restore buffer. Instruction information from the staging latches takes precedence over the normal entry latches of the save and restore buffer because the instruction information from the staging latches can be transferred to and restored by mapper 204 before the instruction information from mapper 204 has traveled far enough in restore pipeline 210 to be in a position to be transferred to mapper 204. With two staging latches, this acceleration process will make the restoration of mapper 204 at least two cycles faster from a flush.

[0058] The top of the processor pipeline diagram 300 indicates the start of the flushing process and can be referred to as the flush plus zero ("Flush+0") cycle. In this part of the pipeline, a broadcast flush command is received by various design elements such as save and restore buffers 206a, 206b, 206c, 206d and mapper 204. At this time, the design elements have received the flush command but have not had time to respond to the flush command. Receiving the flush command causes the design elements to start performing additional actions associated with flushing and flush recovery. Receiving the flush command can trigger notifications to the first staging latches 22a, 22c, 22e, the second staging latches 22b, 22d, 22f, and the entry latches 208a, 208b, 208c, 208d of the save and restore buffers 206a, 206b, 206c, 206d. Using the accelerated flush recovery processor core 200, the notification to the entry latches can cause the save and restore buffer entries corresponding to the instruction information stored in the first and second staging latches to be marked as having been restored, causing the main recovery path to skip these entries. Marking these entries as having been restored within the FL+1 cycle will help avoid reading Lreg information for these entries and will help avoid reading the remaining contents of their stored information in subsequent cycles. The receipt of the flush command by the save and restore buffer also forwards the instruction tag of the command to the first multiplexer 304 to initiate various processes such as flush recovery. The first multiplexer 304 performs the function of selecting the correct flush thread for comparison by the flush comparison in the comparison logic block 306.

[0059] Figure 3The left part shows that for accelerated path refresh, entries from the second scratchpad latch 22b are dispatched at a time equivalent to the main path cycle III. The main path refresh cycle requires comparing instruction tag information with the entry latches of the save and restore buffer, while the accelerated path does not need to wait for the comparison of instruction tag information with the SRB entry latches and can send the information stored in the scratchpad latch in a timely manner in the FL+1 cycle. Figure 3 The first dispatch cycle 301 is shown to indicate the dispatch of the information stored in the second scratchpad latch 22b. This scratchpad latch 22b can be along the stage between the SRB entry latch 208a and the mapper 204 of the restore pipeline 210a. The scratchpad latch can be used to store recently dispatched instructions because according to the process of the main path, this scratchpad latch is empty and / or unused in these initial stages of the refresh restore. The second scratchpad latch 22b can generally be in the output position of a multiplexer (such as the scratchpad latch multiplexer 320a). In this refresh+1 cycle, the information saved in the second scratchpad latch 22b can be compared with the thread information of the refresh mask to confirm thread matching. If thread matching is confirmed, the information saved in the second scratchpad latch 22b can be continued to be used for refresh restore. If thread matching is not confirmed, the information in this specific scratchpad latch 22b is irrelevant to the current refresh restore and can be put into a dormant state temporarily in the scratchpad latch 22b.

[0060] This information from the second scratchpad latches 22b, 22d, 22f can be provided from the first scratchpad latch for mapper restore in the first processor cycle. In at least some embodiments, the advantage of two cycles is achieved through accelerated processing. The main path refresh cycle does not have instruction information supply until the third main path cycle (III), while the accelerated path has instruction information supply in the first stage / cycle after receiving the refresh command. In the accelerated path cycle, the instruction information from the second scratchpad latch 22b, especially the lreg information of the relevant instruction, can be provided / read in the first processor cycle and is used to look up the instruction restore information at the correct address from the entry latches of the save and restore buffer.

[0061] In at least some embodiments, the information in the second scratchpad latch 22b was previously in the first scratchpad latch 22a and was passed to the second scratchpad latch 22b due to the arrival of a new instruction. Therefore, in some embodiments, the information in the second scratchpad latch 22b can represent the second latest dispatched instruction, and the information in the first scratchpad latch 22a can represent the latest dispatched instruction.

[0062] Also during the first processor cycle, instruction information from the first scratchpad latch 22a can be multiplexed into the recovery pipeline via the scratchpad latch multiplexer 320a. Prior to this time point when no information has yet been forwarded from the entry latch of the mapper 204, the scratchpad latch multiplexer 320a will default to forwarding the information received from the first scratchpad latch 22a because no other information is present at this scratchpad latch multiplexer 320a to compete with the information received from the first scratchpad latch 22a. Figure 3 The scratchpad latch multiplexer 320a is labeled in Figure 3 . Each of the first scratchpad latches 22a, 22c, 22e can be input into a separate scratchpad latch multiplexer that selects between the information from the respective first scratchpad latch and information flowing along the main recovery pipeline.

[0063] Figure 3 The second dispatch cycle 303 is shown in Figure 3 to indicate that the dispatch of the information stored in the first scratchpad latch 22a occurs within a time interval equivalent to the main path cycle II. During the second processor cycle, for the accelerated path refresh, instruction information retrieved from the entry latch using the address from the second scratchpad latch 22b can also be driven to the mapper 204 while the instruction information obtained from the first scratchpad latch 22a is read because now the first scratchpad latch instruction information is located in the second scratchpad latch 22b or an equivalent location within the recovery pipeline.

[0064] In the third processor cycle, for accelerated path refresh, instruction information retrieved from the entry latch using an address initially from the first scratchpad latch 22a (which, in some embodiments, is passed from the first scratchpad latch 22a to the second scratchpad latch 22b) can be driven to the mapper 204. Thus, in the third processor cycle of the accelerated refresh path, the first instruction information and the second instruction information are driven to the mapper 204. However, in the third processor cycle of the main path refresh, only the information of the entry latch being read from the save - restore buffer is being read. This performance acceleration of the accelerated path refresh shows how the use of scratchpad latches can demonstrate a cycle advantage for implementing mapper recovery, e.g., a two - cycle advantage. When the main path refresh occurs in combination with the accelerated refresh, the main path refresh can provide information from the third most recently dispatched instruction or other earlier instructions to continue the accelerated refresh recovery after a jump start is received from one or more scratchpad latches. The main path refresh can be designed to skip the entries corresponding to the most recently and second - most recently dispatched instructions, thus feeding other information to the middle or later part of the refresh recovery. For such an earlier instruction, e.g., the third instruction, the entry latch can be read in the third cycle after the refresh command is received. The retrieved / read entry can be driven to the mapper in the fourth cycle after the refresh command is received.

[0065] In the refresh + 1 cycle, the tag 302 is a refresh tag or a completion tag and is compared with the tags in the entries of the entry latches 208a, 208b, 208c, 208d. The refresh / completion tag for each thread goes into all entries. Within an entry, each entry multiplexes the correct thread refresh / completion itag and performs the comparison. The first multiplexer 304 can select between the refresh / completion among four threads. For example, one entry can be refreshing thread 0 while another entry is completing thread 1. By performing this selection, the first multiplexer 304 further forwards the refresh - related instructions into the refresh recovery pipeline. The first multiplexer 304 multiplexes the correct thread refresh / completion itag within each entry, so each entry has a multiplexer 304 (although Figure 3Only a single first multiplexer 304 is shown. A comparison logic block 306 is also provided for each entry, such that there will be more than one. However, the components, namely the tag 302, the restore vector 308, and the hold dispatch logic 312, are either sent to all entries of a particular slice or pulled from it. The number of instruction tags may correspond to the number of threads of the entry latches of the save and restore buffer, so the first multiplexer 304 selects an appropriate flush / complete tag to match the information held in the entry. The received information indicates what the save and restore buffer entry is holding. The first multiplexer 304 selects from the potential flush / complete tags 302 to provide to the comparison logic block 306. The comparison logic block 306 will perform a comparison of the flush tag with the entry tag (the evictor) and the previous tag (the evicted). If the evictor is flushed and the evicted is not flushed, the entry must be restored to the mapper. This comparison helps make a retrieval decision based on which instruction is doing the flushing. If an instruction is younger than the flush instruction, the younger instruction will participate in the flush and thus in the flush restore. In the traditional path, all latch entries of the save and restore buffer are read in the flush + 1 cycle to check for relevance to the flush command.

[0066] When an instruction completes, the processor notifies the save and restore buffer such that the entry for the corresponding instruction can be retired from the save and restore buffer. This retirement can be performed by removing the corresponding entry from the save and restore buffer. If one of the instruction tags 302 indicates that the instruction has completed, no Figure 3 of the two flush cycle paths shown needs to be performed for that instruction tag. If such an instruction tag indicates instruction completion, the accelerated flush restore processor core 200 can also compare with the entries of the first scratchpad latches 22a, 22c, 22e and / or the entries of the second scratchpad latches 22b, 22d, 22f, and if the entries match, clear the corresponding information in the scratchpad latches. This clearing of the scratchpad latches will prevent information from the completed instruction from being injected into the flush restore process, since information from a successfully completed instruction is not needed for the flush restore.

[0067] When a flush command is received, the flush command provides information about the instructions involved in the flush. The flush command information can be compared in the comparison logic block 306 with the entries of the entry latches 208a, 208b, 208c, 208d of the save and restore buffer. The comparison logic block 306 receives inputs from the entry latches 208a, 208b, 208c, 208d and also receives the output from the first multiplexer 304 in order to compare the information from two different input streams with each other. The flush tag in the instruction tag 302 indicates which tag is flushed and whether the flush tag is valid. For the flush comparison of the scratchpad latches, the comparison is performed further downstream in the restore pipeline compared to performing the comparison for the entry latches of the save and restore buffer. If the instruction information evicted by the corresponding instruction is not flushed and the corresponding instruction has been flushed, a flush recovery process needs to be performed. Thus, in the FL+1 cycle, the newly received instruction tag is compared with all the entries of the entry latches 208a, 208b, 208c, 208d. In some embodiments, the entry latch 208a may include registers 0-31 and may refer to GPRs, the entry latch 208b may include registers 32-63 and may refer to V0 registers, the entry latch 208c may include registers 64-95 and may refer to V1 registers, and the entry latch 208d may include registers 96:127. The number of collected instruction tags 302 may correspond to the number of entry latch partitions, e.g., for four save and restore buffers 206a, 206b, 206c, 206d, there are four instruction tags for the four entry latch partitions 208a, 208b, 208c, 208d respectively.

[0068] If a need for flush recovery is identified based on the comparison in the comparison logic block 306, the comparison logic block 306 forwards the flush recovery information to both the OR logic 310 and the restore vector 308. The OR (logical OR) logic 310 performs an OR operation on all the instructions involved in the flush recovery and then sends a set of instructions resulting from the OR operation to the hold dispatch logic 312 in order to set the latches of the hold dispatch logic 312. The hold dispatch logic 312 prevents any new dispatch on the relevant thread until the entries in the restore vector 308 are exhausted and are complete. The arrow extending from the hold dispatch logic 312 but without a receiving box indicates that hold instructions are sent to the dispatcher 202 and / or other components of the processor to enforce the hold / prevention of dispatch for a particular thread. The restore vector 308 may include latches having bits corresponding to the entry latches 208a, 208b, 208c, 208d of the save and restore buffers 206a, 206b, 206c, 206d. When an instruction is involved in the flush recovery, the bit of the latch of the restore vector corresponding to the entry of that instruction in the entry latch is set.

[0069] The OR logic 310, hold dispatch logic 312, and restore vector 308 together form a loop that continues until the flush recovery is complete. The relevant entries involved in the flush are retained in the restore vector 308 until these instructions are processed. Instructions will be processed one by one for each restore pipeline each time. This loop helps to avoid the situation where the accelerated flush recovery processor core 200 attempts to perform dispatch and flush recovery on the same mapper entry at the same time. Such a situation may trigger a processor error, and this loop helps to avoid that error.

[0070] In the main phase cycle II, the hold dispatch logic 312 includes bits set by the thread to be sent to the dispatch logic to prevent new dispatches from being dispatched to the thread being recovered. This bit setting is done by ORing the entries for which recovery is requested for the corresponding thread with the OR logic 310. Once all entries have completed recovery, the bits that set the restore vector 308 are cleared, which also clears the path through which it is ORed into the hold dispatch logic 312. This loop is similar to or equivalent to the loop for the accelerated path flush cycle shown on the left side in Figure 3 as described above.

[0071] In the main path cycle phase II, recovery is performed by filtering the recovery requests from the restore vector 308 through the logical registers. The filtered requests are passed to the lookup logic 314, where entries are selected for each register type. In some embodiments, the lookup logic 314 may include multiple partitions corresponding to multiple partitions of the save and restore buffer. For example, the lookup logic 314 may include a first lookup logic for general-purpose registers GPR, a second lookup logic for floating-point registers / low vector scalar registers V0, and a third lookup logic for high vector scalar registers V1.

[0072] Once the lookup logic 314 has selected the entries to be recovered from each partition, the entry numbers (e.g., the bits or numbers representing the selected restore vector) are encoded by the encoding logic 316 to reduce the number of bits passed down through the recovery logic.

[0073] The information of the relevant entries is sent as the output 320 from the second cycle phase. This output 320 is provided as an input to the third cycle phase.

[0074] These cycle phases I and II in the main path refresh cycle include functions and elements that are also performed in cycle phases I and II of the accelerated path refresh cycle. However, the main path refresh cycle does not include the acceleration activities related to the first staging latch and the second staging latch. The accelerated path refresh cycle in phases I and II includes dispatching hold loops and instruction tag comparisons with the entry latch, but these actions will be used for any subsequent instructions that need to be refreshed, rather than for one or two most recently dispatched instructions.

[0075] In the refresh +3 cycle in the main path, instruction entries are read from entry latches 208a, 208b, 208c, 208d in order to obtain information therefrom, so as to send this information to the mapper 204 for restoring the mapper 204. At this position, an lreg retrieved in the refresh +2 phase and part of the output 320 of each entry is used to find the correct address within the entry latches 208a, 208b, 208c, 208d. Then the corresponding address is read to obtain the correct instruction information. The read information is fed to the information multiplexer 322, and the information multiplexer 322 is fed to the stage three multiplexer 324 and the optimization multiplexer 326. The information multiplexer 322 selects an appropriate SRB interface layer 328 for a specific instruction information type. The optimization multiplexer 326 receives an input from the second optimization latch 331, and the second optimization latch 331 receives an input from the first optimization latch 329. The optimization latches 329, 331 help with customization based on port availability for simultaneous multithreading, which allows multiple instruction streams (threads) to run concurrently on the same physical processor.

[0076] For the accelerated path, retrievals / readings have occurred at the corresponding pipeline positions in the refresh +1 cycle and at the corresponding pipeline positions in the refresh +2 cycle. The retrieval in the refresh +1 cycle is for the entry of the second staging latch for information from the hold second most recently dispatched instruction. The retrieval in the refresh +2 cycle is for the entry of the first staging latch for information from the hold most recently dispatched instruction.

[0077] In the refresh +4 cycle in the main path, the retrieved / read instruction information is driven from the stage three multiplexer 324 and the optimization multiplexer 326 to the SRB interface layer 328 to drive the information to the mapper 204. The SRB interface layer 328 can drive this information by transmitting it to the mapper interface layer 330. This transmission can occur through Figure 2 the pipeline wiring 220 shown. The mapper interface layer 330 receives this transmission. In some cases, this pipeline wiring 220 can be referred to as the restore cycle pipeline. The SRB interface layer 328 and the mapper interface layer 330 can include latches specific to the restore pipeline partition.

[0078] For the acceleration path, the driving of information representing the read / retrieval of the instruction representing the second most recently dispatched instruction has occurred at the corresponding pipeline position in the flush +2 cycle. The driving of information representing the read / retrieval of the instruction representing the most recently dispatched instruction has occurred at the corresponding pipeline position in the flush +3 cycle.

[0079] In the flush +5 cycle in the main path, the mapper interface layer 330 feeds the received instruction information to the mapper multiplexer 332 for selection into the appropriate mapper partition 334. After being received, the received / retrieved instruction information is stored in the appropriate mapper partition 334.

[0080] For the acceleration path, the feeding of information representing the read / retrieval of the instruction representing the second most recently dispatched instruction to the mapper multiplexer 332, the selection into the appropriate mapper partition, and the storage have occurred at the corresponding pipeline positions in the flush +3 cycle. The feeding, selection, and storage of information representing the read / retrieval of the instruction representing the most recently dispatched instruction have occurred at these corresponding pipeline positions in the flush +4 cycle.

[0081] In the flush +6 cycle, the read / retrieved information saved in the mapper 204 can now be read from the mapper 204. Unless other instructions are also being restored for that particular flush, the mapper restoration is complete. The above steps can be repeated for each additional instruction involved in that particular flush operation.

[0082] For the acceleration path, the availability of the read / retrieval of information representing the instruction representing the second most recently dispatched instruction from the mapper 204 has occurred in the flush +4 cycle. The availability of the read / retrieval of information representing the instruction representing the most recently dispatched instruction from the mapper 204 has occurred in the flush +5 cycle.

[0083] During write-back, it is detected whether the write-back occurs on the same thread. If the write-back occurs on the same thread, the write-back instruction tag is compared with the instruction tag of the instruction in the acceleration recovery scratch latch. If this comparison indicates a match, the read bit (e.g., the W bit) of the mapper 204 is set to equal one. The write-back affects the data that can be restored to the mapper held in the save and restore buffers. The write-back function is the write bit used to indicate that a register has been written. This write bit can be read to see if the data is ready. In response to receiving a new instruction to write to the register of the mapper 204 from the dispatcher 202, the write bit is cleared from the mapper 204.

[0084] After completing the flush recovery, the mapper 204 can release the dispatch hold, and new instructions can be written to the mapper 204.

[0085] It can be understood that Figure 2 andFigure 3 Only illustrations of some embodiments are provided, and no limitations on how to implement different embodiments are implied. Many modifications can be made to the depicted embodiments (e.g., the depicted sequence of steps or the arrangement of processor components) according to design and implementation requirements.

[0086] Table 1 below summarizes the advantages of the acceleration path compared to the main path for utilizing the first scratchpad latch in the recovery mapper. Table 2 below summarizes the advantages of the acceleration path compared to the main path for utilizing two scratchpad latches during operation.

[0087]

[0088] Table 1

[0089]

[0090]

[0091] Table 2

[0092] Tests performed using the accelerated refresh recovery processor core 200 show an improvement in processing speed.

[0093] Figure 4 is a block diagram showing a portion of an alternative accelerated refresh recovery processor core 400, which includes many of the same components, structures, and arrangements as the Figure 2 shown accelerated refresh recovery processor core 200, but with the order of the first and second scratchpad latches reversed. Figure 5 shows an alternative processor pipeline schematic 500 that can be executed by the Figure 4 depicted alternative accelerated refresh recovery processor core 400 according to at least one embodiment. Figure 4 and Figure 5 show that for this alternative embodiment, the positions of the first and second scratchpad latches are swapped such that the alternative first scratchpad latch 44a is deployed downstream of the alternative second scratchpad latch 44b in the mapper recovery direction of the recovery pipeline. This alternative can include additional first and second scratchpad latches as implemented for the Figure 2 and Figure 3 embodiments, but for simplicity, only the alternative first and second scratchpad latches 44a, 44b are labeled in Figure 4 and Figure 5 respectively.

[0094] Figure 4It is shown that through this alternative arrangement, the cyclic flow transfer between the scratchpad latches occurs in the upstream direction opposite to the mapper recovery direction of the recovery pipeline. The cyclic flow transfer refers to the transfer of instruction information between the latches when a new instruction is dispatched, where the information of the new instruction is to be stored in the save and restore buffer and the scratchpad latches. As explained above, the first instruction can initially be located in the first scratchpad latch and then be pushed into another scratchpad latch when a new instruction enters. Specifically, Figure 4 The alternative latch transfer arrow 419 shown in Figure 4 illustrates an example of the alternative cyclic flow transfer between the alternative first scratchpad latch 44a and the alternative second scratchpad latch 44b. In contrast, Figure 2 The latch transfer arrow 219 shown in Figure 2 illustrates the cyclic flow transfer that occurs in the downstream direction with respect to the mapper recovery direction of the recovery pipeline between the first scratchpad latch 22a and the second scratchpad latch 22b.

[0095] Figure 4 and Figure 5 This alternative embodiment of Figure 5 achieves the following advantage: in an embodiment where there are two scratchpad latches for each instruction type in each slice, the most recently dispatched instruction receives the front position (downstream in the mapper recovery direction) of the scratchpad latch. In Figure 2 and Figure 3 In the embodiment shown, the second most recently dispatched instruction holds the front position (downstream in the mapper recovery direction) of the scratchpad latch. Therefore, Figure 4 and Figure 5 The alternative embodiments of Figure 5 emphasize the concept of achieving faster recovery by assuming that the most recently dispatched instruction is most likely to be involved in the mapper recovery. As shown in Figure 5 The alternative embodiment may include a cyclic multiplexer 55 in the flow path for the cyclic transfer between the alternative first scratchpad latch 44a and the alternative second scratchpad latch 44b. The cyclic multiplexer 55 can select the correct second scratchpad latch for the feedback instruction information popped from the first scratchpad latch 44a due to the arrival of a new instruction. Figure 2 and Figure 3 The embodiments shown in Figure 3 may be advantageous in that it is not necessary to add such an additional cyclic multiplexer to the normal main flow path components.

[0096] For this alternative embodiment, the most recently dispatched instruction will be driven to the mapper before the second most recently dispatched instruction. In the embodiments of Figure 2 and Figure 3 previously explained, the second most recently dispatched instruction is driven to the mapper before the most recently dispatched instruction. Figure 4 and Figure 5The remaining components of the acceleration path refresh period and the main path refresh period of the alternative embodiment shown operate substantially in the same manner as these components and periods described in the embodiments with respect to Figure 2 and Figure 3 .

[0097] It will be appreciated that Figure 4 and Figure 5 are provided only as illustrations of certain embodiments and do not imply any limitations as to how different embodiments may be implemented. Many modifications may be made to the depicted embodiments (e.g., the depicted sequence of steps or the arrangement of processor components) in accordance with design and implementation requirements.

[0098] Figure 6 is a block diagram of internal and external components of a computer system in which one or more of the processors described herein may be implemented. Computing environment 600 shows an example of one or more computers having processors that perform accelerated refresh recovery. Computing environment 600 includes, for example, computer 601, wide area network (WAN) 602, end user device (EUD) 603, remote server 604, public cloud 605, and private cloud 606. In this embodiment, computer 601 includes a set of processors 610 (including processing circuitry 620 and cache 621), communication fabric 611, volatile memory 612, persistent storage 613 (including operating system 622 and software program 616), a set of peripheral devices 614 (including a set of user interface (UI) devices 623, storage 624, and a set of Internet of Things (IoT) sensors 625), and network module 615. Remote server 604 includes remote database 630. Public cloud 605 includes gateway 640, cloud orchestration module 641, a set of host physical machines 642, a set of virtual machines 643, and a set of containers 644. The various computers shown may each include one or both of the accelerated refresh recovery processor cores 200 and alternative accelerated refresh recovery processor cores 400 described above.

[0099] Computer 601 may take the form of a desktop computer, laptop computer, tablet computer, smart phone, smart watch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or later developed that is capable of running programs, accessing a network, or querying a database such as remote database 630. The various computers may execute methods distributed among multiple computers and / or multiple locations. On the other hand, in this presentation of computing environment 600, the discussion focuses on a single computer, specifically computer 601, to keep the presentation as simple as possible. Computer 601 may be located in the cloud, even though Figure 6 is not shown in the cloud. On the other hand, computer 601 need not be in the cloud.

[0100] The processor set 610 includes one or more computer processors of any type now known or later to be developed, which are configured to execute the accelerated refresh recovery process described in this disclosure. The processing circuitry 620 may be distributed across multiple packages, such as multiple cooperative integrated circuit chips. The processing circuitry 620 may implement multiple processor threads and / or multiple processor cores. The cache 621 is memory located within the processor chip package and is generally used for data or code that should be available for rapid access by threads or cores running on the processor set 610. Cache memory is typically organized into multiple levels based on relative proximity to the processing circuitry. Alternatively, some or all of the cache in the processor set may be located "off-chip". In some computing environments, the processor set 610 may be designed to work with qubits and perform quantum computing. The processor set 610 may include one or more of the above-described accelerated refresh recovery processor cores 200 and alternative accelerated refresh recovery processor cores 400.

[0101] Computer-readable program instructions are typically loaded onto the computer 601 to cause a series of operational steps to be executed by the processor set 610 of the computer 601. These computer-readable program instructions are stored in various types of computer-readable storage media, such as the cache 621 and other storage media discussed below. The processor set 610 accesses the program instructions as well as associated data to control and direct the execution of the inventive method.

[0102] The communication structure 611 is a signal conduction path that allows some components of the computer 601 to communicate with each other. Generally, this structure consists of switches and conductive paths, such as switches and conductive paths that make up a bus, a bridge, a physical input / output port, etc. Other types of signal communication paths may be used, such as fiber optic communication paths and / or wireless communication paths.

[0103] The volatile memory 612 is any type of volatile memory now known or later to be developed. Examples include dynamic random access memory (RAM) or static RAM. Generally, the volatile memory 612 is characterized by random access, but this is not required unless explicitly stated. In the computer 601, the volatile memory 612 is located within a single package and within the computer 601, but alternatively or additionally, the volatile memory may be distributed across multiple packages and / or located external to the computer 601.

[0104] Persistent storage 613 is any form of non-volatile storage for a computer that is known now or will be developed in the future. The non-volatility of this storage means that the stored data is retained regardless of whether power is supplied to computer 601 and / or directly to persistent storage 613. Persistent storage 613 can be read-only memory (ROM), but typically at least a portion of the persistent storage allows for writing of data, deletion of data, and re-writing of data. Some common forms of persistent storage include disk and solid-state storage devices. Operating system 622 can take several forms, such as various known proprietary operating systems or operating systems of the open-source portable operating system interface type that employ a kernel.

[0105] Peripheral device set 614 includes a collection of the peripheral devices of computer 601. The data communication connections between the peripheral devices and the other components of computer 101 can be implemented in various ways, such as Bluetooth connections, near-field communication (NFC) connections, connections made by a cable (such as a universal serial bus (USB)-type cable), plug-in connections (e.g., Secure Digital (SD) card), connections made through a local communication network, and even connections made through a wide area network such as the Internet. In various embodiments, UI device set 623 can include components such as a display screen, speakers, microphones, wearable devices (such as glasses and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage 624 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. Storage 624 can be persistent and / or volatile. In some embodiments, storage 624 can take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 601 needs to have a large amount of storage (e.g., in the case where computer 601 locally stores and manages a large database), this storage can be provided by a peripheral storage device designed to store a very large amount of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. IoT sensor set 625 consists of sensors that can be used in Internet of Things applications. For example, one sensor can be a thermometer, while another sensor can be a motion detector. Various sensors and / or UI devices can each respectively include a package having one or more enhanced processors as described herein.

[0106] The network module 615 is a collection of computer software, hardware, and firmware that allows the computer 601 to communicate with other computers via the WAN 602. The network module 615 can include hardware such as a modem or a Wi-Fi signal transceiver, software for packetizing and / or depacketizing data transmitted over a communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control function and the network forwarding function of the network module 615 are executed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control function and the forwarding function of the network module 615 are executed on physically separate devices such that the control function manages several different network hardware devices. The computer-readable program instructions for performing the methods of the present invention can generally be downloaded to the computer 601 from an external computer or an external storage device via a network adapter or a network interface included in the network module 615.

[0107] The WAN 602 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances via any technology now known or later developed for transmitting computer data. In some embodiments, the WAN 602 can be replaced and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area, such as a Wi-Fi network. The WAN and / or LAN generally includes computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0108] The end-user device (EUD) 603 is any computer system used and controlled by an end user (e.g., a customer of an enterprise operating the computer 601) and can take any form discussed above in connection with the computer 601. The EUD 603 typically receives helpful and useful data from the operation of the computer 601. For example, in the hypothetical case where the computer 601 is designed to provide recommendations to an end user, the recommendation will typically be transmitted from the network module 615 of the computer 601 to the EUD 603 via the WAN 602. In this way, the EUD 603 can display or otherwise present the recommendation to the end user. In some embodiments, the EUD 603 can be a client device such as a thin client, a thick client, a mainframe computer, a desktop computer, etc. Each EUD 603 can separately include a package having an enhanced processor as described herein.

[0109] The remote server 604 is any computer system that provides at least some data and / or functionality to the computer 601. The remote server 604 can be controlled and used by the same entity that operates the computer 601. The remote server 604 represents a machine that collects and stores helpful and useful data used by other computers such as the computer 601. For example, in the hypothetical case where the computer 601 is designed and programmed to provide recommendations based on historical data, then that historical data can be provided to the computer 601 from the remote database 630 of the remote server 604. The remote server 604 can include an enclosure having one or more enhanced processors as described herein.

[0110] The public cloud 605 is any computer system that can be used by multiple entities and provides on-demand availability of computer system resources and / or other computer capabilities (notably data storage (cloud storage) and computing power) without direct active management by the user. Cloud computing typically exploits the sharing of resources to achieve economies of scale and consistency. The direct and active management of the computing resources of the public cloud 605 is performed by the computer hardware and / or software of the cloud orchestration module 641. The computing resources provided by the public cloud 605 are typically implemented by virtual computing environments running on various computers that make up the set of host physical machines 642, which is the universe of physical computers within and / or available to the public cloud 605. Each of these physical computers can include a corresponding enclosure having one or more enhanced processors as described herein.

[0111] The virtual computing environment (VCE) generally takes the form of virtual machines from the set of virtual machines 643 and / or containers from the set of containers 644. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after instantiation of the VCE. The cloud orchestration module 641 manages the transfer and storage of the images, deploys new instantiations of the VCE, and manages the active instantiations of the VCE deployment. The gateway 640 is a collection of computer software, hardware, and firmware that allows the public cloud 605 to communicate via the WAN 602.

[0112] Some further explanations of virtualized computing environments (VCEs) will now be provided. A VCE can be stored as an "image". New active instances of the VCE can be instantiated from the image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows for multiple isolated user space instances, called containers. From the perspective of the programs running within them, these isolated user space instances generally appear as actual computers. A computer program running on a normal operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU capabilities, and quantifiable hardware capabilities. However, a program running within a container can only use the contents of the container and the devices allocated to the container, which is a feature known as containerization.

[0113] A private cloud 606 is similar to a public cloud 605, except that the computing resources are only available to a single enterprise. Although the private cloud 606 is depicted as communicating with the WAN 602, in other embodiments, the private cloud can be completely disconnected from the Internet and only accessible through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), typically implemented by different vendors. Each of the multiple clouds remains as a separate and discrete entity, but the larger hybrid cloud architecture is tied together through standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple constituent clouds. In this embodiment, both the public cloud 605 and the private cloud 606 are part of the larger hybrid cloud.

[0114] In the computing environment 600, the computer 601 is shown connected to the Internet (see WAN 602). However, in many embodiments, the computer 601 will be isolated from communicating with a communication network and not connected to the Internet, operating as a stand-alone computer. In these embodiments, to ensure isolation and prevent external communication from entering the computer 601, the network module 615 of the computer 601 may not be needed or even desired. Stand-alone computer embodiments may be advantageous in at least some applications of the present invention because they are generally more secure. In other embodiments, the computer 601 is connected to a secure WAN or a secure LAN instead of the WAN 602 and / or the Internet. In embodiments where there is a network connection (i.e., not stand-alone), the system designer may wish to take appropriate security measures, now known or developed in the future, to reduce the risk that incoming network communication does not result in a security vulnerability.

[0115] One or more of the computer 601, computer components of the wide area network 602, the end-user device 603, the remote server 604, computers in the public cloud 605, and computers in the private cloud 606 may include one or more processors that perform an accelerated refresh recovery process as described herein.

[0116] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present invention. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the figures. For example, two blocks shown in succession may actually be completed as one step, executed simultaneously, executed substantially simultaneously, executed in a partially or fully time-overlapping manner, or depending on the functions involved, these blocks may sometimes be executed in the reverse order. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0117] The terms used herein are for the purpose of describing particular embodiments and are not intended to limit the invention. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "comprises", "comprising", "includes", "including", "has", "having", "with", etc. specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0118] The description of the various embodiments of the invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology found in the marketplace, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for restoring a mapper of a processor core, the method comprising: Storing first information in a scratch latch, the first information representing a newly dispatched first instruction of the processor core; Storing the first information in an entry latch of a save and restore buffer of the processor core; And In response to receiving a flush command for the processor core, starting the restoration of the mapper with the first information from the scratch latch without waiting for a comparison of a flush tag of the flush command with the entry latch of the save and restore buffer.

2. The method according to claim 1, further comprising: Further in response to receiving the flush command, marking the first information in the entry latch of the save and restore buffer as restoration complete.

3. The method according to claim 1, further comprising: Further in response to receiving the flush command, processing remaining instruction information from the save and restore buffer for transmission to the mapper.

4. The method according to claim 1, further comprising: In response to the dispatch of a second instruction after the dispatch of the first instruction: Transferring the first information from the scratch latch to another scratch latch, Storing second information representing the second instruction in the scratch latch, and Storing the second information in another entry latch of the save and restore buffer; And Further in response to receiving the flush command for the processor core, continuing the restoration of the mapper with the second information from the scratch latch without waiting for a comparison of a flush tag of the flush command with the said another entry latch of the save and restore buffer.

5. The method according to claim 4, wherein in a first processor cycle after receiving the flush command, a first logical register according to the first information in the scratch latch is used to read entry data from the entry latch of the save and restore buffer.

6. The method according to claim 5, wherein transferring the first information from the scratch latch to the said another scratch latch occurs in the first processor cycle.

7. The method according to claim 4, wherein transferring the first information from the scratch latch to the said another scratch latch occurs via a multiplexer within a recovery pipeline of the processor core.

8. The method according to claim 4, wherein in a second processor cycle after receiving the flush command, the first entry data read from the entry latch is driven to the mapper.

9. The method according to claim 8, wherein in the second processor cycle, a second logical register according to the second information in the scratch latch is used to read second entry data from the said another entry latch of the save and restore buffer.

10. The method according to claim 9, wherein in a third processor cycle after receiving the flush command, the read second entry data is driven to the mapper.

11. The method according to claim 9, wherein in a fourth processor cycle after receiving the flush command, third entry data read from the save and restore buffer is driven to the mapper.

12. The method according to claim 4, wherein the second instruction is from the most recently dispatched instruction, and the first instruction is from the second most recently dispatched instruction.

13. The method according to claim 4, wherein transferring the first information from the scratchpad latch to the other scratchpad latch occurs in a direction opposite to the mapper recovery direction of the recovery pipeline.

14. The method according to claim 13, wherein the first instruction is from the most recently dispatched instruction, and the second instruction is from the second most recently dispatched instruction.

15. The method according to claim 1, further comprising confirming that a thread of a flush command and the first information match before starting the recovery of the mapper with the first information from the scratchpad latch.

16. The method according to claim 1, further comprising: Clearing the first information from the scratchpad latch in response to completion of the first instruction.

17. The method according to claim 1, further comprising: Clearing the first information from the scratchpad latch in response to completion of a flush corresponding to the flush command.

18. The method according to claim 1, wherein the first instruction is from the most recently dispatched instruction.

19. A computer system, comprising a processor core, wherein the processor core is configured to: Save first information in a scratchpad latch, the first information representing a newly dispatched first instruction of the processor core; Save the first information in an entry latch of a save and restore buffer; and In response to receiving a flush command, start the recovery of the mapper of the processor core with the first information from the scratchpad latch without waiting for a comparison of a flush tag of the flush command with the entry latch of the save and restore buffer.

20. A processor core, comprising a dispatcher, a mapper, a save and restore buffer including an entry latch and connected to the mapper via at least one pipeline, and a first register arranged in the at least one pipeline, wherein the processor core is configured to: Save first information in the first register, the first information representing a newly dispatched first instruction from the dispatcher; Save the first information in an entry latch of the save and restore buffer; and In response to receiving a flush command for the processor core, use the first information saved in the first register in the at least one pipeline to recover the first instruction in the mapper.

21. A computer program product, comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 18.