Apparatus, system, chip-containing product, and non-transitory computer readable medium

By introducing preprocessed operand data buffers and registers to the data processing system, the power consumption problem caused by the increase in distance or the need for preprocessing is solved, and efficient reuse of operand data is achieved, and dynamic power consumption is saved from 5% to 10%.

CN120216449APending Publication Date: 2025-06-27ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411901181.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-12-23
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In a data processing system, power consumption increases significantly as the distance between the register file and the execution circuit system increases or when preprocessing operand data is required, especially when operands are reused between instructions, power is wasted.

Method used

A device is provided that includes a preprocessed operand data buffer separate from the register file, can be accessed by an execution circuit system, and stores preprocessed operand data corresponding to a subset of the register file. At the same time, a register is introduced and then used to detect the detection circuit system to detect whether subsequent instructions can use the preprocessing operand data of the previous instructions, and when the chance of reuse is detected, the preprocessing action of the operand data is suppressed.

Benefits of technology

By reducing repeated preprocessing of operand data, significant power consumption is saved, such as 5% to 10% for dynamic power consumption under typical matrix processing workloads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216449A_ABST
    Figure CN120216449A_ABST
Patent Text Reader

Abstract

The invention relates to an apparatus, a system, a chip-containing product, and a non-transitory computer readable medium. Execution circuitry in the apparatus performs a data processing operation on the pre-processed operand data in response to an instruction referencing a given source register of the register file. A buffer separate from the register file stores the pre-processed operand data. When ensuring that there is no intermediate instruction that would result in a write to the reuse source register, register reuse detection circuitry detects a register reuse opportunity that references a subsequent instruction of the reuse source register that is also referenced by a previous instruction for which there is no intermediate instruction that would result in a write to the reuse source register. Preprocessed operand data corresponding to the reuse source register is written to the buffer. In response to detecting the register reuse opportunity, the data processing operation for the subsequent instruction may be performed using pre-processed operand data stored in the buffer, and the pre-processing action is suppressed with respect to the stored operand data from the reuse source register.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art Technical Field

[0001] The present technology relates to the field of data processing. Technical Background

[0003] A data processing apparatus may have: a register file including registers for storing operand data for instructions; and execution circuitry for performing a data processing operation using the stored operand data from a given register referenced as a source register. Summary of the Invention

[0004] At least some examples of the present technology provide an apparatus including:

[0005] A register file including a plurality of registers for storing operand data for instructions;

[0006] Execution circuitry for performing a data processing operation on preprocessed operand data in response to an instruction referencing a given source register, the preprocessed operand data being obtained after performing a preprocessing action using the stored operand data from the given source register of the register file;

[0007] A preprocessed operand data buffer separate from the register file, the preprocessed operand data buffer being accessible by the execution circuitry and configured to store preprocessed operand data corresponding to a subset of the plurality of registers; and

[0008] A register reuse detection circuitry for:

[0009] Detecting a register reuse opportunity for a subsequent instruction when ensuring that no intermediate instruction between a previous instruction and the subsequent instruction referencing a reuse source register also referenced by the previous instruction will cause a write to the reuse source register, and writing the preprocessed operand data corresponding to the reuse source register to the preprocessed operand data buffer for the previous instruction; and

[0010] In response to detecting the register reuse opportunity, controlling the execution circuitry to perform the data processing operation for the subsequent instruction using the preprocessed operand data corresponding to the reuse source register stored in the preprocessed operand data buffer, and suppressing the performance of the preprocessing action for the subsequent instruction with respect to the stored operand data of the reuse source register from the register file.

[0011] At least some examples of the present technology provide an apparatus including:

[0012] A register renaming circuit system for performing register renaming to map an architectural register identifier specified by an instruction to a physical register identifier indicating a corresponding part of a hardware register storage device;

[0013] A register recycling circuit system for determining when a previously allocated physical register identifier is free to be reallocated to a new architectural register identifier specified by an instruction waiting to be renamed;

[0014] A recycling delay circuit system for preventing the actual reallocation of a given physical register identifier during a protection delay period after the register recycling circuit system indicates that the given physical register identifier is free to be reallocated; and

[0015] A protection delay period adjustment circuit system for dynamically adjusting the duration of the protection delay period based on at least one feedback indication.

[0016] At least some examples of the present technology provide a system including:

[0017] An apparatus according to any one of the above two examples, implemented in at least one packaged chip;

[0018] At least one system component; and

[0019] A board,

[0020] wherein the at least one packaged chip and the at least one system component are assembled on the board.

[0021] At least some examples of the present technology provide a chip-containing product including the above system, which is assembled on another board with at least one other product component.

[0022] At least some examples provide a non-transitory computer-readable medium for storing computer-readable code for fabricating an apparatus including:

[0023] A register file including a plurality of registers for storing operand data for instructions;

[0024] An execution circuit system for performing a data processing operation on preprocessed operand data in response to an instruction that references a given source register, the preprocessed operand data being obtained after performing a preprocessing action using the operand data stored in the given source register of the register file;

[0025] A preprocessed operand data buffer that is separate from the register file, the preprocessed operand data buffer being accessible by the execution circuit system and configured to store preprocessed operand data corresponding to a subset of the plurality of registers; and

[0026] A register reuse detection circuit system for:

[0027] Detecting a register reuse opportunity for a subsequent instruction when there is no intervening instruction between a previous instruction and the subsequent instruction that references a reuse source register also referenced by the previous instruction that will cause a write to the reuse source register, and for the previous instruction, writing the preprocessed operand data corresponding to the reuse source register to the preprocessed operand data buffer; and

[0028] In response to detecting the register reuse opportunity, controlling the execution circuit system to perform the data processing operation for the subsequent instruction using the preprocessed operand data corresponding to the reuse source register stored in the preprocessed operand data buffer and suppressing the performance of the preprocessing action for the subsequent instruction with respect to the operand data stored in the reuse source register of the register file.

[0029] Additional aspects, features, and advantages of the present technique will be apparent from the following description of examples read in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Examples illustrating an apparatus having a register reuse detection circuit system for detecting a register reuse opportunity when preprocessed operand data stored in a preprocessed operand data buffer is available for processing a subsequent instruction that references the same source register as an earlier instruction;

[0031] Figure 2 Examples illustrating different ways of addressing array register storage devices within a register file;

[0032] Figure 3 Examples illustrating detecting a register reuse opportunity based on the number of loops separating two instructions that reference the same source register;

[0033] Figure 4Illustrates an example of a register protection delay circuit system for imposing a variable delay on a register overwrite enable action;

[0034] Figure 5 Illustrates an example instruction sequence;

[0035] Figure 6 Illustrates an example of imposing a variable delay on the issue of an instruction that overwrites a given register read by an earlier instruction;

[0036] Figure 7 Illustrates an example of imposing a variable delay on releasing a given physical register for reallocation to the destination architectural register of a newly renamed instruction;

[0037] Figure 8 Illustrates steps for storing a preprocessed operand in a preprocessed operand data buffer and reusing the preprocessed operand for a subsequent instruction;

[0038] Figure 9 Illustrates steps for detecting whether a register reuse opportunity is available;

[0039] Figure 10 Illustrates steps for adjusting the duration of a register protection delay period; and

[0040] Figure 11 Illustrates a system and a chip-containing product. Detailed Description

[0041] Traditionally, in the field of processor design, a register file has generally been considered the fastest-access type of storage device available to an execution circuit system because the register file is located extremely close to the execution circuit system on a chip, such that there is no need to buffer operand data from the register file between the point where the register file is read and the point where the execution circuit system performs a corresponding data processing operation. It is generally assumed that the execution circuit system can perform on operands read directly from the register file.

[0042] However, the present inventors have recognized that there are increasingly data processing systems where the register file may be large and / or may be located at a relatively long distance from the execution circuitry, or where some reformatting or recoding of operand data may be performed after reading the operand data from the register but before the operand data is processed by the execution circuitry. Thus, a relatively significant amount of power may be incurred when performing various preprocessing actions on the stored operand data before the stored operand data read from the register file is used by the execution circuitry. The present inventors have also recognized that there may be a software workload if a reasonable portion of the executed instructions happen to reuse the same operand data that was also used in an earlier instruction. For example, a matrix processing workload may process the same input in several different combinations, such that the likelihood of operand reuse between instructions may be relatively high. In such cases, obtaining the operand data from the register file whenever needed may waste power by repeatedly performing the same preprocessing actions on the same stored data.

[0043] These problems can be solved by providing an apparatus that includes a preprocessed operand data buffer that is separate from the register file, accessible by the execution circuitry, and stores preprocessed operand data corresponding to a subset of the registers of the register file. A register reuse detection circuitry is provided to detect a register reuse opportunity for a subsequent instruction when it is guaranteed that no intervening instruction between a previous instruction and the subsequent instruction that references a reuse source register that is also referenced by the previous instruction will cause a write to the reuse source register, and for which the preprocessed operand data corresponding to the reuse source register is written to the preprocessed operand data buffer. In response to detecting the register reuse opportunity, the register reuse detection circuitry controls the execution circuitry to use the preprocessed operand data corresponding to the reuse source register stored in the preprocessed operand data buffer to perform a data processing operation for the subsequent instruction, and inhibits performing the preprocessing action for the subsequent instruction with respect to the stored operand data of the reuse source register from the register file. By this method, a relatively significant amount of power can be saved. For example, based on simulations of a typical matrix processing workload across different types of workloads, it is estimated that the dynamic power consumption of a processor may be saved by 5% to 10% due to the reuse of preprocessed operand data.

[0044] One challenge with this method can be associated with detecting whether it is guaranteed that no intervening instructions between a previous instruction and a subsequent instruction will result in a write to a reused source register. Some embodiments may provide circuit logic that identifies the source / destination registers referenced by each instruction and keeps track of when an intervening instruction writes to a register referenced as a source register by an earlier instruction, such that a register reuse opportunity for a later instruction that references that register as a source register can be prevented. However, such register comparison logic can be complex to implement and incur a significant circuit area cost. In practice, it may be relatively difficult to implement the register comparison logic, especially where there are multiple different execution units that execute different subsets of instructions and where it is possible that intervening instructions may be processed in an execution unit different from the instruction that references the source register.

[0045] Accordingly, the inventors have recognized that a simpler method of detecting whether it can be guaranteed that there will be no intervening overwrite instructions can be based on an analysis of the number of cycles in the interval between a previous instruction and a subsequent instruction that both reference a reused source register. In particular, when a potential reuse opportunity is identified when the subsequent instruction is identified as referencing the same reused source register as the previous instruction and the operand data preprocessed for that reused source register is stored in a preprocessed operand data buffer, the register reuse detection circuitry can determine whether the number of cycles between the previous instruction that references the reused source register and the subsequent instruction that references the reused source register is less than a threshold number of cycles. In response to determining that the number of cycles between the previous instruction and the subsequent instruction is less than the threshold number, the register reuse detection circuitry can determine that there is a register reuse opportunity for the subsequent instruction because it is guaranteed that no intervening instructions between the previous instruction and the subsequent instruction will result in a write to the reused source register.

[0046] In particular, the threshold number of cycles can correspond to the minimum number of cycles that could occur between two instructions that reference the same source register when the two instructions are separated by an intervening instruction that results in a write to the same source register. Thus, if a previous instruction and a subsequent instruction that both access the reused source register are separated by fewer than the minimum number of cycles, it can be inferred that there cannot be any intervening instructions that write to the same source register and thus the stored operand data for the reused source register has not changed in the register file between the previous instruction and the subsequent instruction, and thus it is safe to use the preprocessed operand data corresponding to the reused source register obtained from the preprocessed operand data buffer.

[0047] Thus, by using a threshold comparison of the number of cycles in the interval between a previous instruction and a subsequent instruction as a means of detecting whether it can be guaranteed that there are no intervening instructions, the circuit area and power cost of implementing the register reuse detection circuitry can be significantly reduced compared to other more direct techniques that compare the registers read / written by different instructions.

[0048] However, in some specific implementations, the processing pipeline microarchitecture for supplying instructions to the execution circuitry can be such that: the minimum number of cycles of the interval between two instructions when two instructions referencing the same register are separated by an intermediate write to that register can be a relatively small number of cycles, thus giving too small a window within which a register reuse opportunity can be available. To address this, a delay can be imposed on the execution of the register overwrite enable action, which register overwrite enable action needs to have been executed such that a later instruction can cause an overwrite of a given register read by an earlier instruction.

[0049] Accordingly, the apparatus can include a register overwrite enable circuitry for performing a register overwrite enable action that needs to have been executed after an earlier instruction reads a given register and before a later instruction has a possibility of causing an overwrite of the given register, where the register overwrite enable action depends on satisfying at least one condition. A register protection delay circuitry can impose a register protection delay period after determining that at least one condition will be satisfied to prevent the register overwrite enable circuitry from performing the register overwrite enable action for at least the register protection delay period after it has been determined that at least one condition has been satisfied.

[0050] This method of deliberately delaying the register overwrite enable action can be seen as highly counterintuitive because it deliberately delays the action required for the instruction to make progress beyond the time when the conditions for progress have been met. It would typically be assumed that once any conditions that need to be satisfied for the execution of the action have been met, any action required by a given instruction should proceed as quickly as possible.

[0051] However, the present inventors have recognized that by preventing the register overwrite enable circuitry from performing the register overwrite enable action for at least the register protection delay period after it has been determined that at least one condition has been satisfied, this will increase the minimum delay between two instructions when two instructions reusing a given register as a source register are separated by an intermediate write to the same register, such that the threshold for the number of cycles for the interval used to compare for detecting a register reuse opportunity can be increased. This means that it is more likely that two instructions reusing the same source register can be detected as being separated by less than the minimum number of cycles, thereby allowing more frequent utilization of the register reuse opportunity by reusing preprocessed operand data to save power.

[0052] The register overwrite enable action can take various forms. In some examples, the register overwrite enable circuitry includes a register renaming circuitry that performs register renaming to map an architectural register identifier specified by an instruction to a physical register identifier that identifies a register of a register file, and the register overwrite enable action includes: after the register retirement circuitry has indicated that a given physical register identifier identifying a given register is free to be remapped to a new architectural register identifier, the register renaming circuitry remaps the given physical register identifier to the destination architectural register of a newly renamed instruction. At least one condition can include: the register retirement circuitry indicates that the released physical register identifier is free to be remapped. Thus, with this method, by delaying the ability to remap a previously allocated physical register identifier to the destination architectural register of a newly renamed instruction, the earliest possible cycle at which a later instruction can overwrite a physical register previously read by an earlier instruction can be delayed. In other words, a further delay is imposed on register retirement outside of the cycle in which the register retirement circuitry additionally indicates that a given physical register will be free to be remapped. This increases the time window within which two references to the same physical register identifier as a source operand can be trusted to actually refer to the same operand data (as opposed to potentially different operand data due to an intervening write to the corresponding physical register). With this method, a relatively long time window can be provided within which reuse opportunities can occur, such that the threshold number of cycles between a previous instruction and a subsequent instruction can be relatively high, thereby giving more opportunities to save power by reusing preprocessed operand data.

[0053] In other examples, the register overwrite enable action includes: issuing a later instruction that overwrites a given register read by an earlier instruction; and at least one criterion includes: the issue circuitry determining that the later instruction is ready to be issued except that a register protection delay period has not elapsed. Although the method can be applied to out-of-order processor implementations, the method can be particularly useful in in-order implementations that do not support register renaming and thus the above-described delayed register retirement method would not be available. By delaying the earliest cycle in which a subsequent overwrite instruction that overwrites a given register read by an earlier instruction can be issued, likewise, the threshold number of cycles between a previous instruction and a subsequent instruction can be higher, thereby giving more opportunities to save power by reusing preprocessed operand data. Although in particular an instruction that writes to a destination register previously used as a source register by an earlier instruction will be delayed in its issue timing, in practice, this is likely to be the case for most instructions processed, because after the initial period in which each register is written for the first time since reset, it is likely that most subsequent instructions will write to registers previously used as source operands. Thus, in some implementations, the delay imposed on the issue of instructions can be imposed on all instructions (which will include instructions that overwrite a given register read by an earlier instruction) without checking whether each instruction is actually overwriting a register read by an earlier instruction.

[0054] Although this method of delaying the register overwrite enable action can contribute to power savings by increasing the opportunity for preprocessed operand data reuse, not all workloads may have a large number of instructions that actually reuse the same source operands as earlier instructions. While for some workloads, delaying the register overwrite enable action can work well and give increased power savings, justifying any small performance loss associated with delaying some instructions due to the delayed register overwrite enable action, for other workloads, if there is little operand reuse between instructions anyway, delaying register retirement or instruction issue can degrade performance without a corresponding benefit in terms of power savings.

[0055] Accordingly, while some embodiments may implement the register protection delay period as a fixed number of cycles, it can be beneficial to provide circuitry that dynamically adjusts the duration of the register protection delay period. Thus, the register protection delay circuitry may dynamically adjust the duration of the register protection delay period based on at least one feedback indication. By considering the feedback collected during the processing of an instruction, the actual needs of the workload being executed can be analyzed and a delay duration more suitable for the particular workload can be selected (e.g., a shorter delay can be selected if there are relatively few register reuse opportunities or if it is found that a longer delay adversely affects performance, and a longer delay can be selected if there are more reuse opportunities and it has been found that the longer delay will not significantly adversely affect performance).

[0056] For example, in response to detecting a risk of forward progress stalling due to an inability to perform a register overwrite enable action for any instruction, the register overwrite enable circuitry may provide a feedback indication to request that the register protection delay circuitry reduce the duration of the register protection delay period. Thus, this provides a safeguard against the implemented register protection delay period adversely affecting performance, as the duration of the register protection delay period can be reduced to reduce the risk of future stalls if forward progress is likely to stall. For example, in an example based on register renaming, the feedback indication provided by the register overwrite enable action can be a measure based on the number of free registers available for remapping in a given cycle. For example, the register renaming circuitry may issue a feedback indication when the number of free registers available for remapping is less than a threshold (or several thresholds may be implemented such that a stronger feedback cue that the register protection delay circuitry should reduce the duration of the register protection delay period is successfully triggered). Similarly, in a method that imposes a delay on the issue of instructions, the feedback can be based on a detection of whether the number of instructions available for issue is less than a threshold number (likewise, in some examples, multiple thresholds can be defined such that when the number of issueable instructions drops below a higher threshold, this triggers a weaker form of feedback for early warning that may optionally be ignored by the register protection delay circuitry, while if the number of issueable instructions available in a given cycle drops below a lower threshold, this can trigger a stronger feedback indication that can require the register protection delay circuitry to reduce the duration of the register protection delay period). It should be understood that there can be a wide variety of ways to implement the feedback mechanism.

[0057] Another type of feedback for adjusting the duration of the register protection delay period can be based on an instruction interval tracking circuitry that tracks the interval between consecutive instructions that read a tracked register for at least one of a plurality of registers. The register protection delay circuitry can adjust the duration of the register protection delay period in response to an interval tracking feedback indication that depends on the intervals tracked by the instruction interval tracking circuitry. By considering the maximum interval seen between consecutive instructions that read the same tracked register, this can help detect the degree of operand data reuse between instructions to help set the duration of the register protection delay period at an appropriate duration given the current workload being executed. The interval tracking feedback indication can depend on a comparison between a threshold set based on the current duration of the register protection delay period and the maximum interval tracked by the instruction interval tracking circuitry for any of the at least one tracked registers. For example, if the threshold is less than the maximum interval, the duration of the register protection delay period can be increased (since some examples of reuse opportunities that could not be exploited due to the register protection delay period currently being too short have been seen), while if the threshold set based on the current duration is greater than the maximum interval, the duration of the register protection delay period can be decreased (to reduce the risk that such a delay period could adversely affect performance when implementing a currently relatively long delay that provides little benefit since no instruction pairs separated by that long delay are observed to access the same register). By this method, a better balance between power savings and performance can be achieved.

[0058] It may be sufficient to track the interval between consecutive read instructions only for a subset of registers to limit the power and circuit area overhead of the instruction interval tracking circuitry. For example, the tracked registers can be the same subset of registers for which a preprocessed operand data buffer stores preprocessed operand data. Matching the tracked registers to the protected registers for which preprocessed operand data is stored in the buffer can be useful because this allows sharing the same set of counters for tracking intervals for the purpose of adjusting the register protection delay period as well as for tracking the interval between consecutive accesses to the same register for the purpose of detecting whether there are register reuse opportunities.

[0059] A selection circuit system can be provided that is configured to select, as a subset of a plurality of registers for which preprocessed operand data is stored in a preprocessed operand data buffer, one or more registers that are referenced as source registers by instructions of at least one predetermined instruction class. For example, instructions of the predetermined class can include certain vector instructions and / or matrix processing instructions that operate on one-dimensional or two-dimensional data arrays. The inventors have recognized that there are certain instruction classes (e.g., outer product instructions, instructions for processing multiple long vectors, multi-vector indexed dot product instructions, etc.) that are more likely to involve substantial reuse of operands between instructions, such that by selecting, based on which registers are referenced by instructions of at least one predetermined class, the registers whose data will be buffered in the preprocessed operand data buffer, significant power savings are more likely to be obtained by reusing the preprocessed operand data.

[0060] Preprocessing actions can include a wide variety of actions taken between reading stored operand data from a register file and processing the operand data by an execution circuit system. When preprocessed operand data is reused from the preprocessed operand data buffer for an input to a data processing operation to be performed by the execution circuit system, any one or more of these actions can be inhibited.

[0061] For example, a preprocessing action that can be inhibited in response to detecting a register reuse opportunity can include reading stored operand data from a given source register of a register file. Reading from the register file can incur a given power cost, so if the register read can be inhibited because the preprocessed operand data can be obtained closer to the execution circuit system in the preprocessed operand data buffer, power can be saved.

[0062] Similarly, in some examples, stored operand data read from a register file can be reused to form operand data for processing (e.g., because the register file supports multiple different access modes by which an addressable register storage can be defined to specify operands for an instruction). This reuse can also incur a certain amount of power, so it can be useful to inhibit the reuse if it is not needed because the preprocessed operand data (which has already undergone reuse in processing an earlier instruction) can be reused for a later instruction.

[0063] However, in some examples, a register reuse detection circuit system can be implemented at a point in a pipeline beyond which register reads (and, if necessary, reuse of operands) from the register file have already occurred. However, there can still be several downstream preprocessing actions that can be inhibited to save power.

[0064] For example, even after reading from the register file, the physical transfer of the stored operand data to the execution circuitry can incur a power cost. The execution circuitry can be located at a long distance from the register file (especially in an example implementing a matrix array register, which can include a relatively large number of storage devices and thus require a physically large area on the chip, increasing the distance between the register file and the execution logic). Thus, when transferred to the execution circuitry, the operand data may need to pass through multiple flip-flops or repeaters, which will consume dynamic power. Power can be saved if the physical transfer can be suppressed for a later instruction because pre-processed operand data that has been previously transferred can be obtained from a buffer local to the execution circuitry. Thus, in some examples, the pre-processing action suppressed in response to detecting a register reuse opportunity includes transferring the stored operand data from a given source register to the execution circuitry.

[0065] Another example can be a pre-processing action that is suppressed in response to detecting a register reuse opportunity: including reformatting the stored operand data to generate pre-processed operand data. There can be some examples where once the operand data has been transferred to the execution circuitry, the execution circuitry can initially perform a certain reformatting (e.g., recoding) of the operand data, which incurs a power cost. For example, for an operation involving multiplication (e.g., which is very common in vector and matrix processing workloads), the reformatting can include performing Booth encoding on the operand data for the multiplication operation. Booth encoding is a technique where the multiplication operands are recoded based on the detection of runs of consecutive 1s in the operand data, which helps reduce the number of partial products that need to be added to obtain the multiplication result. Booth encoding (or other similar operand reformatting applied as a preparatory step for performing arithmetic operations at the execution circuitry) can be relatively power expensive. Thus, power can be saved if pre-processed operand data that has undergone Booth encoding or other reformatting can be reused from the pre-processed operand data buffer.

[0066] Some examples may provide an apparatus that includes: register renaming circuitry that performs register renaming to map architectural register identifiers specified by instructions to physical register identifiers indicative of corresponding portions of a hardware register storage; register retirement circuitry that determines when a previously allocated physical register identifier is available to be reallocated to a new architectural register identifier specified by an instruction waiting to be renamed; retirement delay circuitry that, in response to the register retirement circuitry indicating that a given physical register identifier is available to be reallocated, prevents the given physical register identifier from actually being reallocated during a guard delay period after the register retirement circuitry indicates that the given physical register identifier is available to be reallocated; and guard delay period adjustment circuitry that dynamically adjusts a duration of the guard delay period based on at least one feedback indication.

[0067] The method is highly counterintuitive because it would be assumed that deliberately delaying the timing when a released physical register identifier can be reallocated to a new architectural register identifier would harm processing performance. Typically, processor design focuses on retiring registers as soon as it is safe to do so. However, as described above, the inventors recognized that sometimes delaying the timing when a given physical register identifier can be reallocated to a new architectural register specifier can help save power by increasing the window of opportunity to reuse operand data for preprocessing. By providing dynamic adjustment of the guard delay period, the balance between power savings and performance can be adjusted according to the needs of the software workload being executed, to achieve power savings when possible but mitigate the impact on performance.

[0068] Specific examples will now be described with reference to the accompanying drawings.

[0069] Figure 1 An example of an apparatus 2 including execution circuitry 4 that performs data processing operations in response to instructions is illustrated. A register file 6 is provided that includes a plurality of registers for storing operand data for instructions to be executed by the execution circuitry 4. The operand data stored in the register file 6 may undergo a number of preprocessing actions performed by operand preprocessing circuitry 8 located between the register file 6 and the execution circuitry 4. For example, as Figure 1 shown, the preprocessing actions may include:

[0070] · register read circuitry 10 reads the stored operand data from the register file 6;

[0071] · operand assembly multiplexing circuitry 12 multiplexes the operand data read from portions of the register file to form an operand value to be processed by the execution circuitry 4;

[0072] · Physically transfer the multiplexed operand to the execution circuitry using an operand transfer circuitry 14 (e.g., including repeater elements and flip-flops to ensure that operand data can travel a relatively large distance across the chip within multiple processing cycles); and

[0073] · Reformat the operand data by an operand reformatting circuitry 16 (such as performing Booth encoding on the operand data to prepare for multiplication operations using Booth-encoded operand data by the execution circuitry).

[0074] These preprocessing actions consume a relatively large amount of power, especially when the register file 6 includes an array register as shown in Figure 2 which can be used to represent a two-dimensional data structure, such as a part of a matrix. Figure 2 Each small box within represents a vector element of a given size (e.g., 8-bit, 16-bit, 32-bit, or 64-bit). In the example of Figure 2 32 array registers ZA[0] to ZA

[31] are provided, as shown, which can be addressed in a wide variety of ways:

[0075] - In the examples ZA6H.D[0], ZA0H.H[7], ZA2H.S[5], ZA12H.Q[1], horizontal slices of data from a single array register ZA[i] can be accessed as vector operands, where different sizes of vector elements are indicated by the.D,.H,.S,.Q notations;

[0076] - In the example ZA0V.B

[22] , a vertical slice of data at a given column position

[22] within each of the 32 ZA registers ZA[0] to ZA

[31] can be accessed as a vector operand. Similarly, similar to the horizontal slices shown in Figure 2 it will be possible to provide different data element sizes for vector operands accessed as vertical slices.

[0077] - As shown in the examples ZA7V.D[3], ZA3V.S[4], ZA1V.H[1], ZA8V.Q[0], it is also possible to access a block of elements as a single vector operand, where each part of the block is extracted from a different register among the ZA registers ZA[0] to ZA

[31] . For example, ZA7V.D[3] includes 4 sets of 8 elements selected from column positions [31:24] of each of ZA[7], ZA

[15] , ZA

[23] , ZA

[31] ; ZA3V.S[4] includes from ZA[3],

[0078] Eight sets of 4 elements each selected from column positions [19:16] of ZA[7], ZA

[11] , ZA

[15] , ZA

[19] , ZA

[23] , ZA

[27] , and ZA

[31] ; ZA1V.H[1]

[0079] Sixteen sets of 2 elements each selected from column positions [3:2] of each of the odd-numbered ZA registers, and ZA8V.Q[0] includes two sets of 16 elements each selected from column positions [15:0] of ZA[8] and ZA

[24]

[0080] respectively.

[0081] It should be understood that this is only a subset of the available addressing modes.

[0082] Given the relatively large size of the arrays (and thus the relatively long distances at which portions of the arrays can be positioned relative to the execution circuitry 4) and the wide variety of available addressing modes (requiring relatively complex multiplexing by the operand assembly multiplexing circuitry 12), the preprocessing circuitry 8 can consume a relatively large amount of power to read, multiplex, and transfer operand data. In addition, the nature of the operations applied to this matrix array data such that it can contain a large number of multiplications, such that the Booth encoding overhead for reformatting the operand data can also incur a large amount of power consumption.

[0083] Referring again to Figure 1 , to reduce the power consumption of the device 2, a preprocessed operand data buffer 20 is provided locally to the execution circuitry 4 to store preprocessed operand data for a subset of the registers of the register file 6. A register reuse detection circuitry 22 is provided to detect when there is a register use opportunity, i.e., when a subsequent instruction references the same reuse source register as that referenced by a previously executed instruction, and when it is determined that the preprocessed operand data for that reuse source register is stored in the preprocessed operand data buffer 20 and it is guaranteed that there is no intervening instruction that writes to that source register between the previous instruction and the subsequent instruction (if there is such an intervening write, the preprocessed operand data cannot be reused because the underlying data in the register file may have changed). This takes advantage of the fact that for matrix processing algorithms, there can be many such operand reuses in instruction patterns such as:

[0084] FMOPA ZA0.S,P0 / M,P2 / M,Z0.S,Z2.S

[0085] FMOPA ZA1.S,P0 / M,P3 / M,Z0.S,Z3.S

[0086] FMOPA ZA2.S,P1 / M,P2 / M,Z1.S,Z2.S

[0087] FMOPA ZA3.S,P1 / M,P3 / M,Z1.S,Z3.S。

[0088] Here, the operands Z0, Z1, Z2, and Z3 are all vector operands for more than one instruction (P0 to P3 refer to the associated predicate operands for indicating which elements of the vector operands are active or inactive, where the / M suffix indicates the use of a combined predicate, and where the part of the result corresponding to the inactive vector elements retains the previous value of the corresponding part of the ZA destination register). In this case, the FMOPA instruction is an instruction that performs an outer product and accumulation operation on two one-dimensional vector operands using floating-point arithmetic to generate the corresponding elements of a two-dimensional structure that is written back to the block array ZA. It should be understood that this is not the only type of instruction that can reuse operands in this way.

[0089] Thus, when the register reuse detection circuitry 22 detects a register reuse opportunity, the register reuse detection circuitry 22 controls the execution circuitry to reuse the operand data that has been preprocessed for the reuse source register and stored in the buffer 20 to perform data processing operations for subsequent instructions, and controls the circuitry 8 for operand preprocessing to suppress at least one of the preprocessing actions 10, 12, 14, 16 for the stored operand data for the reuse source register stored in the register file 6. It should be understood that not all of the preprocessing actions 10, 12, 14, 16 can be suppressed. The register reuse detection circuitry 22 can be implemented at various different points within the processing pipeline, and thus depending on the point at which the register reuse detection circuitry 22 is executed, it may not be possible to suppress all preprocessing actions in time. However, even so, when a register reuse opportunity is detected, power can still be saved by suppressing at least one preprocessing action.

[0090] Figure 3 A more detailed example of the apparatus 2 is shown, as in Figure 1 which includes a register file 6, circuitry 8 for operand preprocessing, execution circuitry 4, and a buffer 20 for preprocessed operand data. Figure 3 More details of the circuitry used by the register reuse detection circuitry 22 to detect whether there is a register reuse opportunity are shown. The difficulty in detecting whether it is safe to reuse the preprocessed operand data from an earlier instruction that references the same source register can lie in: detecting whether there is any risk that an intervening instruction between the subsequent instruction that may potentially reuse the preprocessed operand data and the earlier instruction that caused the preprocessed operand data to be written to the buffer 20 may have overwritten the data in that source register. In fact, comparing the source register of the reuse instruction with the destination registers of other potential overwriting instructions can be complex to implement. Therefore, instead,Figure 3 The method shown uses a method of counting the number of cycles based on detecting the intervals between pairs of instructions that reference the same register.

[0091] As Figure 3 shown, the instruction pipeline 30 for supplying instructions to the execution circuitry 4 includes a register reuse detection circuitry 22 at a given stage of the pipeline (e.g., at the register renaming stage, at the instruction issue stage, at a backpressure stage, etc., the backpressure stage being for applying backpressure to issued instructions for execution to request a reduced rate of supplying instructions from earlier stages in the case where the execution circuitry 4 is being flooded).

[0092] In addition, the pipeline includes a pipeline stage 32 that includes a register overwrite enable circuitry 34 that performs at least one register overwrite enable action that is required to have been performed before a subsequent instruction has a possibility of overwriting a register read as a source register by an earlier instruction. For example, as discussed with respect to Figure 6 and Figure 7 in subsequent examples, the register overwrite enable action can be issuing an instruction that designates the same register used as a source register by an earlier instruction as its destination register, or in a particular implementation that supports register renaming, can be remapping a previously allocated physical register to allocate it to a new destination architectural register. Although Figure 3 the stage 32 that includes the register overwrite enable circuitry 34 is shown as a pipeline stage separate from the pipeline stage that includes the register reuse detection circuitry 22, in other examples, the register reuse detection circuitry 22 can be in the same stage of the pipeline as the register overwrite enable circuitry 34. The pipeline stage 32 also includes: a condition evaluation circuitry 36 for evaluating whether at least one condition required as a prerequisite for performing the register overwrite enable action is satisfied; and a delay circuitry 38 for imposing a register protection delay after the cycle in which the at least one condition is satisfied to delay the register overwrite enable circuitry 34 from performing the register overwrite enable action for a duration of at least a register protection delay period after the cycle in which it is first determined that the at least one condition will be satisfied.

[0093] A delay control circuitry 40 is provided to dynamically adjust the duration of the register protection delay imposed by the delay circuitry 38 based on various feedback indications provided by portions of the pipeline 30, as described in more detail below. The delay control circuitry 40 sends an indication of the current delay selected for the register protection delay period to both the delay circuitry 38 and the register reuse detection circuitry 22. Figure 4An example of the delay circuit system 38 is shown, which includes a number of delay elements 50 (e.g., flip-flops), each delay element for applying a one-cycle delay to a signal indicating that at least one condition is satisfied output by the condition evaluation circuit system 36. The multiplexer 52 selects between the delay signal paths sampled after a variable number of delay elements 50. The control input of the multiplexer 52 is the "current delay" control signal output by the delay control circuit system 40. Thus, a variable number of cycles can be selected as the duration of the register protection delay period based on the dynamic control of the delay control circuit system 40.

[0094] Returning to Figure 3 the discussion of, a register selection circuit system 44 is provided, which is used to select a subset of registers for which the preprocessed operand data is to be stored in the preprocessed operand data buffer 20, and this subset of registers will be referred to as protected registers. For example, the selected registers can be up to N registers that have been specified as source registers by instructions of one or more instruction classes (e.g., one of these classes can be the FMOPA instruction described above). The classes can be selected by the system designer as the types of instructions that are expected to relatively likely reuse operand data in an access pattern similar to the access pattern of the FMOPA instruction sequence shown earlier. N can be any number greater than or equal to 1, but in practice does not need to be particularly high (e.g., making N = 2 or 3 may be sufficient so that at most 2 or 3 registers store their operand data in the buffer 20). An instruction interval tracking circuit system 42 is provided, which is used to track an interval metric for each register in the selected subset of registers selected for protection in the operand data buffer 20, and this interval metric indicates the number of processing cycles of the interval seen between an instance of an instruction that references the corresponding register as a source operand and another instance of an instruction that references the register as its source operand. The register selection circuit system 44 can use the interval metric to evaluate whether it is worth continuing to protect a given register or whether it is preferably possible to switch to protecting a different register (e.g., if the interval delay between consecutive references to a previously selected register is too long to give sufficient register reuse opportunities - the longer the interval, the more likely an intermediate write will overwrite the register between consecutive register reads of the register). When attempting to allocate a new register as one of the protected registers, the register selection circuit system 44 can implement a replacement strategy (such as least recently used) to decide which of the previously protected registers should be replaced.

[0095] The register reuse detection circuit system 22 detects a register reuse opportunity when a subsequent instruction is detected as referencing the same source register as a previous instruction, where the reused source register is one of the protected registers currently selected by the register selection circuit system 44, and the preprocessed operand data has been stored in the buffer 20 for that register, and the number of cycles of the interval between the previous instruction and the subsequent instruction is less than a minimum threshold number of cycles, which is selected based on the current latency, and the current latency is selected by the latency control circuit system for the register protection latency period imposed by the latency circuit system 38 during a latency register overwrite enable action. The number of cycles of the interval between the previous instruction and the subsequent instruction can be determined by the register reuse detection circuit system 22 based on the interval count value tracked by the instruction interval tracking circuit system 42 (thus, the same counter for counting the number of cycles since the previous access to each protected register can be shared for both the purpose of determining by the register reuse detection circuit system 22 whether there is a register reuse opportunity and for generating a backend feedback indication that is provided to the latency control circuit system 40 for setting the duration of the register protection latency period imposed by the latency circuit system 38).

[0096] As Figure 5 shown, the method is based on an assessment of the minimum number of cycles between two consecutive reads A and C of the same register X when separated by an intermediate write B to the same register X. For a given pipeline design, in the absence of any additional latency period imposed by the latency circuit system 38, this minimum number of cycles can be a specific number of cycles P, and P will depend on the specific pipeline implementation.

[0097] For example, in one example shown in the following timing example 1, Figure 5 the pipelining of the instructions A, B, C shown can be such that: instruction A takes 3 cycles to execute, and any potential overwrite instruction B that overwrites the register written by A cannot start until instruction A is in its third cycle, but instruction C can start one cycle after A or B. In this example, the minimum number of cycles between A and C in the case of having an intermediate write between A and C can be 3 cycles, and "pipe stage 3" is the last stage where an instruction will potentially be cancelled.

[0098] Timing Example 1:

[0099] Loop Pipeline Stage 1 Pipeline Stage 2 Pipeline Stage 3 0 A 1 A 2 B A 3 C B 4 C B 5 C

[0100] Thus, if it is actually detected that there are only 2 loops between instructions A and C, as in timing example 2, it can be deduced that there cannot be an intermediate write B between instructions A and C that reference the same register, because the latency between A and C is less than the minimum possible latency that could occur in the case of an intermediate write to a reused register.

[0101] Timing Example 2:

[0102] Loop Pipeline Stage 1 Pipeline Stage 2 Pipeline Stage 3 0 A 1 A 2 C A 3 C 4 C

[0103] Thus, by comparing the number of loops in the interval between an earlier instruction / a later instruction that reference the same source register with a threshold, this provides a technique by which, without actually comparing the register identifiers of the destination register of an instruction with the source registers of other instructions, it can be detected that there is guaranteed to be no intermediate write to a reused register, thereby confirming that it is safe to use the preprocessed operand data written to buffer 20 by instruction A when processing instruction C.

[0104] However, in practice, for many pipeline implementations, the minimum latency P between such instructions A and C with an intermediate register write B may be too small to allow many opportunities to reuse preprocessed operand data. For example, in some cases, P can be as small as 1 and even back-to-back instructions may not necessarily guarantee the absence of an intermediate write.

[0105] Thus, as previously described, by using the delay circuit system 38 to impose a delay on the register overwrite enable action, the timing at which overwrite instruction B can be executed relative to A can be delayed, thereby increasing the minimum latency period during which it can be safely assumed that there is no intermediate write between A and C.

[0106] For example, as shown in timing example 3, if an additional 2-loop delay is imposed on instruction B (after the earliest loop at which instruction B could originally be executed) by delaying the register overwrite enable action required to execute B and write to the same register as A, then the minimum number of loops possible between A and C in the presence of an intermediate write becomes 5 loops:

[0107] Timing Example 3:

[0108] Loop Pipeline Stage 1 Pipeline Stage 2 Pipeline Stage 3 0 A 1 A 2 A 3 4 B 5 C B 6 C B 7 C

[0109] This means that in the absence of an intermediate write, there are more opportunities to detect that two instructions that reference the same register can reuse preprocessed operand data between them, as in timing example 4:

[0110] Timing Example 4:

[0111] Loop Pipeline Stage 1 Pipeline Stage 2 Pipeline Stage 3 0 A 1 A 2 A 3 4 C 5 C 6 C

[0112] Now, since the number of cycles between A and C is 4 cycles, which is less than the minimum of 5 cycles possible in the presence of an intervening write to B (longer than the normal minimum of 3 cycles in the case where no additional 2-cycle register protection delay is imposed on the delayed instance of processing B), a register reuse opportunity that would otherwise be unavailable becomes available.

[0113] Thus, in some examples, the threshold number of cycles for the separation of two instructions that reference the same source register (which the register reuse detection circuitry 22 applies to detect the presence of a register reuse opportunity) can be P + V, where:

[0114] · P is a fixed number of cycles corresponding to the minimum delay between instructions A and C in the presence of an intervening write, in the case where the delay circuitry 38 does not introduce any additional delay in the register overwrite enabling actions required for B to execute after A; and

[0115] · V is a variable number of cycles depending on the additional delay imposed by the delay circuitry 38 as selected by the delay control circuitry 40.

[0116] It should be understood that the specific number of cycles shown in the above timing examples is for illustration only, and in practice, the actual delay between instructions can be shorter or longer depending on the specific pipeline implementation. However, the specific number of cycles is used to illustrate the principle that it is possible to check the risk of an intervening write purely based on an assessment of the separation between two reads of the same source register, and why it can be beneficial to impose a delay on the actions required to process an intervening write instruction, which would otherwise be counterintuitive since delaying the instruction can be seen as harmful to performance.

[0117] As Figure 3As shown, the delay control circuit system 40 receives several feedback indications from a portion of the pipeline 30 for dynamically adjusting the current duration of the delay imposed by the delay circuit system 38. A front-end feedback indication from the register overwrite enable circuit system 34 may indicate whether the current delay causes a risk of pipeline stalls due to an insufficient number of instructions for which the register overwrite enable circuit system 34 is targeted. If a front-end feedback indication is received (or if a sufficient number of front-end feedback indications are received within a given time period, depending on the dynamic control mechanism used), the delay control circuit system 40 may reduce the length of the delay period. Additionally, a back-end feedback indication is received from the instruction interval tracking circuit system 42 to give feedback based on the actual interval seen between consecutive instructions that read the same protected register in the protected register selected by the register selection circuit system 44. Based on a comparison between the maximum isolation seen for any of the tracked registers in the tracked registers and a threshold set based on the current delay period set by the delay control circuit system 40, the delay control circuit system 40 may adjust the delay period to be larger when the instruction interval tracking circuit system 42 encounters reuse of the source register between instructions with a larger maximum interval than when it detects reuse of the source register between instructions with a smaller maximum interval. In this way, if the maximum interval between instructions that can reuse preprocessed operand data is relatively small, the delay control circuit system 40 may reduce the length of the delay period because even with a longer delay, the instructions will not benefit from a greater register reuse opportunity, and thus reducing the delay imposed by the circuit system 38 will help improve performance by avoiding unnecessary delays to the instructions. On the other hand, if the tracking performed by the instruction interval tracking circuit system 42 identifies the existence of reuse opportunities that may already exist but are not being utilized (because the current delay is too short and thus the instructions are separated by a greater number of cycles compared to the current threshold used by the register reuse detection circuit system 22 to detect register reuse opportunities), the duration of the delay imposed by the circuit system 38 may be increased.

[0118] Figure 6 and Figure 7 Shows two specific examples of the register overwrite enable circuit system 34.

[0119] In Figure 6 (which includes having been targeted for Figure 3Among all the components described, in this specific example, the register overwrite enable circuit system 34 includes an issue circuit system 60 for issuing instructions for execution by the execution circuit system 4. The condition evaluation circuit system 36 includes a dependency checker 66 that checks for register dependency hazards between instructions to control the issue of a given instruction at a certain timing at which it is predicted that the operands for the instruction will be available in the register file 6 when the instruction reaches the register read stage of the pipeline. The issue circuit system 60 can access an issue queue 62 that queues instructions waiting to be issued and includes an issue control circuit system 64 that selects, from the queued instructions in the queue 62, the instruction to be issued for which the dependency checker 66 has indicated that the dependency check has been completed, such that all the dependency check conditions required for the instruction to be issued have been satisfied. For instructions that do not write to a destination register that is the same as the source register of an earlier instruction, there is no need to delay the issue of the given instruction outside of the loop in which the dependency checker 66 determines that all dependency checks have passed for the given instruction. However, in practice, many instructions can write to a destination register that is the same as the source register of an earlier instruction. At least for those instructions that write to a destination register that is the same as the source register of an earlier instruction (and in some cases, for all instructions, to avoid the overhead of checking whether the destination register of an instruction is the same as the source register of an earlier instruction), the delay circuit system 38 can impose an additional register protection delay outside of the loop in which the dependency checker 66 confirms passing the dependency check for the given instruction to delay the earliest loop in which the given instruction can be issued by at least one additional loop after the loop in which the dependency checker signals that the instruction is ready to be issued (in addition to the register protection delay period imposed by the delay circuit system 38). If the issue circuit system 60 detects that, due to the delay imposed by the delay circuit system 38, there are insufficient instructions ready to be issued in a given loop (but the dependency checker has indicated that there is at least one instruction that has passed its dependency check conditions), then the issue circuit system 60 can issue an issue starvation feedback indication that there is a risk of pipeline stalls due to the insufficient number of instructions available for issue in the given loop (as an Figure 3 example of the front-end feedback indication mentioned in

[0120] Thus, by Figure 6In this method, adding an additional delay during the issuance of instructions for execution (outside the loop where the original instructions would be ready for issuance) can be used to increase the possible minimum loop count between two instructions A and C that reuse the same register X when separated by an intervening instruction B that writes to source register X, in order to increase the window of opportunity to reuse preprocessed operand data. This method can be used in an in-order implementation where register renaming cannot be used as a mechanism to control the delay between instructions A and C.

[0121] Figure 7 Another example of an out-of-order processor suitable for supporting register renaming is shown. In this example, the register overwrite enable circuitry 34 includes: a renaming circuitry 70 that performs register renaming to map an architectural register identifier (ATAG) specified by an instruction defined according to an instruction set architecture to a physical register identifier (PTAG) that identifies a register of the physical register file 6; and a register reclamation circuitry 72 that includes a reclamation condition detection circuitry 75 (an example of the earlier mentioned condition evaluation circuitry 36) that detects when a register reclamation condition is satisfied for a given physical register identifier that has been assigned to a particular architectural register identifier, such that the given physical register identifier can be released to be reassigned to a destination architectural register identifier specified by a newly renamed instruction. The renaming circuitry 70 and the reclamation circuitry 72 use several renaming data structures 74 that include a speculative renaming table (SRT) 76, an architectural renaming table (ART) 78, a free register list 80, and a register commit queue (RCQ) 82.

[0122] SRT 76 designates a speculative set of register mappings associated with the most recent speculative point of program execution reached by rename stage 70. When an instruction reaches the rename circuitry 70, any source architectural register specifiers designated by the instruction are mapped to the physical register identifiers identified in the entries of SRT 76 corresponding to those source architectural register specifiers. Each destination architectural register identifier designated by the instruction is mapped to an available physical register identifier selected from among those physical register identifiers identified as available for reallocation by free list 80. Free list 80 is updated to mark the selected physical register identifier as no longer available for reallocation. In addition, the entry (or entries, if there are multiple destinations in an instruction) of SRT 76 corresponding to the destination architectural register identifier is updated to specify the physical register identifier now mapped to that architectural register identifier. Further, for each new mapping generated, an RCQ entry indicating the mapping between the destination architectural register identifier and the selected physical architectural register identifier is pushed onto RCQ 82, which acts as a first-in, first-out buffer representing an ordered record of changes made to SRT 76 in response to successive instructions. A pointer associated with RCQ 82 tracks the point in the queue corresponding to the commit point of the program flow (the point in the program flow corresponding to the earliest uncommitted instruction). When an instruction commits, one or more entries of RCQ 82 representing the register mappings generated by rename circuitry 70 for any destination registers of the instruction are popped from RCQ 82 and used to update ART 78, which represents the register mapping at the most recent commit point of execution (i.e., the register mappings set for the instruction that are now not speculative). In addition, any physical registers overwritten in ART 78 based on the popped RCQ entries are generally available for reallocation when they are no longer indicated in ART 78. When instructions from the pipeline are flushed due to a mis-speculation event such as a branch mis-prediction, the register state can be rolled back to an earlier point in the program flow by copying the ART 78 contents to SRT 76 and then reconstructing ART 78 based on the RCQ entries associated with the instructions between the commit point represented by ART 78 and the flush point to which program execution must be rolled back to reach before the mis-speculation.

[0123] Accordingly, when the submitted instruction causes the SRT entry that maps architectural register X to physical register Y to be overwritten such that architectural register X is now mapped to physical register Z, physical register Y can generally be freed for reallocation because physical register Y is no longer needed to potentially restore the architectural state after a flush caused by mis-speculation. Accordingly, the reclaim condition detection circuitry 75 analyzes the changes to the ART 78 triggered by the submitted RCQ entry and, when it detects that a given physical register is overwritten in the ART 78, can issue a signal to trigger an update to the "free state" of the given physical register in the free list 80.

[0124] However, in Figure 7 the method of, the latency control circuitry 40 controls the register protection latency circuitry 38 to provide an additional register protection latency period outside of the loop in which the reclaim condition detection circuitry 75 determines that a given register can be freed for reallocation. For example, the free list update signal issued by the reclaim condition detection circuitry 75 can first pass through the latency circuitry 38 before being sent to the free list 80 to trigger an update to the free list. Accordingly, this provides an additional latency of a variable number of loops (where the specific number of latency loops is selected based on the feedback hint provided to the latency control circuitry 40, as discussed earlier) before a given physical register used as a source operand by one instruction can be freed for reallocation to be remapped to the destination architectural register of a later instruction. Since this freeing of the given physical register will be a necessary action after an earlier instruction reads the given physical register and before any subsequent instruction can write to the given physical register, this additional latency will increase the possible minimum number of loops between two instructions that both read the same physical register separated by an intervening instruction that writes to the same physical register. Accordingly, this provides a longer time window during which two instructions that both read the same physical register can be implicitly detected as having no intervening write to the register between these instructions, thus giving a good window of opportunity for reusing the pre-processed operand data stored in the buffer 20.

[0125] In this example, the front-end feedback hint can be issued by the renaming circuitry 70 based on the number of registers indicated as free in the free list 80. If the number of free registers drops below a threshold (e.g., indicating a risk of stalling forward progress due to insufficient available free registers), a hint can be issued indicating that the latency imposed by the latency circuitry 38 can be decreased. In some examples, several alternative thresholds can be defined, e.g.:

[0126] · If the number of free registers drops below a first threshold, the rename circuitry 70 may issue a weak request for a shorter latency, and the latency control circuitry 40 may choose to either ignore the request (without reducing the number of cycles of the register protection latency period) or follow the request (reduce the number of cycles of the register protection latency period).

[0127] · If the number of free registers drops below a second threshold (lower than the first threshold), the rename circuitry 70 may issue a stronger request for a shorter latency period, and the latency control circuitry 40 has no discretion to ignore this stronger type of request.

[0128] In other respects, Figure 7 the operation of Figure 3 is similar to the operation already described for Figure 3 and the components shown with the same reference numerals as in

[0129] Figure 8 illustrates steps performed to enable reuse of preprocessed operand data. At step 100, the register reuse detection circuitry 22 determines whether a register use opportunity is detected for the current instruction. If not, then at step 102, at least one preprocessing action is performed on the stored operand data of a given source register from the register file 6. At step 104, the register reuse detection circuitry 22 determines whether the given source register is currently selected by the register selection circuitry 44 as a protected register. If so, then at step 106, the preprocessed operand data obtained by performing the preprocessing action is buffered in the preprocessed operand data buffer 20. Regardless of whether the given source register is selected as a protected register, at step 108, the execution circuitry 4 performs a data processing operation on the preprocessed operand data.

[0130] On the other hand, if a register reuse opportunity is detected for the current instruction at step 100, then at step 110, execution of at least one preprocessing action on the stored operand data from the given source register is inhibited. For example, as previously mentioned, it may be inhibited: reading the stored operand data from the register file 6, assembling the read operand data into an operand by the multiplexing circuitry 12, transferring the operand across the integrated circuit to the execution circuitry 4, and / or reformatting the operand (e.g., Booth encoding). At step 112, instead, preprocessed operand data corresponding to the given source register is obtained from the preprocessed operand data buffer 20. At step 108, the data processing operation required by the instruction is then performed using the preprocessed operand data from the buffer 20.

[0131] Regarding step 104, the selection of which registers are protected registers is performed by register selection circuitry 44. In one example, when an instruction of one of several instruction classes / types (e.g., outer product, multi-vector length type instruction, multi-vector index dot product, etc.) is detected to have specified a given register as a source register, register reuse detection circuitry 22 may issue a request to make the given register a protected register. There may be a certain maximum number (e.g., 2, 3, or more) of registers that may be referred to as protected registers at a given time. Thus, if the maximum number of registers has been protected, a replacement strategy may be used to determine which of the previously protected registers should be replaced by the newly requested register. For example, the least recently allocated protected register may be replaced by the newly requested protected register. Alternatively, a replacement strategy based on information tracked by instruction gap tracking circuitry 42 may be used (e.g., information regarding the frequency with which each protected register is reused may be maintained to attempt to identify which registers are most useful to protect in the future based on the preferential retention of more frequently reused protected registers).

[0132] Figure 9 Illustrates the steps for determining whether there is a register reuse opportunity at step 100 of Figure 8 At step 120, register reuse detection circuitry 22 determines whether the current instruction references a protected register in which the operand data to be preprocessed is stored in buffer 20 as a source register. Additionally, at step 122, register reuse detection circuitry 22 determines whether the number of loops separating the previous instruction that will reference the source register from the current instruction is less than a threshold number based on the instruction gap tracking provided by instruction gap tracking circuitry 42 for the protected register. The threshold number corresponds to the minimum possible number of loops between instruction A and instruction C that read the same register and are separated by an intervening write B to the same register. For example, the threshold number may correspond to P + V, where P is a certain minimum number of loops between instruction A and instruction C even if delay circuitry 38 does not impose any additional delay, and V is the variable delay imposed by delay circuitry 38 as currently selected by delay control circuitry 40. If the current instruction references a protected register as its source register and the number of loops between the previous instruction that references the same source register and the current instruction is less than the threshold, then at step 124, it is determined that there is a register reuse opportunity. Otherwise, if the current instruction does not reference a protected register as its source register, or the number of loops between the previous instruction and the current instruction is greater than or equal to the threshold, register reuse detection circuitry 22 determines that there is no register reuse opportunity.

[0133] Figure 10Steps for controlling the length of a register protection delay period are illustrated. At step 140, if the delay control circuitry 40 receives a feedback indication from the register overwrite enable circuitry 34 (e.g., the issue circuitry 60 or the rename circuitry 70) that indicates a risk of forward progress stalling due to an inability to perform a register overwrite enable action, then at step 142, the delay control circuitry 40 reduces the duration of the register protection delay period. This reduces the likelihood of future stalls due to an inability to issue an instruction or release a physical register for reallocation.

[0134] At step 144, the delay control circuitry 40 evaluates a condition based on an interval tracking feedback indication from the instruction interval tracking circuitry (the interval tracking feedback indication being based on interval tracking values maintained for each protected register, each interval tracking value indicating the number of cycles separating the most recent read of the register from the previous instruction that read the register). The interval tracking feedback indication is based on a comparison of the maximum interval among the intervals tracked by the instruction interval tracking circuitry 42 for the respective protected registers with a threshold that depends on the current delay period imposed by the delay circuitry 38. If the interval tracking feedback indication indicates that the maximum interval is greater than the threshold based on the current delay, then at step 146, the duration of the register protection delay period is increased to increase the likelihood that two instructions reading the same register separated by the maximum interval number of cycles at a future time can benefit from the reuse of preprocessed operand data. However, if the interval tracking feedback indication indicates that the maximum interval is less than the threshold based on the current delay, then at step 148, the duration of the register protection delay period is decreased because this would mean that the current delay is longer than the delay required to reflect the actual number of cycles between consecutive reads of the same register, such that performance can be improved by reducing the delay imposed on the register overwrite enable action being performed (e.g., thereby allowing a subsequent instruction to be issued earlier, or allowing a previously mapped physical register to be reclaimed for reallocation earlier).

[0135] The concepts described herein may be embodied in a system including at least one packaged chip. The previously described apparatus is implemented in the at least one packaged chip (implemented in one particular chip of the system, or distributed across more than one packaged chip). The at least one packaged chip is assembled on a board together with at least one system component. A chip-containing product may include a system assembled on another board with at least one other product component. The system or chip-containing product may be assembled into a housing or onto a structural support (such as a frame or blade).

[0136] As Figure 11As shown, one or more packaged chips 400 are fabricated by a semiconductor chip manufacturer, where the above-described apparatus is implemented on one chip or distributed across two or more of these chips. In some examples, the chip product 400 fabricated by the semiconductor chip manufacturer may be provided as a semiconductor package that includes a protective housing (e.g., made of metal, plastic, glass, or ceramic) that houses semiconductor devices implementing the above-described apparatus and connectors such as pads, solder balls, or pins for connecting the semiconductor devices to the external environment. In cases where more than one chip 400 is provided, the chips may be provided as separate integrated circuits (provided as separate packages), or may be packaged by the semiconductor provider into a multi-chip semiconductor package (e.g., using an interposer, or by using 3D integration to provide a multi-layer chip product including two or more vertically stacked integrated circuit layers).

[0137] In some examples, a collection of die (i.e., small modular chips having specific functionality) may themselves be referred to as chips. Die may be individually packaged in semiconductor packages and / or packaged together with other die into a multi-die semiconductor package (e.g., using an interposer, or by using 3D integration to provide a multi-layer die product including two or more vertically stacked integrated circuit layers).

[0138] One or more packaged chips 400 are assembled on a board 402 together with at least one system component 404 to provide a system 406. For example, the board may include a printed circuit board. The board substrate may be made of any of a variety of materials, such as plastic, glass, ceramic, or a flexible substrate material such as paper, plastic, or textile material. The at least one system component 404 includes one or more external components that are not part of the one or more packaged chips 400. For example, the at least one system component 404 may include, for example, any one or more of the following: another packaged chip (e.g., provided by a different manufacturer or fabricated at a different process node), an interface module, a resistor, a capacitor, an inductor, a transformer, a diode, a transistor, and / or a sensor.

[0139] Manufacturing a chip-containing product 416, the chip-containing product including a system 406 (including a board 402, one or more chips 400, and at least one system component 404) and one or more product components 412. The product components 412 include one or more additional components that are not part of the system 406. As a non-exhaustive list of examples, the one or more product components 412 may include user input / output devices such as keyboards, touchscreens, microphones, speakers, display screens, haptic devices, etc.; wireless communication transmitters / receivers; sensors; actuators for actuating mechanical motion; thermal control devices; additional packaged chips; interface modules; resistors; capacitors; inductors; transformers; diodes; and / or transistors. The system 406 and the one or more product components 412 may be assembled on an additional board 414.

[0140] The board 402 or the additional board 414 may be disposed on or within a device housing or other structural support (e.g., a frame or a blade) to provide a product that can be manipulated by a user and / or is intended for operational use by a person or a company.

[0141] The system 406 or the chip-containing product 416 may be at least one of an end-user product, a machine, a medical device, a computing or telecommunications infrastructure product, or an automated control system. For example, as a non-exhaustive list of examples, the chip-containing product may be any one of the following: a telecommunications device, a mobile phone, a tablet computer, a laptop computer, a computer, a server (e.g., a rack-mounted server or a blade server), an infrastructure device, networking equipment, a vehicle or other automotive product, an industrial machine, a consumer device, a smart card, a credit card, smart glasses, avionics equipment, a robotic device, a camera, a television, a smart TV, a DVD player, a set-top box, a wearable device, a household appliance, a smart meter, a medical device, a heating / lighting control device, a sensor, and / or a control system for controlling public infrastructure equipment (such as a smart highway or traffic lights).

[0142] The concepts described herein may be embodied in computer-readable code for making an apparatus embodying the described concepts. For example, the computer-readable code may be used in one or more stages of a semiconductor design and fabrication process, including the electronic design automation (EDA) stage, to fabricate an integrated circuit including an apparatus embodying these concepts. The above computer-readable code may additionally or alternatively enable the definition, modeling, simulation, verification, and / or testing of an apparatus embodying the concepts described herein.

[0143] For example, computer-readable code for fabricating an apparatus embodying the concepts described herein may be embodied in code in a hardware description language (HDL) representation that defines these concepts. For example, the code may define a register-transfer-level (RTL) abstraction of one or more logic circuits for defining an apparatus embodying these concepts. The code may define an HDL representation of one or more logic circuits of the apparatus in Verilog, SystemVerilog, Chisel, or VHDL (Very High-Speed Integrated Circuit Hardware Description Language), as well as intermediate representations such as FIRRTL. The computer-readable code may provide a definition of the concepts or other behavioral representations of the concepts embodied using system-level modeling languages such as SystemC and SystemVerilog, which may be interpreted by a computer to enable simulation, functional, and / or formal verification and testing of the concepts.

[0144] Additionally or alternatively, the computer-readable code may define a low-level description of an integrated circuit component embodying the concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSII. One or more netlists or other computer-readable representations of the integrated circuit component may be generated by applying one or more logic synthesis processes to the RTL representation to generate a definition for fabricating an apparatus embodying the present invention. Alternatively or additionally, one or more logic synthesis processes may generate a bitstream to be loaded into a field-programmable gate array (FPGA) from the computer-readable code to configure the FPGA to embody the described concepts. The FPGA may be deployed for the purpose of validating and testing the concepts before fabricating an integrated circuit, or the FPGA may be directly deployed in a product.

[0145] The computer-readable code may include a mixture of code representations for fabricating an apparatus, such as a mixture including one or more of an RTL representation, a netlist representation, or another computer-readable definition used in a semiconductor design and fabrication process for fabricating an apparatus embodying the present invention. Alternatively or additionally, the concepts may be defined in a combination of a computer-readable definition used in a semiconductor design and fabrication process for fabricating an apparatus and computer-readable code defining instructions that will be executed by the defined apparatus once fabricated.

[0146] Such computer-readable code may be disposed on any known transient computer-readable medium (such as a wired or wireless transmission of the code over a network) or non-transient computer-readable medium such as a semiconductor, disk, or optical disc. An integrated circuit fabricated using the computer-readable code may include components such as one or more of a central processing unit, a graphics processing unit, a neural processing unit, a digital signal processor, or other components that individually or jointly embody the concept.

[0147] In the present application, the phrase "configured to..." is used to mean that an element of a device has a configuration capable of performing the defined operation. In this context, "configuration" means an arrangement or manner of interconnection of hardware or software. For example, the device may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not mean that the device element needs to be changed in any way to provide the defined operation.

[0148] In the present application, a list of features starting with the phrase "at least one of..." means that any one or more of those features can be provided individually or in combination. For example, "at least one of the following: [A], [B], and [C]" encompasses any of the following options: only A (without B or C), only B (without A or C), only C (without A or B), a combination of A and B (without C), a combination of A and C (without B), a combination of B and C (without A), or a combination of A, B, and C.

[0149] Although the exemplary embodiments of the present invention have been described in detail herein with reference to the accompanying drawings, it should be understood that the present invention is not limited to those exact embodiments, and various changes and modifications may be made by those skilled in the art without departing from the scope of the present invention as defined by the appended claims.

Claims

1. A device, comprising: a register file comprising a plurality of registers for storing operand data for instructions; execution circuitry for, in response to an instruction referencing a given source register, performing a data processing operation on pre-processed operand data obtained after a pre-processing action has been performed using stored operand data of the given source register from the register file; a pre-processed operand data buffer separate from the register file, the pre-processed operand data buffer being accessible by the execution circuitry and configured to store pre-processed operand data corresponding to a subset of the plurality of registers; and A register reuse detection circuit system, the register reuse detection circuit system is used to: detecting a register reuse opportunity of the subsequent instruction while ensuring that no intervening instruction between a previous instruction and a subsequent instruction referencing a reuse source register also referenced by the previous instruction will result in a write to the reuse source register, and for the previous instruction, preprocessed operand data corresponding to the reuse source register is written to the preprocessed operand data buffer; as well as In response to detecting the register reuse opportunity, controlling the execution circuitry to perform the data processing operation for the subsequent instruction using the preprocessed operand data corresponding to the reused source register stored in the preprocessed operand data buffer, and suppressing execution of the preprocessing action for the subsequent instruction with respect to the stored operand data of the reused source register from the register file.

2. The apparatus of claim 1 , wherein the register reuse detection circuitry is configured to: determining whether a number of cycles between the previous instruction referencing the reused source register and the subsequent instruction referencing the reused source register is less than a threshold number of cycles; and In response to determining that the number of cycles between the preceding instruction and the subsequent instruction is less than the threshold number, it is determined that the register reuse opportunity exists for the subsequent instruction because it is guaranteed that no intervening instructions between the preceding instruction and the subsequent instruction will result in a write to the reused source register.

3. The apparatus of claim 2, wherein the threshold number of cycles corresponds to a minimum number of cycles possible between two instructions that reference the same source register when separated by an intervening instruction that results in a write to the same source register.

4. An apparatus according to any preceding claim, comprising: register overwrite enabling circuitry for performing a register overwrite enabling action that is required to have been performed after a given register is read by an earlier instruction and before a later instruction can cause an overwrite of the given register, wherein the register overwrite enabling action is dependent on at least one condition being satisfied; and A register protection delay circuit system is used to apply a register protection delay period after determining that the at least one condition will be met to prevent the register overwrite enable circuit system from performing the register overwrite enable action within at least the register protection delay period after it has been determined that the at least one condition has been met.

5. The device according to claim 4, wherein: The register overwrite enabling circuitry includes register renaming circuitry for performing register renaming to map an architectural register identifier specified by an instruction to a physical register identifier identifying the register of the register file; The register overwrite enabling action comprises: after the register recycling circuitry has indicated that a given physical register identifier identifying the given register is free to be remapped to a new architectural register identifier, the register renaming circuitry remapping the given physical register identifier to a destination architectural register of a newly renamed instruction; and The at least one condition comprises the register reclamation circuitry indicating that a released physical register identifier is free to be remapped.

6. The device according to claim 4, wherein: The register overwrite enable circuitry includes issue circuitry for issuing instructions for execution; The register overwrite enabling action comprises: issuing the later instruction that overwrites the given register read by the earlier instruction; and The at least one criterion comprises the issue circuitry determining that the later instruction is ready to be issued except that the register protection delay period has not elapsed. 7 . The apparatus of claim 4 , wherein the register protection delay circuitry is configured to dynamically adjust a duration of the register protection delay period based on at least one feedback indication.

8. An apparatus according to claim 7, wherein in response to detecting a risk of forward progress being stalled due to an inability to perform the register overwrite enabling action for any instruction, the register overwrite enabling circuit system is configured to provide a feedback indication to request the register protection delay circuit system to reduce the duration of the register protection delay period.

9. The device according to any one of claims 7 and 8, comprising: an instruction interval tracking circuit system for tracking, for at least one tracked register among the plurality of registers, intervals between consecutive instructions that read the tracked register; wherein: The register protection delay circuitry is configured to adjust the duration of the register protection delay period in response to an interval tracking feedback indication that is dependent on the interval tracked by the instruction interval tracking circuitry.

10. An apparatus as claimed in claim 9, wherein the interval tracking feedback indication is dependent on a comparison between a threshold set based on a current duration of the register protection delay period and a maximum interval tracked by the instruction interval tracking circuitry for any tracked register of the at least one tracked register.

11. The apparatus of any one of claims 9 and 10, wherein the at least one tracked register comprises the subset of registers for which the pre-processed operand data buffer stores the pre-processed operand data.

12. An apparatus according to any one of claims 1 to 11, comprising a selection circuit system, the selection circuit system being used to select one or more registers referenced as source registers by instructions of at least one predetermined instruction class as the subset of the plurality of registers for which the pre-processed operand data buffer stores the pre-processed operand data.

13. The apparatus of any one of claims 1 to 12, wherein the pre-processing action that is suppressed in response to detecting the register reuse opportunity comprises: The stored operand data is read from the given source register of the register file.

14. The apparatus of any one of claims 1 to 13, wherein the pre-processing action that is suppressed in response to detecting the register reuse opportunity comprises: The stored operand data is transferred from the given source register to the execution circuitry.

15. The apparatus of any one of claims 1 to 14, wherein the pre-processing action that is suppressed in response to detecting the register reuse opportunity comprises: The stored operand data is reformatted to generate the pre-processed operand data.

16. The apparatus of claim 15, wherein the reformatting comprises: The operand data is Booth encoded for multiplication operations.

17. A device, comprising: register renaming circuitry to perform register renaming to map an architectural register identifier specified by an instruction to a physical register identifier indicative of a corresponding portion of a hardware register storage device; register reclamation circuitry for determining when a previously allocated physical register identifier is free to be reallocated to a new architectural register identifier specified by an instruction awaiting renaming; reclamation delay circuitry for, in response to the register reclamation circuitry indicating that a given physical register identifier is free to be reallocated, preventing actual reallocation of the given physical register identifier during a protection delay period after the register reclamation circuitry indicates that the given physical register identifier is free to be reallocated; and A protection delay period adjustment circuitry is provided for dynamically adjusting a duration of the protection delay period based on at least one feedback indication.

18. A system, comprising: The device according to any one of claims 1 to 17, wherein the device is implemented in at least one packaged chip; at least one system component; and plate, The at least one packaged chip and the at least one system component are assembled on the board.

19. A chip-containing product, comprising the system according to claim 18, wherein the system and at least one other product component are assembled on a separate board.

20. A computer program code for producing an apparatus comprising: a register file comprising a plurality of registers for storing operand data for instructions; execution circuitry for, in response to an instruction referencing a given source register, performing a data processing operation on pre-processed operand data obtained after a pre-processing action has been performed using stored operand data of the given source register from the register file; a pre-processed operand data buffer separate from the register file, the pre-processed operand data buffer being accessible by the execution circuitry and configured to store pre-processed operand data corresponding to a subset of the plurality of registers; and A register reuse detection circuit system, the register reuse detection circuit system is used to: detecting a register reuse opportunity of the subsequent instruction while ensuring that no intervening instruction between a previous instruction and a subsequent instruction referencing a reuse source register also referenced by the previous instruction will result in a write to the reuse source register, and for the previous instruction, preprocessed operand data corresponding to the reuse source register is written to the preprocessed operand data buffer; as well as In response to detecting the register reuse opportunity, controlling the execution circuitry to perform the data processing operation for the subsequent instruction using the preprocessed operand data corresponding to the reused source register stored in the preprocessed operand data buffer, and suppressing execution of the preprocessing action for the subsequent instruction with respect to the stored operand data of the reused source register from the register file.

21. A computer readable storage medium for storing the computer program code according to claim 20.