An adaptive bandwidth scheduling method and device
By adaptively adjusting the number of micro-operations read from the operation queue port, the matching problem between the processor front-end and back-end was solved, thereby improving processor energy efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HYGON INFORMATION TECH CO LTD
- Filing Date
- 2026-03-27
- Publication Date
- 2026-06-12
AI Technical Summary
In traditional processor designs, the reading and dispatching of front-end microoperations cannot match that of the back-end, resulting in hardware energy consumption and an imbalance between performance and power consumption, which limits the energy efficiency improvement of high-performance processors.
By using an adaptive bandwidth scheduling method, the number of micro-operations read by the Operation Queue (UOPQ) port is dynamically adjusted, the number of micro-operations effectively dispatched by the DISPATCH module is controlled, excessive micro-operation dispatch is avoided, and the energy efficiency of the execution unit is improved.
Without reducing pipeline efficiency, the power consumption of the UOPQ and DISPATCH modules was reduced, thus improving the processor's energy efficiency.
Smart Images

Figure CN122195513A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to an adaptive bandwidth scheduling method and apparatus. Background Technology
[0002] In CPU microarchitecture, the pipeline is mainly divided into the front end and the back end. As the types and number of back-end execution units continue to increase, the execution capability of the processor's back end continues to improve, and the bandwidth demand for front-end instruction dispatch is growing. However, in traditional designs, the reading and dispatching of front-end microoperations can no longer match that of the back end. Undispatched microoperations will cause hardware energy consumption, leading to an imbalance between performance and power consumption, which restricts the improvement of energy efficiency of high-performance processors.
[0003] Improving processor energy efficiency is an urgent problem that needs to be solved. Summary of the Invention
[0004] To address the problems in the prior art, embodiments of this specification provide an adaptive bandwidth scheduling method and apparatus that can solve the problem of low energy efficiency caused by the mismatch between the number of micro-operations in the micro-operation queue (UOPQ) port and the number of micro-operations in the DISPATCH module in the processor front end. By controlling the effective number of micro-operations in the DISPATCH module, excess micro-operations are prevented from being dispatched to the execution unit, thereby improving the energy efficiency of the execution unit.
[0005] This specification provides an adaptive bandwidth scheduling method, including: Get the first number of micro-operations read from the UOPQ port of the operation queue, and the second number of micro-operations effectively dispatched by the DISPATCH module; The number of effective read micro-operations of the UOPQ port is adjusted based on the difference between the first and second quantities.
[0006] As a further aspect of this specification, the effective dispatch of micro-operations by the DISPATCH module includes: the micro-operations that are effectively dispatched; the number of micro-operations that the UOPQ port effectively reads includes: the number of micro-operations that the UOPQ port corresponding to the micro-operations that can be effectively dispatched by the DISPATCH module can read.
[0007] As a further aspect of this specification, the second number of micro-operations effectively dispatched by the DISPATCH module includes: the second number of first-type micro-operations effectively dispatched by the DISPATCH module, wherein the difference is the difference between the first number of first-type micro-operations read by the UOPQ port and the second number of first-type micro-operations effectively dispatched by the DISPATCH module.
[0008] As a further aspect of this specification, adjusting the number of effective read micro-operations of the UOPQ port based on the difference between the first and second quantities further includes: When the difference continues to exceed the first threshold, the number of effective read micro-operations of the UOPQ port is reduced. When the difference remains below the second threshold, the number of valid read micro-operations on the UOPQ port is increased.
[0009] As a further aspect of this specification, the differences include: DELTA = First quantity - Second quantity; Where DELTA is the difference between the first quantity and the second quantity.
[0010] As a further aspect of this specification, when the difference continues to exceed the first threshold, reducing the number of valid read micro-operations on the UOPQ port further includes: Analyze the reasons why the difference continues to exceed the first threshold. If the limitation is due to a static constraint, then the number of valid read micro-operations on the UOPQ port is reduced according to the first rule; If the limitation is due to dynamic constraints, then the number of valid read micro-operations on the UOPQ port is reduced according to the second rule.
[0011] As a further aspect of this specification, the static constraints include: hardware limitations of the DISPATCH module and the pipeline back-end execution unit; The dynamic constraints include the working status of the DISPATCH module and the pipeline back-end execution unit.
[0012] As a further aspect of this specification, the specific reasons for the limitation that the difference continues to exceed the first threshold include: During the process of sending the micro-operations read from the UOPQ port to the DISPATCH module through static and dynamic constraint circuits. When the number of micro-operations obtained by the dynamic constraint circuit is less than the number of micro-operations read by the UOPQ port, the reason for the limitation that the difference is continuously greater than the first threshold is static constraint. or, When the number of micro-operations acquired by the static constraint circuit is greater than the number of micro-operations read by the UOPQ port, the reason for the limitation that the difference continues to be greater than the first threshold is static constraint. or, When the number of micro-operations output by the dynamic constraint circuit is less than the number of micro-operations obtained by the dynamic constraint circuit, the reason for the limitation that the difference continues to be greater than the first threshold is dynamic constraint.
[0013] As a further aspect of this specification, the specific reasons for the limitation that the difference continues to exceed the first threshold include: Obtain the available resource tokens for the execution units corresponding to the first type of micro-operation effectively dispatched by the DISPATCH module; If the change of the TOKEN of the execution unit within a predetermined time is less than a stable threshold, then the reason for the limitation that the difference continues to be greater than the first threshold is a static constraint. If the change of the TOKEN of the execution unit within a predetermined time period is greater than a stable threshold, then the reason for the limitation that the difference continues to be greater than the first threshold is dynamic constraint.
[0014] As a further aspect of this specification, the reduction of the number of valid UOPQ port read micro-operations according to the first rule further includes: The number of valid read micro-operations on the UOPQ port is reduced to a specified number based on the hardware limitations. The second rule for reducing the number of valid read micro-operations on the UOPQ port further includes: Based on the operating state, the number of effective read micro-operations of the UOPQ port is reduced within a certain number of clock cycles.
[0015] As a further aspect of this specification, adjusting the number of effective read micro-operations of the UOPQ port includes at least one of the following methods: The number of UOPQ ports can be adjusted by using the UOPQ port enable signal, thereby regulating the number of effective read micro-operations by the UOPQ port; or, By controlling the number of micro-operations written to the UOPQ port by the pipeline front end, and waiting for all micro-operations written to the UOPQ port, the number of UOPQ ports is adjusted, thereby regulating the number of effective micro-operations read by the UOPQ port; or, By controlling the number of micro-operations written to the UOPQ port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ port, and after the micro-operations in the UOPQ have been dispatched, the number of UOPQ ports is adjusted, thereby adjusting the number of micro-operations that the UOPQ port can effectively read; or, By redirecting the REDIRECT action to trigger the FLUSH operation at the front end of the pipeline, the number of UOPQ ports is adjusted, thereby adjusting the number of effective read micro-operations of the UOPQ ports.
[0016] As a further aspect of this specification, after reducing the number of effective read micro-operations on the UOPQ port, it also includes: Monitor the number of micro-operations effectively dispatched by the DISPATCH module; If the number of first-type micro-operations effectively dispatched by the DISPATCH module remains below a preset threshold, the number of micro-operations effectively read by the UOPQ port will be increased.
[0017] As a further aspect of this specification, after reducing the number of effective read micro-operations on the UOPQ port, it also includes: Monitor changes in processor performance IPC; As the DISPATCH module completes the dispatch of micro-operations, if the IPC remains below a preset threshold, the number of effective micro-operations read by the UOPQ port will be increased.
[0018] As a further aspect of this specification, after reducing the number of effective read micro-operations on the UOPQ port, it also includes: Monitor the number of execution unit tokens at the back end of the processor pipeline; As the DISPATCH module completes the dispatch of micro-operations, if the number of TOKENs corresponding to a certain type of micro-operation continues to exceed a preset threshold, and if there are still micro-operations of that type in the DISPATCH module, then the number of micro-operations that are effectively read by the UOPQ port is increased.
[0019] This specification also provides an adaptive bandwidth scheduling device, comprising: The acquisition unit is configured to acquire the first number of micro-operations read from the operation queue UOPQ port, and the second number of micro-operations effectively dispatched by the instruction dispatch DISPATCH module. The decision-maker is configured to adjust the number of effective read micro-operations of the UOPQ port based on the difference between the first number and the second number.
[0020] This specification also provides an instruction dispatching unit that executes the adaptive bandwidth scheduling method described above.
[0021] This specification also provides a processor, including an instruction dispatch unit for performing the above-described methods.
[0022] This specification also provides a computer device, including a memory and a computer program stored in the memory and executable on a processor, including the processor described above.
[0023] This specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the above-described method.
[0024] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.
[0025] By utilizing the embodiments in this specification, the power consumption of the UOPQ module can be reduced without decreasing pipeline efficiency, and the power consumption of the DISPATCH module can also be reduced, thereby improving the processor's energy efficiency. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 The diagram shown is a flowchart of an adaptive bandwidth scheduling method according to an embodiment of this specification. Figure 2 The diagram shown is a schematic of the UOPQ port adjustment circuit in an embodiment of this specification. Figure 3 The diagram shown is a schematic diagram of a blocking write operation in an embodiment of this specification; Figure 4 The diagram shown is a schematic of how to clear the front end of the pipeline to adjust the number of UOPQ ports according to an embodiment of this specification; Figure 5 The diagram shown is a structural schematic of the adaptive bandwidth scheduling device according to an embodiment of this specification. Figure 6 The diagram shown is a structural schematic of the instruction dispatch unit in an embodiment of this specification. Figure 7 The diagram shown is a structural schematic of the intermediate processing module in an embodiment of this specification. Figure 8 The diagram shown is a schematic representation of a computer device provided in an embodiment of this specification. Detailed Implementation
[0028] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0029] like Figure 1 The diagram shows a flowchart of an adaptive bandwidth scheduling method according to an embodiment of this specification. This diagram describes how, in the pipeline front-end of a high-performance processor, particularly during the process of reading micro-operations from the Unified Operation Queue (UOPQ) to the Dispatch port for micro-operation dispatch, the number of micro-operations read from the UOPQ port is dynamically adjusted to match the number of micro-operations at the pipeline front-end and back-end. This reduces ineffective toggle switching of micro-operations between instruction dispatches, reduces the power consumption of execution units caused by excessive instruction dispatch, avoids hardware energy consumption, and improves processor energy efficiency. The method specifically includes: Step 101: Obtain the first number of micro-operations read from the UOPQ port and the second number of micro-operations effectively dispatched by the DISPATCH module; Step 102: Adjust the number of effective read micro-operations of the UOPQ port based on the difference between the first number and the second number.
[0030] The methods described in this specification can reduce the power consumption of the UOPQ module, the power consumption of the DISPATCH module, and the power consumption of the execution unit without reducing pipeline efficiency, thereby improving the energy efficiency of the processor.
[0031] As an embodiment of this specification, the effective dispatch of micro-operations by the DISPATCH module includes: the micro-operations that are effectively dispatched; the number of micro-operations that are effectively read by the UOPQ port includes: the number of micro-operations that can be read by the UOPQ port corresponding to the micro-operations that can be effectively dispatched by the DISPATCH module.
[0032] In this embodiment, the number of micro-operations read by the UOPQ port can be determined based on the hardware resources of the UOPQ port. For example, if the width of the UOPQ port is 8, it means that 8 micro-operations (SLOT-0~SLOT-7) in the queue can be read at once. However, since the number of micro-operations dispatched by the DISPATCH module is limited by the working status of the downstream execution unit (e.g., the computing unit ALU), it may only be able to dispatch 4 micro-operations to the corresponding downstream execution unit. Alternatively, the dispatching capability of the DISPATCH module may also be limited by its own hardware resources (e.g., the number of channels, the number of ports, etc.), and it may only be able to dispatch 4 micro-operations to the downstream execution unit. As a micro-operation reading end, the UOPQ port (e.g., SLOT-0~SLOT-3) can only actually read 4 micro-operations that can be effectively dispatched by the DISPATCH module. That is to say, the effectively dispatched micro-operations refer to the micro-operations that can be executed by the downstream execution unit in a short period of time. The busy / idle status and resource limitations of the DISPATCH module, as well as the busy / idle status and resource limitations of the existing execution unit, will affect the number of micro-operations that the DISPATCH module can effectively dispatch. In the embodiments of this specification, the number of micro-operations read by the UOPQ port corresponding to a micro-operation that can be effectively dispatched by the DISPATCH module is referred to as the number of micro-operations effectively read by the UOPQ port. Since a micro-operation read by the UOPQ port may be parsed into more micro-operations, the number of micro-operations effectively read by the UOPQ port is less than or equal to the number of micro-operations effectively dispatched by the DISPATCH module. By adjusting the number of UOPQ ports, the number of micro-operations effectively read by the UOPQ port can be matched to the number of micro-operations effectively dispatched by the downstream DISPATCH module. This avoids reading micro-operations in the UOPQ that far exceed the processing capacity of the downstream DISPATCH module or execution unit, avoids invalid inversion transitions in the UOPQ, and avoids invalid transitions, invalid circuit buffering, and movement in the downstream DISPATCH module and execution unit module, thereby effectively reducing device power consumption.
[0033] In this embodiment, the UOPQ port read micro-operations include all micro-operations read by the UOPQ port. As described in the previous embodiment, the UOPQ port may read 8 micro-operations, but due to the limited dispatching capability of the DISPATCH module, it may only be able to successfully dispatch 4 of them. These 4 successfully dispatched micro-operations are called UOPQ port valid read micro-operations. The UOPQ port valid read micro-operations include the micro-operations read by the UOPQ port corresponding to the micro-operations that can be effectively dispatched by the DISPATCH module.
[0034] In another embodiment, when the DISPATCH module dispatches micro-operations, the type of the effectively dispatched micro-operation can be obtained. The second quantity of effectively dispatched micro-operations by the DISPATCH module is the number of different types of micro-operations. For example, the number of fixed-point arithmetic micro-operations successfully dispatched to fixed-point arithmetic execution units by the DISPATCH module, and the number of floating-point arithmetic micro-operations successfully dispatched to floating-point arithmetic execution units by the DISPATCH module. Each different type of micro-operation has a different second quantity. The difference between the second quantity and the first quantity of different types of micro-operations can be the same or different, and the threshold corresponding to the difference between different types of micro-operations can be the same or different. The type of micro-operation can, for example, include: floating-point arithmetic micro-operations, fixed-point arithmetic micro-operations, memory access micro-operations, etc. The type of micro-operation can also be further subdivided; for example, floating-point arithmetic micro-operations can further include specific sub-operation types such as addition, multiplication, and vector operations.
[0035] When the available resources (tokens) of the execution unit corresponding to the effectively dispatched micro-operation are more or less, for example, when the available resources of the execution unit corresponding to the floating-point operation type are more, by adjusting the number of micro-operations read by the UOPQ port (for example, adjusting the number of UOPQ ports), the number of micro-operations effectively read by the UOPQ port for a certain type can be matched with the number of tokens of the downstream execution unit that processes the corresponding type of micro-operation, thereby improving the processor's processing efficiency and reducing the power consumption of the execution unit.
[0036] As an embodiment of this specification, the second number of micro-operations effectively dispatched by the DISPATCH module includes: the second number of first-type micro-operations effectively dispatched by the DISPATCH module, wherein the difference is the difference between the first number of first-type micro-operations read by the UOPQ port and the second number of first-type micro-operations effectively dispatched by the DISPATCH module.
[0037] In this embodiment, the difference may be the difference between the first number of micro-operations read by the UOPQ port and the second number of micro-operations effectively dispatched by the DISPATCH module, or it may be the difference between the first number of micro-operations read by the UOPQ port and the second number of a certain type of micro-operations effectively dispatched by the DISPATCH module. This type may be the micro-operation type with the largest number of micro-operations effectively dispatched by the DISPATCH module, the micro-operation type with the smallest number of micro-operations effectively dispatched by the DISPATCH module, or a specified micro-operation type among all types of micro-operations effectively dispatched by the DISPATCH module.
[0038] As an embodiment of this specification, adjusting the number of effective read micro-operations of the UOPQ port based on the difference between the first quantity and the second quantity further includes: When the difference continues to exceed the first threshold, the number of effective read micro-operations of the UOPQ port is reduced. When the difference remains below the second threshold, the number of valid read micro-operations on the UOPQ port is increased.
[0039] In this embodiment, the difference between the first quantity and the second quantity can be obtained by the difference between the first quantity and the second quantity, for example, DELTA = first quantity - second quantity, where DELTA is the difference between the first quantity and the second quantity; it can also be obtained by calculating the proportion of the second quantity to the first quantity, or by other specific calculation methods to obtain the difference between the two. The goal is to adjust the number of micro-operations effectively read by the UOPQ port based on the difference between the number of micro-operations read by the UOPQ port and the number of micro-operations effectively dispatched by the DISPATCH module.
[0040] For example, if the number of micro-operations read by the UOPQ port significantly exceeds the number of micro-operations effectively dispatched by the DISPATCH module over several consecutive clock cycles, it indicates that the DISPATCH module or a certain type of downstream execution unit is subject to static or dynamic constraints, causing congestion in micro-operation dispatch. If the UOPQ port continues to maintain a high read speed for micro-operations, these read micro-operations cannot be dispatched to the corresponding type of backend execution unit in a timely manner, and instead undergo transitions within the UOPQ and simultaneously within the DISPATCH module, resulting in high processor power consumption. In the embodiments of this specification, when the number of micro-operations read by the UOPQ port significantly exceeds the number of micro-operations effectively dispatched by the DISPATCH module, reducing the number of micro-operations effectively read by the UOPQ port can alleviate congestion in micro-operation dispatch, thereby reducing power consumption.
[0041] Similarly, when the number of effective read micro-operations of the UOPQ port is reduced through the aforementioned steps, and the number of effective read micro-operations of the UOPQ port is significantly less than the number of effective dispatch micro-operations of the DISPATCH module in multiple consecutive clock cycles, the number of effective read micro-operations of the UOPQ port can be increased to improve the utilization of device hardware resources and improve processor efficiency.
[0042] The number of valid read micro-operations on the UOPQ port can be reduced or increased. This can be done by reducing the number of read micro-operations on the UOPQ port to a minimum, such as reducing it to 0 or 1; or by increasing the number of read micro-operations on the UOPQ port to a maximum, such as increasing it to the maximum value of the UOPQ port.
[0043] In other embodiments of this specification, reducing or increasing the number of micro-operations read by the UOPQ port may also involve reducing or increasing the number of micro-operations read by the UOPQ port to a specified number, monitoring the number of validly dispatched micro-operations by the DISPATCH module during subsequent clock cycles, and gradually adjusting the number of micro-operations read by the UOPQ port until the number of micro-operations read by the UOPQ port matches the number of validly dispatched micro-operations by the DISPATCH module.
[0044] As an embodiment of this specification, when the difference continues to exceed the first threshold, reducing the number of effective read micro-operations of the UOPQ port further includes: Analyze the reasons why the difference continues to exceed the first threshold. If the limitation is due to a static constraint, then the number of valid read micro-operations on the UOPQ port is reduced according to the first rule; If the limitation is due to dynamic constraints, then the number of valid read micro-operations on the UOPQ port is reduced according to the second rule.
[0045] In this embodiment, when the difference is consistently greater than the first threshold, it indicates that the DISPATCH module is experiencing congestion in dispatching micro-operations. The cause of the current dispatch congestion can be determined using the following method: During the process of sending the micro-operations read from the UOPQ port to the DISPATCH module through static and dynamic constraint circuits. When the number of micro-operations obtained by the dynamic constraint circuit is less than the number of micro-operations read by the UOPQ port, the reason for the limitation that the difference is continuously greater than the first threshold is static constraint. or, When the number of micro-operations acquired by the static constraint circuit is greater than the number of micro-operations read by the UOPQ port, the reason for the limitation that the difference continues to be greater than the first threshold is static constraint. or, When the number of micro-operations output by the dynamic constraint circuit is less than the number of micro-operations obtained by the dynamic constraint circuit, the reason for the limitation that the difference continues to be greater than the first threshold is dynamic constraint.
[0046] In this embodiment, reference may also be made to Figure 7As shown, after the UOPQ port reads multiple micro-operations, the static constraint circuit dispatches and controls these micro-operations based on the number of subsequent execution units in the CPU microarchitecture pipeline. For example, if the UOPQ port reads eight addition micro-operations through eight read ports, but the CPU microarchitecture pipeline only has four addition execution units, the static constraint circuit will send only four addition micro-operations to the subsequent dynamic constraint circuit, which will then output them to the corresponding addition execution units. Based on the number of micro-operations obtained by the dynamic constraint circuit, the constraints of the static constraint circuit can be analyzed, thus determining that the reason for the persistent difference exceeding the first threshold is the static constraint.
[0047] Furthermore, the dynamic constraint circuit can dynamically control and dispatch micro-operations to the downstream execution units of the pipeline based on factors such as the working status of the execution units. For example, if the dynamic constraint circuit acquires four addition micro-operations, and after acquiring the working status of the corresponding addition execution units (e.g., if only two addition execution units are idle), the addition micro-operations are sent to these two addition execution units via the DISPATCH module. Based on the number of micro-operations acquired and output by the dynamic constraint circuit, the constraint conditions of the downstream execution units can be analyzed, thereby determining that the reason for the persistent difference exceeding the first threshold is dynamic constraint.
[0048] In other embodiments, when the difference between the second number of micro-operations effectively dispatched by the DISPATCH module and the first number of micro-operations read by the UOPQ port is the same as the difference between the fourth number of micro-operations acquired by the dynamic constraint circuit and the first number of micro-operations read by the UOPQ port, the limitation can be attributed to static constraints; otherwise, the limitation can be attributed to dynamic constraints. Alternatively, when the difference between the fourth number of micro-operations acquired by the dynamic constraint circuit and the second number of micro-operations effectively dispatched by the DISPATCH module is 0, the limitation is attributed to static constraints; otherwise, the limitation can be attributed to dynamic constraints. Or, when the third number of micro-operations acquired by the static constraint circuit is greater than the first number of micro-operations read by the UOPQ port, the limitation is attributed to static constraints.
[0049] In other embodiments, the cause of the current dispatch congestion can also be determined by the following methods: Obtain the available resource tokens for the execution units corresponding to the first type of micro-operation effectively dispatched by the DISPATCH module; If the change of the TOKEN of the execution unit within a predetermined time is less than a stable threshold, then the reason for the limitation that the difference continues to be greater than the first threshold is a static constraint. If the change of the TOKEN of the execution unit within a predetermined time period is greater than a stable threshold, then the reason for the limitation that the difference continues to be greater than the first threshold is dynamic constraint.
[0050] For example, if the number of tokens in the execution unit is detected to be consistently below the stable threshold (e.g., the number of available tokens in the floating-point unit is 3 for 5 cycles, which is less than the stable threshold of 5), then the limitation is due to static constraints. If the number of tokens in the execution unit is detected to change frequently in a short period of time, and the change is greater than the stable threshold (e.g., the number of available tokens in the floating-point unit is 3 for 2 cycles, and then the number of tokens in the floating-point unit becomes 5 for the next 3 cycles, with the change exceeding the stable threshold), then the limitation is due to dynamic constraints.
[0051] In some embodiments, dispatch congestion can be categorized into two types: static constraints and dynamic constraints. Static constraints include: hardware limitations between the DISPATCH module and the pipeline back-end execution units; for example, a limited number of execution units for a certain type of micro-operation (e.g., 4 fixed-point execution units), a single-cycle processing limit for execution units (e.g., a floating-point execution unit can only process 2 UOPs per cycle), port bandwidth limitations (e.g., the port bandwidth limit for the DISPATCH module to transmit UOPs to execution units), and resource binding limitations (e.g., specific operations can only be bound to specific register resources; when register resources are limited, the corresponding micro-operation cannot be dispatched). Specific static constraints may include: a single-cycle processing limit of 4 UOPs for a fixed-point execution unit (EX), and a single-cycle processing limit of 2 UOPs for a floating-point execution unit (FP) (due to hardware complexity); the Memory Execution Unit (LSU) supports memory access UOPs but not arithmetic UOPs. The dynamic constraints include: the working state between the DISPATCH module and the pipeline back-end execution unit; for example, the resource status of the issue queue (such as whether there are free slots in the issue queue), the hardware resource occupancy status (such as the multiplier of the floating-point execution unit is processing the previous UOP and cannot receive new floating-point multiplication UOPs), and the data dependency status (such as the mutual influence between the execution processes of UOPs); specific dynamic constraints may include: the long execution cycle of floating-point UOPs (3 cycles) can easily lead to congestion of the floating-point execution unit queue (frequent dynamic constraints); the short execution cycle of fixed-point UOPs (1 cycle) results in less congestion of the fixed-point execution unit queue; the memory access execution unit is easily affected by the cache status (dynamic constraints are mostly cache port occupancy).
[0052] When the difference is significant, it may be due to congestion in micro-operation dispatch caused by static constraints or by dynamic constraints. When the congestion is caused by static constraints, the number of micro-operations read by the UOPQ port can be reduced according to the first rule; when the congestion is caused by dynamic constraints, the number of micro-operations read by the UOPQ port can be reduced according to the second rule.
[0053] As an embodiment of this specification, reducing the number of valid read micro-operations of the UOPQ port according to the first rule further includes: The number of valid read micro-operations on the UOPQ port is reduced to a specified number based on the hardware limitations. The second rule for reducing the number of valid read micro-operations on the UOPQ port further includes: Based on the operating state, the number of effective read micro-operations of the UOPQ port is reduced within a certain number of clock cycles.
[0054] In this embodiment, since most of the static constraints are hardware conditions, if the number of micro-operations of a certain type dispatched surges, exceeding the number of execution units of the corresponding type at the back end of the pipeline (i.e., when the constraint is due to static constraints), the number of micro-operations effectively read by the UOPQ port can be reduced to the number of execution units corresponding to that type of micro-operation. For example, if the execution unit only supports 4 micro-operations / cycle, the number of UOPQ ports can be reduced to 4. Since the static constraints of the execution units are fixed, reducing the number of UOPQ ports does not affect the processor's efficiency, and because the number of ports reduced is significant, the resulting power saving effect is more pronounced.
[0055] Since most of the constraints in dynamic constraints relate to the working state of a certain type of execution unit, this working state may change with the number of micro-operations of the corresponding type. Therefore, when the number of effective read micro-operations of the UOPQ port is reduced according to dynamic constraints, the working state of the execution unit may become better due to the sharp decrease in the number of micro-operations of the corresponding type. If the number of effective read micro-operations of the UOPQ port is restricted for a long time, it will affect the efficiency of the processor. Therefore, for cases where the constraint is due to dynamic constraints, after a short period of time (after a certain number of clock cycles) of reducing the number of effective read micro-operations of the UOPQ port, it is necessary to increase the number of effective read micro-operations of the UOPQ port to improve the efficiency of the processor.
[0056] As another embodiment, after reducing the number of effective micro-operations read by the UOPQ port, the method further includes: obtaining the available resources (TOKEN) of the execution unit corresponding to the first type of micro-operation effectively dispatched by the DISPATCH module; when the TOKEN of the execution unit is continuously higher than a preset threshold, and the DISPATCH module still has this type of micro-operation, the number of effective micro-operations read by the UOPQ port is increased.
[0057] In this embodiment, the available resources (TOKEN) of the execution unit corresponding to a certain type of micro-operation at the pipeline back end can be obtained through a TOKEN information collector. The collector collects the available TOKEN counts for each type of execution unit, such as fixed-point execution units, memory access execution units, and scheduled execution units. Here, TOKEN is the resource identifier of the execution unit, and the available count represents the resource idleness. When a certain type of micro-operation dispatched by the DISPATCH module becomes congested due to static or dynamic constraints, after reducing the number of valid micro-operations read by the UOPQ port, the available resources of the execution unit corresponding to that type of micro-operation can be obtained in real time or according to a predetermined clock cycle interval. When the TOKEN remains greater than a certain threshold for a certain number of clock cycles, the number of valid micro-operations read by the UOPQ port can be increased to improve the processor's processing efficiency.
[0058] In other embodiments, available resources of the execution unit corresponding to the micro-operation effectively dispatched by the DISPATCH module can also be obtained. When the token of the execution unit is continuously higher than a preset threshold, the number of micro-operations effectively read by the UOPQ port is increased. That is, instead of monitoring the token of the execution unit corresponding to a specific micro-operation type in real time, the tokens of execution units of all micro-operation types are monitored to increase the number of micro-operations effectively read by the UOPQ port.
[0059] As an embodiment of this specification, adjusting the number of effective read micro-operations of the UOPQ port includes at least one of the following methods: The number of UOPQ ports can be adjusted by using the UOPQ port enable signal, thereby regulating the number of effective read micro-operations by the UOPQ port; or, By controlling the number of micro-operations written to the UOPQ port by the pipeline front end, and waiting for all micro-operations written to the UOPQ port, the number of UOPQ ports is adjusted, thereby regulating the number of effective micro-operations read by the UOPQ port; or, By controlling the number of micro-operations written to the UOPQ port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ port, and after the micro-operations in the UOPQ have been dispatched, the number of UOPQ ports is adjusted, thereby adjusting the number of micro-operations that the UOPQ port can effectively read; or, By redirecting the REDIRECT action to trigger the FLUSH operation at the front end of the pipeline, the number of UOPQ ports is adjusted, thereby adjusting the number of effective read micro-operations of the UOPQ ports.
[0060] In this embodiment, the number of UOPQ ports can be controlled by the enable signal of the UOPQ port, thereby adjusting the number of effective read micro-operations of the UOPQ port. This method can achieve UOPQ port adjustment without latency. Figure 2 The diagram shown is a schematic of the UOPQ port adjustment circuit according to an embodiment of this specification. This diagram illustrates the adjustment circuit within the UOPQ, such as... Figure 2 The 5-level registers and write / read control logic shown receive the decision-maker instruction for UOPQ port adjustment (decrease or increase), and complete the port width switching within one cycle (or switch directly without delay). When adjusting in the opposite direction, the operation is abandoned to prioritize the adjustment speed.
[0061] UOPQ includes five register levels: register 1, register 2, register 3, register 4, and register 5; the register groups are register 2 and register 3, where: Register 1 is an enable signal, indicating whether to reduce the number of UOPQ ports. The default value is 0 (increase the number of ports). Register 1 can also indicate whether to reduce the number of valid UOPQ read ports. If the enable signal is 1, the number of UOPQ read ports is directly reduced (that is, the number of ports in the current UOPQ port that read micro-operations is reduced), which directly reduces the power consumption waste introduced by invalid UOPQ transitions.
[0062] Register 2 is a buffer for reading micro-operations from the UOPQ port, including 8 SLOTs (valid0-7 marks whether it is valid, uops0-7 stores the micro-operation), which directly outputs UOP to subsequent modules, as shown in Table 9 in the figure.
[0063] Register 3 is the main storage array for UOPQ, as shown in Table 8 in the figure, which stores more UOPs to be buffered.
[0064] Registers 4 and 5 are RdPtr (read pointer) and RdPtrPlus (read pointer + 1) respectively, which control the range of SLOT read from register 3. Register 5 can solve the timing problem of the circuit.
[0065] Read controls 6, 7, and 12 select valid UOPs in register 3 based on RdPtr / RdPtrPlus and output them to subsequent modules. Specifically, based on registers 4 and 5, the number of valid SLOTs read from UOPQ is updated, and new UOP data information is selected through the updated read controls 6 and 7, as shown in read control 12.
[0066] Upon receiving the read control enable information, the read control 10 outputs the valid UOPs in Table 9 according to the adjusted valid read ports as UOPQs.
[0067] Write control 13 implements the write logic of UOP. Type 1 is: if the number of valid micro-operations in register 3 is less than the number of valid read ports of UOPQ, the micro-operations in register 3 and the newly written micro-operations are both written to register 2. Type 2 is: when the number of valid micro-operations in register 3 is greater than the number of valid read ports of UOPQ, the valid micro-operations in register 3 are selected and written to register 2. At the same time, the valid micro-operation data in write operation 13 will be written to register 3.
[0068] When controlling the number of UOPQ ports to decrease the number of UOPQ ports, the number of UOPQ ports is reduced through the following steps: Step 1: Trigger a reduction in the number of UOPQ ports by setting the En signal.
[0069] In this step, when the peripheral circuit determines that the number of valid read micro-operations of the UOPQ port needs to be adjusted, it inputs Enable=1 into register 1 (indicating a reduction in the number of UOPQ ports). At this time, the circuit starts polling to check the status of valid0~7 in register 2 (to determine whether the micro-operations of the current UOPQ port can be cleared).
[0070] Step 2: Enter the emptying state and stop receiving new micro-operations.
[0071] In this step, the circuit control write control 13 pauses writing new micro-operations to register 2 (regardless of the number of valid micro-operations in register 3, register 2 will no longer be filled); at the same time, read control 6, 7, and 12 continue to read the UOPs in register 2 according to the original bandwidth logic until all valid0~7 in register 2 are set to 0, that is, all 8 UOPs in register 2 are dispatched, and the emptying is completed.
[0072] Step 3: Fill in the UOP corresponding to the adjusted UOPQ port.
[0073] In this step, after register 2 is emptied, the circuit switches write control 13 to type 1 or type 2. If the number of valid micro-operations in register 3 is greater than 2, the new micro-operation is first written to register 3; then, only 2 valid UOPs are taken from register 3 and filled into the first 2 SLOTs of register 2 (valid0~1 set to 1, uops0~1 stores UOP), while the last 6 SLOTs (valid2~7) remain invalid. If the number of valid data in register 3 is less than 2, and there is a valid micro-operation in register 3, 1 valid micro-operation in register 3 and the newly written micro-operation are combined to form 2 valid UOPs, which are then filled into the first 2 SLOTs of register 2 (valid0~1 set to 1, uops0~1 stores UOP), while the last 6 SLOTs (valid2~7) remain invalid. If the number of valid data in register 3 is 0, 2 valid UOPs are taken from the newly written micro-operation and filled into the first 2 SLOTs of register 2 (valid0~1 set to 1, uops0~1 stores UOP), while the last 6 SLOTs (valid2~7) remain invalid.
[0074] Step 4: Turn off the unused register clock.
[0075] In this step, the RdPtr / RdPtrPlus range of registers 4 and 5 is reduced to 0~1, and UOP is read only from the first two SLOTs of register 2 to achieve read control update; the circuit turns off the register clock corresponding to SLOT2~7 in register 2, and the registers of these SLOTs stop working and no longer consume power; at this time, register 2 only outputs the UOP of the first two SLOTs to the DISPATCH module, thereby reducing the number of UOPQ ports and thus reducing the number of effective read micro-operations of UOPQ ports.
[0076] When controlling the number of UOPQ ports to increase the number of UOPQ ports, the number of UOPQ ports is increased through the following steps: Step 1: Trigger increase, set the En signal to 0.
[0077] In this step, when the peripheral circuit determines that an additional port needs to be added, it inputs Enable=0 into register 1. The circuit then polls the valid0~1 status of register 2 and waits for it to be cleared.
[0078] Step 2: Empty the micro-operations of the current two ports.
[0079] In this step, write control 13 pauses writing data to register 2, and simultaneously pauses reading data from register 3 and writing it into register 2. Read control continues reading the first two slots of register 2 until valid0~1 is set to 0 (both UOPs are dispatched).
[0080] Step 3: Prefetch 8 UOPs and fill 8 SLOTs in register 2.
[0081] In this step, write control 13 switches back to type 1. If the number of valid micro-operations in register 3 is less than 8, all micro-operations in register 3 and the newly read micro-operations are written to register 2. Eight valid UOPs are taken from register 3 or the input terminal and filled into the eight SLOTs of register 2 (valid0~7 are all set to 1). Alternatively, write control 13 switches back to type 2. Write control 13 writes the latest data into register 3, and at the same time, writes the micro-operations in register 3 into the eight SLOTs of register 2 (valid0~7 are all set to 1).
[0082] Step 4: Turn on the clock and restore the 8 bandwidths to work.
[0083] In this step, RdPtr / RdPtrPlus is restored to the range of 0~7, covering all SLOTs to achieve read control update; the register clock of SLOT2~7 in register 2 is turned on, and all 8 SLOTs resume operation; register 2 outputs the UOP of the 8 SLOTs to the DISPATCH module to increase the number of UOPQ ports, thereby increasing the number of effective read micro-operations of the UOPQ ports.
[0084] Through register 3 in the UOPQ port adjustment circuit of this specification embodiment, and the read / write control logic for reducing and increasing UOPQ ports, new micro-operations can be buffered using register 3. The read control logic reads the micro-operations in register 2 and sends them to the downstream processing logic of the pipeline (e.g., the DISPATCH module) for dispatching micro-operations. When the micro-operations in register 2 are emptied, the write control logic can obtain the adjusted number of valid read micro-operations and write them into register 2, or obtain the adjusted number of valid read micro-operations from register 3 and write them into register 2, or obtain the adjusted number of valid read micro-operations from register 3 and the write control logic and write them into register 2. By controlling the number of UOPQ ports, the number of valid read micro-operations of the UOPQ ports can be adjusted.
[0085] In other embodiments, it may also be as follows: Figure 3The diagram illustrates a blocking write operation in an embodiment of this specification. By controlling the number of micro-operations written to the UOPQ port by the pipeline front end, and waiting for all micro-operations written to the UOPQ port, the number of UOPQ ports is adjusted, thereby adjusting the number of effective micro-operations read by the UOPQ port. When an instruction to adjust the number of effective micro-operations read by the UOPQ port is generated (e.g., reducing the number of effective micro-operations read from 8 to 2), a STALL signal is sent to the instruction decoding channel and the OC (Op Cache) instruction fetch channel; wait for N cycles (e.g., N=3 for a 3-stage pipeline) to ensure that there are no micro-operations not written to UOPQ in the pipeline; then adjust the UOPQ port, switch the read pointer to SLOT0~1, and turn off the clocks of SLOT2~7; after confirming that the UOPQ port configuration is complete, set the STALL signal low, allowing the instruction decoding channel and the OC instruction fetch channel to write new micro-operations to UOPQ; UOPQ reads micro-operations at 2 ports. For example, the decision maker determines that the number of micro-operations read from the UOPQ port needs to be reduced to 2, and sends a STALL signal to the instruction decoding channel; waits for 3 cycles, and all 3 in-flight micro-operations (micro-operations that have not been dispatched) in the pipeline are written to UOPQ and dispatched; adjusts the read pointer to SLOT0~1, and turns off the clock of the remaining 6 SLOTs; sets the STALL signal low, and writes the new micro-operations to UOPQ in 2 ports to avoid conflicts caused by data residue during switching.
[0086] In other embodiments, such as Figure 3 As shown, the number of micro-operations written to the UOPQ port by the pipeline front end can also be controlled. After all the micro-operations at the pipeline front end are written to the UOPQ port, and after the micro-operations in the UOPQ are dispatched, the number of UOPQ ports can be adjusted, thereby adjusting the number of micro-operations effectively read by the UOPQ port. The difference from the previous embodiment is that in this embodiment, after all the micro-operations at the pipeline front end are written to the UOPQ port, it is also necessary to wait for the micro-operations in the UOPQ to be dispatched before adjusting the number of UOPQ ports. For the sake of simplicity, the similarities between this embodiment and the previous embodiment will not be repeated.
[0087] In other embodiments, it may also be as follows: Figure 4The diagram illustrates an embodiment of this specification that uses a pipeline front-end clearing operation to adjust the number of UOPQ ports. A REDIRECT action triggers a FLUSH operation at the pipeline front-end, clearing all micro-operations in the UOPQ ports. The number of UOPQ ports is then adjusted via their enable signals, thereby regulating the number of effective micro-operations read by each UOPQ port. By using the REDIRECT action to FLUSH the pipeline front-end (branch prediction unit, instruction fetch unit, etc.), all unfinished micro-operations read by the UOPQ ports are cleared, and the execution start point is redirected, readjusting the number of UOPQ ports and simplifying the adjustment circuit design. For example, the decision-maker determines that the number of valid micro-operations read from the UOPQ port needs to be increased to 8, triggering the REDIRECT action; the REDIRECT signal triggers FLUSH, clearing any pending micro-operations or instructions in the instruction fetch unit, instruction decoding channel, and OC instruction fetch channel at the front end of the pipeline; while clearing UOPQ, the UOPQ port is adjusted, the read pointer covers slots 0~7, and all SLOT clocks are turned on; REDIRECT specifies a new execution start point (the current program counter PC), and the pipeline fetches and decodes instructions from this address again; the new micro-operations are written to the UOPQ port with an 8-port bandwidth, the system returns to normal, and the port adjustment is completed.
[0088] Specifically, the current state is as follows: the number of valid micro-operations read by the UOPQ port is 2 (i.e., 2 SLOTs), the historical FLUSH recovery time is 3 cycles, the current IPC is 95% of the baseline value, and the maximum token capacity of the execution unit is 8. The decision-maker determines that the UOPQ port bandwidth for 8 micro-operations needs to be restored, triggering the REDIRECT action; FLUSH clears the 3 pending instructions in the instruction fetch unit and the 2 undisassembled micro-operations in the instruction decoding channel; the UOPQ port is adjusted to 8, the read pointer covers all 8 SLOTs, and the clock is fully enabled; the pipeline fetches instructions again from the current PC value, 8 new micro-operations are written to UOPQ, the DISPATCH module dispatches according to the bandwidth of 8 micro-operations, and the performance is fully restored after 3 cycles, with no data interference during the adjustment process.
[0089] As one embodiment of this specification, after reducing the number of effective read micro-operations of the UOPQ port, the following is also included: Monitor the number of micro-operations effectively dispatched by the DISPATCH module; If the number of first-type micro-operations effectively dispatched by the DISPATCH module remains below a preset threshold, the number of micro-operations effectively read by the UOPQ port will be increased.
[0090] In this embodiment, when the number of effective micro-operations read by the UOPQ port is reduced, the number of effective micro-operations dispatched by the DISPATCH module remains below a preset threshold. In order to prevent the processor performance from degrading, the number of effective micro-operations read by the UOPQ port can be increased to restore the full micro-operation read capability of the UOPQ port.
[0091] In another embodiment, if the difference between the number of micro-operations effectively read by the UOPQ port and the number of micro-operations effectively dispatched by the DISPATCH module continuously exceeds a certain preset threshold after the number of micro-operations effectively read by the UOPQ port is reduced, the number of micro-operations effectively read by the UOPQ port can be increased, for example, increased to the maximum. Here, "continuously" can be a predetermined time period, multiple consecutive clock cycles, or multiple consecutive effective dispatches; the micro-operations effectively dispatched by the DISPATCH module refer to the micro-operations dispatched by the DISPATCH module to the downstream execution units of the pipeline that can be executed; and the micro-operations that the DISPATCH module can dispatch refer to the micro-operations that the downstream execution units of the DISPATCH module can still receive and process, i.e., the available resources of the execution units.
[0092] By monitoring the number of micro-operations effectively dispatched by the DISPATCH module, the number of effective micro-operations read by the UOPQ port can be increased in real time, thereby improving the processor's processing efficiency and avoiding the problem of excessive adjustment of effective micro-operations read by the UOPQ port, which would reduce the processor's processing efficiency.
[0093] As one embodiment of this specification, after reducing the number of effective read micro-operations of the UOPQ port, the following is also included: Monitor changes in processor performance (Instructions Per Cycle, IPC); As the DISPATCH module completes the dispatch of micro-operations, if the IPC remains below a preset threshold, the number of micro-operations effectively read by the UOPQ port will be increased.
[0094] In this step, before adjusting the number of valid read micro-operations on the UOPQ port, the IPC value under stable processor conditions is recorded in the history. For example, an average IPC of 5 over 10 consecutive cycles serves as a preset threshold for judging performance degradation. If the IPC falls below the preset threshold within a certain duration or clock cycle, performance degradation can be considered, triggering an error correction mechanism to increase the number of valid read micro-operations on the UOPQ port, for example, by increasing it to the maximum.
[0095] In another embodiment, after the number of valid read micro-operations of the UOPQ port is reduced, the number of execution units (TOKENs) at the back end of the processor pipeline is monitored; as the DISPATCH module completes the dispatch of micro-operations, if the number of TOKENs corresponding to a certain type of micro-operation continues to be higher than a preset threshold, and if there are still micro-operations of that type in the DISPATCH module, then the number of valid read micro-operations of the UOPQ port is adjusted to be increased.
[0096] By monitoring changes in processor performance IPC through this embodiment, the number of effective read micro-operations of the UOPQ port can be increased in real time, thereby improving the processor's processing efficiency and avoiding the problem of excessive adjustment of effective read micro-operations of the UOPQ port, which would reduce the processor's processing efficiency.
[0097] like Figure 5 The diagram shown is a schematic representation of the adaptive bandwidth scheduling device according to an embodiment of this specification. This diagram illustrates a device for operating the methods described in the above embodiments. Each functional module can be implemented through circuits, or the functions of each module can be accomplished through a combination of instruction sets and circuits. Specifically, the device includes: The acquisition unit 501 is configured to acquire the first number of micro-operations read from the operation queue UOPQ port and the second number of micro-operations effectively dispatched by the instruction dispatch DISPATCH module. Decision maker 502 is configured to adjust the number of effective read micro-operations of the UOPQ port based on the difference between the first number and the second number.
[0098] like Figure 6 The diagram shown is a structural schematic of the instruction dispatch unit according to an embodiment of this specification. The internal structure of the instruction dispatch unit is illustrated in this diagram. The functional components can be implemented using hardware circuitry or a combination of instruction sets and hardware circuitry. Not all functional modules are required in this diagram; only a portion of the functional modules are needed to achieve the desired purpose. Specifically, the instruction dispatch unit includes: UOPQ module 601, intermediate processing module 602, dispatch module 603, information collection module 604, TOKEN module 605, type recognition module 606, history record module 607, decision module 608, performance monitoring module 609, execution unit 610.
[0099] The UOPQ module 601 is configured to read a specified number of micro-operations and transmit them to the intermediate processing module 602 under the control of the decision module 608. The intermediate processing module 602 is configured to send the micro-operations read by the UOPQ module 601 to the dispatch module 603. The intermediate processing module 602 includes, in addition to the initial decoding circuit, a static constraint circuit and a dynamic constraint circuit. The static constraint circuit is configured to adjust the number of transmitted micro-operations according to the hardware limitations of the DISPATCH module and the pipeline back-end execution unit. The dynamic constraint circuit is configured to adjust the number of transmitted micro-operations according to the operating state of the DISPATCH module and the pipeline back-end execution unit. Figure 7 The diagram shown is a structural schematic of the intermediate processing module in an embodiment of this specification. It includes a micro-operation initial decoding stage circuit, a static constraint circuit, and a dynamic constraint circuit. Between reading micro-operations from the UOPQ port and dispatching micro-operations to the back-end execution unit, there exist static and dynamic constraint circuits. Specifically, the number of micro-operations read from the UOPQ port is the first quantity; the number of micro-operations obtained after initial decoding is the third quantity; the number of micro-operations processed by the static constraint circuit is the fourth quantity; the number of micro-operations processed by the dynamic constraint circuit is the fifth quantity; and the number of micro-operations dispatched by the dispatch module is the second quantity. The relationship between these five quantities is as follows: First quantity ≤ Third quantity ≥ Fourth quantity ≥ Fifth quantity = Second quantity.
[0100] Among them, when the micro-operations are processed by the initial decoding circuit, there are micro-operations that are split and inserted into FIX OP (fixed-point arithmetic). Such operations will result in the number of micro-operations (third number) after passing through the initial decoding circuit being greater than or equal to the number of micro-operations read by the UOPQ port (first number).
[0101] Because the execution units within a processor core have upper limits—for example, the maximum execution capacity of a fixed-point arithmetic execution unit for addition operations—a static constraint circuit exists during the dispatch process. This static constraint circuit ensures that the micro-operations sent in the current cycle do not exceed the execution capacity of the corresponding micro-operation type execution unit in each cycle. The circuit for this static constraint stage is called the static constraint circuit, which limits the number of dispatched micro-operations based on the capacity of the back-end execution unit in a single cycle. This also includes static constraint restrictions that some special micro-operations can only be placed in dispatched slot-0.
[0102] During processor operation, congestion may occur in specific execution units. Congested micro-operations will be placed into the corresponding ISSUE QUEUE, such as the ALU QUEUE storing addition micro-operations. Therefore, to prevent ALU QUEUE overflow (leading to execution errors), dynamic constraints are required. During the dynamic constraint phase, the dynamic constraint circuit dispatches micro-operations based on the number of QUEUEs required for the dispatched micro-operations and the number of available QUEUE ENTRYs. Micro-operations that have passed the dynamic constraint phase are then directly dispatched to the execution units via the dispatch module.
[0103] In this embodiment, the decision module 608 can analyze the source of the difference between the number of micro-operations read from the UOPQ port (first quantity) and the number of micro-operations dispatched (second quantity) through the initial decoding circuit, static constraint circuit, and dynamic constraint circuit of the micro-operation.
[0104] The dispatch module 603 is configured to dispatch micro-operations to the corresponding execution unit 610; The information collection module 604 is configured to obtain the first number of micro-operations read by the UOPQ module 601, and the second number of micro-operations effectively dispatched by the dispatch module 603; and to obtain the static and dynamic constraints of the dispatch module 603 and the execution unit 610. The TOKEN module 605 is configured to acquire available resources of the execution unit 610; The type identification module 606 is configured to identify the type of micro-operation that has been effectively dispatched; The historical record module 607 is configured to record information acquired by the information collection module 604 and the type identification module 606, including information on micro-operations dispatched by the UOPQ module 601 and the dispatch module 603, as well as static and dynamic constraints of the intermediate processing module 602. The decision module 608 is configured to adjust the number of micro-operations that the UOPQ port of the UOPQ module 601 can effectively read based on the difference between the first quantity and the second quantity obtained by the information collection module 604.
[0105] In another embodiment, the decision module 608 is configured to adjust the number of effective read micro-operations of the UOPQ port of the UOPQ module 601 according to the intermediate processing module 602, the dispatch module 603, the TOKEN module 605, the type recognition module 606, and the history module 607.
[0106] The performance monitoring module 609 is configured to monitor the system performance of the processor and feed back the changes in system performance to the decision module 608, so that the decision module 608 can adjust the number of effective read micro-operations of the UOPQ port of the UOPQ module 601 according to the changes in system performance.
[0107] The following is a description of the process steps of an instruction dispatch unit with the above structure, which reduces the number of UOPQ ports, thereby reducing the number of effective read micro-operations of the UOPQ ports: Step 1: Collect data from multiple modules.
[0108] In this step, the information collection module 604, the TOKEN module 605, the type recognition module 606, the historical record module 607, and the performance monitoring module 609 work synchronously to output full-dimensional data to the decision-making module 608 to ensure accurate decision-making.
[0109] Specifically, the information collection module 604 collects in real time the first number of read micro-operations from the UOPQ module 601 and the second number of effectively dispatched micro-operations from the dispatch module 603. It also obtains the static and dynamic constraints from the intermediate processing module 602.
[0110] The TOKEN module 605 iterates through fixed-point (EX), floating-point (FP), and memory access (LSU) execution units to count the available TOKENs for each execution unit.
[0111] The type identification module 606 parses the micro-operations of the DISPATCH GROUP of the dispatch module 603, classifies them by function (fixed-point / floating-point / memory access, which may include subdivided types such as addition / multiplication / vector), and marks the micro-operation type.
[0112] The history module 607 retrieves the number of micro-operations effectively dispatched by the dispatch module 603 for multiple historical cycles, the number of TOKENs in each execution unit, the micro-operation type, and port adjustment records.
[0113] The performance monitoring module 609 collects real-time values of the current IPC (instructions per cycle) and the number of DISPATCH micro-operations, compares them with historical baseline values (stable values before adjustment), and outputs the performance fluctuation status.
[0114] All module output data is synchronously transmitted to decision module 608 via the internal bus.
[0115] Step 2: The decision module 608 calculates the difference between the first number of read micro-operations of the UOPQ module 601 and the second number of effectively dispatched micro-operations of the dispatch module 603.
[0116] In this step, the decision module 608, in conjunction with the first and second quantities collected by the information collection module 604 (or the historical record module 607), can subtract the first quantity from the second quantity to obtain the difference.
[0117] Step 3: The decision module 608 determines whether the difference is continuously greater than the first threshold or continuously less than the second threshold.
[0118] In this step, if the difference is consistently greater than the first threshold for a predetermined number of clock cycles, it indicates that the effective dispatch micro-operations of the DISPATCH module are continuously congested, and the number of effective read micro-operations of the UOPQ port needs to be reduced; if the difference is consistently less than the second threshold for a predetermined number of clock cycles, it indicates that the effective dispatch micro-operations of the DISPATCH module are continuously unimpeded, and the number of effective read micro-operations of the UOPQ port can be increased.
[0119] Step 4: When the effective dispatching of micro-operations by the DISPATCH module continues to be congested, analyze the reasons for the congestion.
[0120] In this step, the reasons for the continuous congestion in the effective dispatch of micro-operations in the DISPATCH module may include static constraints or dynamic constraints. The causes of congestion can be analyzed based on the static and dynamic constraint information obtained from the intermediate processing module 602.
[0121] When the number of micro-operations obtained by the dynamic constraint circuit is less than the number of micro-operations read by the UOPQ port, the reason for the limitation that the difference is continuously greater than the first threshold is static constraint. or, When the number of micro-operations acquired by the static constraint circuit is greater than the number of micro-operations read by the UOPQ port, the reason for the limitation that the difference continues to be greater than the first threshold is static constraint. or, When the number of micro-operations output by the dynamic constraint circuit is less than the number of micro-operations obtained by the dynamic constraint circuit, the reason for the limitation that the difference continues to be greater than the first threshold is dynamic constraint.
[0122] Specifically, the decision module 608 can call the DISPATCH module of the history record module 607 to effectively distribute micro-operations with limited fluctuation ranges because of static or dynamic constraints.
[0123] If the fluctuation range is less than or equal to the stability threshold (e.g., 5%), and the congestion cause for effective micro-operation dispatch is the static constraint for 4 consecutive cycles (e.g., N=4), then the cause of congestion is the static constraint.
[0124] If the fluctuation amplitude is greater than the stability threshold, and there are static and dynamic constraints among the congestion causes of effective dispatch of micro-operations for M consecutive cycles (e.g., M=3), then the cause of congestion can be determined to be dynamic constraints.
[0125] The congestion of effective dispatch micro-operations in the Display module can also be analyzed based on the TOKEN of the execution unit to determine whether it is due to static or dynamic constraints.
[0126] Decision module 608 analyzes the causes of congestion based on the number of available tokens collected by token module 605: If the change in the number of available tokens is less than the stability threshold within a predetermined time period, for example, if the stability threshold for a floating-point unit is 2, and the number of available tokens for the execution unit in five consecutive cycles is cycle0=4, cycle1=4, cycle2=5, cycle3=4, cycle4=3, and the change in the number of tokens is less than the stability threshold of 2, then the cause of the congestion can be attributed to static constraints.
[0127] If the change in the number of available tokens exceeds a stable threshold within a predetermined time period, for example, if the stable threshold for a floating-point unit is 2, and the number of available tokens for the execution unit is 6 in cycle 0 and 2 in cycle 1, and the change in the number of tokens exceeds the stable threshold of 2, then the cause of the congestion is dynamic constraints.
[0128] Step 5: The decision module 608 adjusts the number of ports of the UOPQ module 601 according to the cause of congestion.
[0129] In this step, the decision module 608 sends an instruction to the UOPQ module 601 to reduce the number of ports, and the intermediate processing module 602 synchronously adapts to the new port width.
[0130] Specifically, the decision module 608 outputs an instruction and simultaneously sends a port adjustment notification to the intermediate processing module 602 (informing that the new reading range is slot0~3).
[0131] The Enable signal of register 1 of UOPQ module 601 is set to 1 (start port reduction); the valid status of register 2 is polled and the emptying stage is entered; the range is adjusted to SLOT0~3 through register 4 / 5, and SLOT4~7 is no longer read; the clock gating module turns off the register clock of SLOT4~7 to reduce invalid power consumption; after emptying is completed, 4 UOPs are prefetched from the main memory array (register 3) to fill SLOT0~3.
[0132] After receiving the port adjustment notification, the intermediate processing module 602 adjusts the working range of the Mapping and Arbiter, processing only the UOPs of SLOT0~3 output by the UOPQ module 601 to avoid data flow outside the range and reduce circuit transitions. The number of ports of the UOPQ module 601 becomes 4, outputting only 4 UOPs / cycle, matching the effective dispatch of the dispatch module 603 and eliminating bandwidth differences.
[0133] Step 6: The decision module 608 restores the UOPQ module 601 port based on the monitoring results of the performance monitoring module 609.
[0134] In this step, the performance monitoring module 609 acquires processor performance data in real time, including IPC and the number of micro-operations effectively dispatched by the dispatch module 603. The decision module 608 decides whether to increase the port resources of the UOPQ module 601 based on the processor performance data.
[0135] Specifically, when the available resources (TOKEN) of the execution unit are continuously greater than a preset threshold, the decision module 608 can adjust the number of ports of the UOPQ module 601. For specific methods, please refer to the aforementioned embodiments.
[0136] The methods and apparatus described in the embodiments of this specification can reduce the power consumption of the processor without changing its performance.
[0137] In one embodiment of this specification, a processor including the above-described instruction dispatch unit can execute the aforementioned adaptive bandwidth scheduling method to reduce the processor's power consumption.
[0138] like Figure 8The diagram illustrates a computer device according to an embodiment of this specification. The computer device in this embodiment may include the processor described in this specification and utilize the aforementioned adaptive bandwidth scheduling method. This method can also be run on the computer device in this embodiment to execute the methods described in this specification. The computer device 802 may include one or more processors 804, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 802 may also include any memory 806 for storing information of any kind, such as code, settings, data, etc. Non-limitingly, for example, memory 806 may include any type of RAM, any type of ROM, flash memory, hard disk, optical disk, etc. More generally, any memory can use any technology to store information. Further, any memory can provide volatile or non-volatile retention of information. Further, any memory can represent a fixed or removable component of the computer device 802. In one case, when the processor 804 executes associated instructions stored in any memory or combination of memories, the computer device 802 can perform any operation of the associated instructions. The computer device 802 also includes one or more drive mechanisms 808 for interacting with any memory, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.
[0139] Computer device 802 may also include an input / output module 810 (I / O) for receiving various inputs (via input device 812) and providing various outputs (via output device 814). A specific output mechanism may include a presentation device 816 and an associated graphical user interface (GUI) 818. In other embodiments, the input / output module 810 (I / O), input device 812, and output device 814 may be omitted, and the device may function solely as a computer device within a network. Computer device 802 may also include one or more network interfaces 820 for exchanging data with other devices via one or more communication links 822. One or more communication buses 824 couple the components described above together.
[0140] Communication link 822 can be implemented in any way, such as via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 822 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0141] This specification also provides computer-readable instructions, wherein when a processor executes the instructions, the program therein causes the processor to perform the methods described above.
[0142] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the methods described above.
[0143] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.
[0144] It should be understood that in the various embodiments of this specification, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this specification.
[0145] It should also be understood that, in the embodiments of this specification, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this specification generally indicates that the preceding and following related objects have an "or" relationship.
[0146] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this specification.
[0147] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0148] In the several embodiments provided in this specification, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or units, or they may be electrical, mechanical, or other forms of connection.
[0149] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments described in this specification, depending on actual needs.
[0150] Furthermore, the functional units in the various embodiments of this specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0151] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this specification, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this specification. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0152] This specification uses specific embodiments to illustrate the principles and implementation methods of this specification. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this specification. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this specification. Therefore, the content of this specification should not be construed as a limitation of this specification.
Claims
1. An adaptive bandwidth scheduling method, characterized in that, include: Get the first number of micro-operations read from the UOPQ port of the operation queue, and the second number of micro-operations effectively dispatched by the DISPATCH module; The number of effective read micro-operations of the UOPQ port is adjusted based on the difference between the first and second quantities.
2. The method according to claim 1, characterized in that, The effective dispatch of micro-operations by the DISPATCH module includes: the micro-operations that are effectively dispatched; the number of micro-operations that are effectively read by the UOPQ port includes: the number of micro-operations that can be read by the UOPQ port corresponding to the micro-operations that can be effectively dispatched by the DISPATCH module.
3. The method according to claim 2, characterized in that, The second number of micro-operations effectively dispatched by the DISPATCH module includes: the second number of first-type micro-operations effectively dispatched by the DISPATCH module, wherein the difference is the difference between the first number of first-type micro-operations read by the UOPQ port and the second number of first-type micro-operations effectively dispatched by the DISPATCH module.
4. The method according to claim 3, characterized in that, Adjusting the number of effective read micro-operations of the UOPQ port based on the difference between the first and second quantities further includes: When the difference continues to exceed the first threshold, the number of effective read micro-operations of the UOPQ port is reduced. When the difference remains below the second threshold, the number of valid read micro-operations on the UOPQ port is increased.
5. The method according to claim 4, characterized in that, The differences include: DELTA = First quantity - Second quantity; Where DELTA is the difference between the first quantity and the second quantity.
6. The method according to claim 4, characterized in that, When the difference continues to exceed the first threshold, reducing the number of effective read micro-operations on the UOPQ port further includes: Analyze the reasons why the difference continues to exceed the first threshold. If the limitation is due to a static constraint, then the number of valid read micro-operations on the UOPQ port is reduced according to the first rule; If the limitation is due to dynamic constraints, then the number of valid read micro-operations on the UOPQ port is reduced according to the second rule.
7. The method according to claim 6, characterized in that, The static constraints include: hardware limitations on the DISPATCH module and the pipeline back-end execution unit; The dynamic constraints include the working status of the DISPATCH module and the pipeline back-end execution unit.
8. The method according to claim 6, characterized in that, The specific reasons for the limitation that the difference continues to be greater than the first threshold include: During the process of sending the micro-operations read from the UOPQ port to the DISPATCH module through static and dynamic constraint circuits. When the number of micro-operations obtained by the dynamic constraint circuit is less than the number of micro-operations read by the UOPQ port, the reason for the limitation that the difference is continuously greater than the first threshold is static constraint. or, When the number of micro-operations acquired by the static constraint circuit is greater than the number of micro-operations read by the UOPQ port, the reason for the limitation that the difference continues to be greater than the first threshold is static constraint. or, When the number of micro-operations output by the dynamic constraint circuit is less than the number of micro-operations obtained by the dynamic constraint circuit, the reason for the limitation that the difference continues to be greater than the first threshold is dynamic constraint.
9. The method according to claim 6, characterized in that, The specific reasons for the limitation that the difference continues to be greater than the first threshold include: Obtain the available resource tokens for the execution units corresponding to the first type of micro-operation effectively dispatched by the DISPATCH module; If the change of the TOKEN of the execution unit within a predetermined time is less than a stable threshold, then the reason for the limitation that the difference continues to be greater than the first threshold is a static constraint. If the change of the TOKEN of the execution unit within a predetermined time period is greater than a stable threshold, then the reason for the limitation that the difference continues to be greater than the first threshold is dynamic constraint.
10. The method according to claim 7, characterized in that, The reduction of the number of valid read micro-operations on the UOPQ port according to the first rule further includes: The number of valid read micro-operations on the UOPQ port is reduced to a specified number based on the hardware limitations. The second rule for reducing the number of valid read micro-operations on the UOPQ port further includes: Based on the operating state, the number of effective read micro-operations of the UOPQ port is reduced within a certain number of clock cycles.
11. The method according to claim 1, characterized in that, Adjusting the number of effective read micro-operations of the UOPQ port includes at least one of the following methods: The number of UOPQ ports can be adjusted by using the UOPQ port enable signal, thereby regulating the number of effective read micro-operations by the UOPQ port; or, By controlling the number of micro-operations written to the UOPQ port by the pipeline front end, and waiting for all micro-operations written to the UOPQ port, the number of UOPQ ports is adjusted, thereby regulating the number of effective micro-operations read by the UOPQ port; or, By controlling the number of micro-operations written to the UOPQ port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ port, and after the micro-operations in the UOPQ have been dispatched, the number of UOPQ ports is adjusted, thereby adjusting the number of micro-operations that the UOPQ port can effectively read; or, By redirecting the REDIRECT action to trigger the FLUSH operation at the front end of the pipeline, the number of UOPQ ports is adjusted, thereby adjusting the number of effective read micro-operations of the UOPQ ports.
12. The method according to claim 4, characterized in that, Reducing the number of valid read micro-operations on the UOPQ port also includes: Monitor the number of micro-operations effectively dispatched by the DISPATCH module; If the number of first-type micro-operations effectively dispatched by the DISPATCH module remains below a preset threshold, the number of micro-operations effectively read by the UOPQ port will be increased.
13. The method according to claim 4, characterized in that, Reducing the number of valid read micro-operations on the UOPQ port also includes: Monitor changes in processor performance IPC; As the DISPATCH module completes the dispatch of micro-operations, if the IPC remains below a preset threshold, the number of effective micro-operations read by the UOPQ port will be increased.
14. The method according to claim 4, characterized in that, Reducing the number of valid read micro-operations on the UOPQ port also includes: Monitor the number of execution unit tokens at the back end of the processor pipeline; As the DISPATCH module completes the dispatch of micro-operations, if the number of TOKENs corresponding to a certain type of micro-operation continues to exceed a preset threshold, and if there are still micro-operations of that type in the DISPATCH module, then the number of micro-operations that are effectively read by the UOPQ port is increased.
15. An adaptive bandwidth scheduling device, characterized in that... include: The acquisition unit is configured to acquire the first number of micro-operations read from the operation queue UOPQ port, and the second number of micro-operations effectively dispatched by the instruction dispatch DISPATCH module. The decision-maker is configured to adjust the number of effective read micro-operations of the UOPQ port based on the difference between the first number and the second number.
16. An instruction dispatching unit, characterized in that, The instruction dispatching unit executes the adaptive bandwidth scheduling method according to any one of claims 1-14.
17. A processor, characterized in that... Includes the instruction dispatching unit as described in claim 16.
18. A computer device comprising a memory and a computer program stored in the memory and executable on a processor, characterized in that, Includes the processor as described in claim 17 above.