An adaptive bandwidth scheduling method and device

By dynamically adjusting the number of read ports in the UOPQ module, the problem of ineffective dispatching of micro-operations in the UOPQ module was solved, thereby improving processor energy efficiency.

CN122195514APending Publication Date: 2026-06-12HYGON INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HYGON INFORMATION TECH CO LTD
Filing Date
2026-03-27
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In traditional CPU microarchitectures, micro-operations read from the Micro-Operations Queue Module (UOPQ) port cannot be effectively dispatched, resulting in invalid flips and increased power consumption, thus reducing processor performance.

Method used

By obtaining the bit value difference between the UOPQ module and the instruction dispatch module (DISPATCH), the number of read ports of the UOPQ module is dynamically adjusted to match the dispatch requirements, reduce invalid flips and power consumption.

Benefits of technology

Without reducing pipeline efficiency, reduce invalid flips of the UOPQ module to improve processor energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195514A_ABST
    Figure CN122195514A_ABST
Patent Text Reader

Abstract

The present specification relates to the technical field of computers, and particularly relates to an adaptive bandwidth scheduling method and device. The method comprises obtaining a first bit value of each read port of a micro-operation queue (UOPQ) module and a second bit value of each dispatch port of an instruction dispatch (DISPATCH) module; and adjusting the number of read ports of the UOPQ module according to the difference between the first bit value and the second bit value. By using the embodiment of the present specification, the power consumption of the UOPQ module, the power consumption of the DISPATCH module, and the power consumption of the execution unit can be reduced without reducing the pipeline efficiency, and the energy efficiency of the processor can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to an adaptive bandwidth scheduling method and apparatus. Background Technology

[0002] In CPU microarchitecture, the pipeline is mainly divided into the front end and the back end. As the types and number of back-end execution units continue to increase, the processor's back-end execution capability continues to improve, and the bandwidth demand for front-end instruction dispatch is growing. However, in traditional designs, when the number of micro-operations read from the micro-operation queue (UOPQ) module's read port exceeds the number of micro-operations that the instruction dispatch module (DISPATCH) can dispatch, the data read from the UOPQ module will repeatedly flip over over multiple clock cycles, causing jumps in the UOPQ's internal registers. The DI0 stage (data input / output stage) circuit jumps increase the power consumption of the UOPQ module and do not contribute to the reading, parsing, and dispatching of micro-operations, thus reducing processor performance.

[0003] Improving processor energy efficiency is an urgent problem that needs to be solved. Summary of the Invention

[0004] To address the problems in the prior art, this specification provides an adaptive bandwidth scheduling method and apparatus that can solve the problem that micro-operations read from the micro-operation queue (UOPQ) port in the processor front end cannot be effectively dispatched, and that invalid flips generated by the UOPQ module cause low processor energy efficiency.

[0005] This specification provides an adaptive bandwidth scheduling method, including: The first bit value of each read port of the micro-operation queue UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module are obtained. The first bit value of each read port of the UOPQ module includes the bit value corresponding to at least a portion of the bits in the read port, and the second bit value of each dispatch port of the DISPATCH module includes the bit value corresponding to the corresponding bit in the dispatch port. The number of read ports of the UOPQ module is adjusted based on the difference between the first bit value and the second bit value.

[0006] As a further aspect of this specification, obtaining the first bit value of each read port of the operation queue UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module further includes: During a dispatch micro-operation, the first bit value of each read port of the UOPQ module and the second bit value of each dispatch port of the DISPATCH module are obtained.

[0007] As a further aspect of this specification, the bit value corresponding to the corresponding bit in the dispatch port includes: the bit value corresponding to the corresponding bit in the dispatch port where the dispatched micro-operation is located, or the preset bit value in the dispatch port where the undispatched micro-operation is located.

[0008] As a further aspect of this specification, the bit values ​​corresponding to at least some bits in the read port include: The bit values ​​corresponding to multiple consecutive bits in the read port; or, The bit values ​​corresponding to multiple non-contiguous bits in the read port.

[0009] As a further aspect of this specification, the differences include: The number of bits whose first bit value at the read port differs from the second bit value at the dispatch port; or... The proportion of bits whose first bit value of the read port is different from the second bit value of the dispatch port to the total number of bits in all dispatch ports.

[0010] As another further aspect of this specification, Adjusting the number of read ports of the UOPQ module based on the difference between the first bit value and the second bit value further includes: When the number of differences exceeds the first threshold, the number of read ports of the UOPQ module is adjusted to a state of excessive power consumption.

[0011] As a further aspect of this specification, adjusting the number of read ports of the UOPQ module to prevent excessive power consumption further includes: If the power consumption exceeds the limit and the clock cycle lasts for more than the second threshold, then the number of read ports of the UOPQ module is reduced. If the clock cycle of the excessive power consumption state is less than the third threshold, then the number of read ports of the UOPQ module is increased.

[0012] As a further aspect of this specification, when the excessive power consumption state persists for more than a second threshold clock cycle, reducing the number of read ports of the UOPQ module further includes: If, during multiple micro-operations, the power consumption exceeds the second threshold for several consecutive clock cycles, the number of read ports of the UOPQ module will be reduced.

[0013] As a further aspect of this specification, before adjusting the number of read ports of the UOPQ module based on the difference between the first bit value and the second bit value, the following is also included: Match the first bit value and the second bit value according to the thread corresponding to the first bit value and the second bit value.

[0014] As a further aspect of this specification, adjusting the number of read ports of the UOPQ module includes at least one of the following methods: The number of UOPQ module read ports is adjusted by using the adjustment enable signal of the UOPQ module read port; or, By controlling the number of micro-operations written to the UOPQ module read port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ module read port, the number of UOPQ module read ports is adjusted; or, By controlling the number of micro-operations written to the UOPQ module read port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ module read port, and after the micro-operations in the UOPQ module have been dispatched, the number of UOPQ module read ports is adjusted; or, The number of read ports of the UOPQ module is adjusted by triggering the FLUSH operation at the front end of the pipeline through the REDIRECT action.

[0015] As a further aspect of this specification, reducing the number of read ports of the UOPQ module also includes: Monitor the number of micro-operations effectively dispatched by the DISPATCH module; If the number of micro-operations effectively dispatched by the DISPATCH module remains below a preset threshold, the number of read ports of the UOPQ module will be increased.

[0016] As a further aspect of this specification, reducing the number of read ports of the UOPQ module also includes: Monitor changes in processor performance IPC; As the DISPATCH module completes the dispatch of micro-operations, if the IPC remains below a preset threshold, the number of ports read by the UOPQ module will be increased.

[0017] As a further aspect of this specification, reducing the number of read ports of the UOPQ module also includes: Monitor the number of execution unit tokens at the back end of the processor pipeline; As the DISPATCH module completes the dispatch of micro-operations, if the number of TOKENs corresponding to a certain type of micro-operation continues to exceed a preset threshold, and if there are still micro-operations of that type in the DISPATCH module, then the number of read ports of the UOPQ module is increased.

[0018] This specification also provides an adaptive bandwidth scheduling device, comprising: The acquisition unit is configured to acquire the first bit value of each read port of the operation queue UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module, wherein the first bit value of each read port of the UOPQ module includes the bit value corresponding to at least a portion of the bits in the read port, and the second bit value of each dispatch port of the DISPATCH module includes the bit value corresponding to the corresponding bit in the dispatch port. The decision-maker is configured to adjust the number of read ports of the UOPQ module based on the difference between the first bit value and the second bit value.

[0019] This specification also provides an instruction dispatching unit that executes the adaptive bandwidth scheduling method described above.

[0020] This specification also provides a processor, including an instruction dispatch unit for performing the above-described methods.

[0021] This specification also provides a computer device, including a memory and a computer program stored in the memory and executable on a processor, including the processor described above.

[0022] This specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the above-described method.

[0023] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.

[0024] By comparing the bit values ​​of the dispatch port of the DISPATCH module and the read port of the UOPQ module using the embodiments in this specification, it can be determined whether the UOPQ module is causing power loss. By adjusting the number of UOPQ ports, the invalid flips of the UOPQ module can be reduced without reducing pipeline efficiency, thereby improving the energy efficiency of the processor. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1The diagram shown illustrates an invalid flipping of a micro-operation within the UOPQ module in an embodiment of this specification. Figure 2 The diagram shown is a flowchart of an adaptive bandwidth scheduling method according to an embodiment of this specification. Figure 3 The diagram shown is a schematic of the UOPQ module read port adjustment circuit in an embodiment of this specification. Figure 4 The diagram shown is a schematic diagram of a blocking write operation in an embodiment of this specification; Figure 5 The diagram shown is a schematic of how to clear the front end of the pipeline to adjust the number of read ports of the UOPQ module according to an embodiment of this specification; Figure 6 The diagram shown is a structural schematic of the adaptive bandwidth scheduling device according to an embodiment of this specification. Figure 7 The diagram shown is a structural schematic of the instruction dispatch unit in an embodiment of this specification. Figure 8 The diagram shown is a schematic representation of a computer device provided in an embodiment of this specification. Detailed Implementation

[0027] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0028] like Figure 1 The diagram illustrates an invalid flip of micro-operations within the UOPQ module in an embodiment of this specification. It describes how, when multiple read ports of the UOPQ module SLOT read multiple micro-operations, the DISPATCH module cannot dispatch all read micro-operations to the execution unit in a single dispatch process. In this case, the micro-operations are invalidally flipped within the read ports of the UOPQ module. In this embodiment, eight read ports are used as an example for explanation: In Cycle 0, the eight read ports SLOT of the UOPQ module read eight micro-operations (UOPs): SLOT 0-SLOT 3 are UOP(0)-UOP(3), and SLOT 4-SLOT 7 are UOP(4)-UOP(7). Due to static constraints (such as the dispatch port limitation of the DISPATCH module) or dynamic constraints (such as the execution unit being busy), only UOP(0)-UOP(3) of SLOT 0-SLOT 3 conforms to the dispatch rules, is successfully sent to the execution unit, and is consumed and cleared. The read ports SLOT 4-SLOT 7 of the UOPQ module read micro-operations UOP(4)-UOP(7), but they are not effectively dispatched. They are retained according to the FIFO rule and wait for the next dispatch process to process.

[0029] The UOP(4)-UOP(7) read ports SLOT 4-SLOT 7 of Cycle 0 that were not dispatched are moved forward as a whole to the read ports SLOT 0-SLOT 3 of Cycle 1, waiting for the second dispatch; the read pointer of the UOPQ module advances forward according to the FIFO rule, reads the new 4 micro-operations UOP(8)-UOP(11) from the queue, and fills them into the UOPQ module read ports SLOT 4-SLOT 7 of Cycle 1; for the registers of these 4 physical ports SLOT 4-SLOT 7, Cycle 0 stores UOP(4)-UOP(7), and Cycle 1 is refreshed to UOP(8)-UOP(11); in the hardware circuit, when the value of the register changes from one state to another (0 to 1, 1 to 0), a bit flip will occur, and each flip will consume dynamic power consumption; the micro-operations UOP(4)-UOP(7) are just read from the read ports SLOT 4-SLOT 7. 7. Moved to read ports SLOT 0-SLOT 3, the newly filled micro-operations UOP(8)-UOP(11) still cannot be dispatched in the second dispatch process. The flip of the register circuit of the read port of the UOPQ module does not bring any performance benefit and is an invalid flip. Similar to the first dispatch process of Cycle0, only the micro-operations UOP(4)-UOP(7) of read ports SLOT 0-SLOT 3 are effectively dispatched in the second dispatch process. The micro-operations UOP(8)-UOP(11) of read ports SLOT 4-SLOT 7 need to be moved to read ports SLOT 0-SLOT 3 again.

[0030] The dispatch process is repeated in subsequent clock cycles. The UOPQ module reads the four physical ports of SLOT 4-SLOT 7, and performs invalid toggling in each clock cycle, continuously generating invalid power consumption.

[0031] like Figure 2 The diagram shows a flowchart of an adaptive bandwidth scheduling method according to an embodiment of this specification. The diagram illustrates that in the pipeline front-end of a high-performance processor, particularly during the process of the Micro-Operation Queue (UOPQ) module reading micro-operations and sending them to the dispatch port of the instruction dispatch module, the number of read ports of the UOPQ module is dynamically adjusted to match the number of effectively dispatched micro-operations between the pipeline front-end and back-end. This reduces invalid flips of micro-operations in the UOPQ module, reduces the power consumption of the UOPQ module caused by excessive micro-operations, and improves the processor's energy efficiency. The method specifically includes: Step 201: Obtain the first bit value of each read port of the micro-operation queue UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module, wherein the first bit value of each read port of the UOPQ module includes the bit value corresponding to at least some bits in the read port, and the second bit value of each dispatch port of the DISPATCH module includes the bit value corresponding to the corresponding bit in the dispatch port. Step 202: Adjust the number of read ports of the UOPQ module according to the difference between the first bit value and the second bit value.

[0032] The method described in this specification identifies instances of excessive invalid power consumption in the UOPQ module. This reduces invalid flips in the UOPQ module without compromising pipeline efficiency, thereby lowering the power consumption of the UOPQ module and improving the processor's energy efficiency.

[0033] As an embodiment of this specification, the first bit value of each read port of the UOPQ module includes bit values ​​corresponding to at least a portion of the bits in the read port. For example, the read port SLOT 1 of the UOPQ module is 128 bits, and the bit value of the read port can be the 128-bit UOP micro-operation bit value read by the port; or, it can be the UOP micro-operation bit value of the lower 16 bits (or a portion of consecutive bits such as the higher 32 bits) read by the port; or, it can be a multi-bit bit value composed of bit values ​​corresponding to non-consecutive bits starting from the lower (or higher) bit, such as a multi-bit bit value composed of one bit value obtained every 4 bits (or other non-consecutive bits).

[0034] The second bit value of each dispatch port of the DISPATCH module includes the bit value corresponding to the corresponding bit in the dispatch port. The microstructure meaning represented by the bit in this port is the same as the microstructure meaning represented by the bit sampled by the read port of the UOPQ module. That is, the sampled bit can be a set of consecutive bits, or a set of non-consecutive bits starting from the least significant bit. Furthermore, there is a one-to-one correspondence between the dispatch port and the read port of the UOPQ module.

[0035] By comparing the first bit value and the second bit value, it can be determined whether there is an invalid flip in the UOPQ module, which would result in wasted power consumption.

[0036] In this embodiment, the number of micro-operations read by the UOPQ module read port can be determined based on the hardware resources of the UOPQ module read port. For example, if the width of the UOPQ module read port is 8, it means that 8 micro-operations (SLOT 0-SLOT 7) in the queue can be read at once. However, since the number of micro-operations dispatched by the DISPATCH module is limited by the working state of the downstream execution unit (e.g., the computing unit ALU), it may only be able to dispatch 4 micro-operations to the corresponding downstream execution unit. Alternatively, the dispatching capability of the DISPATCH module may also be limited by its own hardware resources (e.g., the number of channels, the number of ports, etc.), and it may only be able to dispatch 4 micro-operations to the downstream execution unit. Therefore, the UOPQ module read port (e.g., SLOT0-SLOT 3) acting as the micro-operation reading end can only output 4 micro-operations that can be effectively dispatched by the DISPATCH module. In the embodiments of the specification, the number of micro-operations read by the UOPQ module read port corresponding to the micro-operations that can be effectively dispatched by the DISPATCH module is referred to as the number of micro-operations effectively read by the UOPQ module read port. By adjusting the number of read ports of the UOPQ module that read micro-operations, the number of micro-operations effectively read by the UOPQ module can be matched with the number of micro-operations effectively dispatched by the downstream DISPATCH module (the actual number of micro-operations read by the UOPQ module matches the number of effectively read micro-operations read by the UOPQ module). This avoids the UOPQ module reading micro-operations far exceeding the processing capacity of the downstream DISPATCH module or execution unit, avoids invalid toggling transitions in the UOPQ module, and avoids invalid transitions, invalid circuit buffering, and movement in the downstream DISPATCH module and execution unit module, thereby effectively reducing the device's power consumption.

[0037] In another embodiment, when the DISPATCH module dispatches micro-operations, the type of the effectively dispatched micro-operation can be obtained. For example, a fixed-point arithmetic micro-operation is successfully dispatched to a fixed-point arithmetic execution unit by the DISPATCH module, and a floating-point arithmetic micro-operation is successfully dispatched to a floating-point arithmetic execution unit by the DISPATCH module. The type of the micro-operation may include, for example, floating-point arithmetic micro-operations, fixed-point arithmetic micro-operations, memory access micro-operations, etc. The type of micro-operation can also be further subdivided; for example, floating-point arithmetic micro-operations may further include specific sub-operation types such as addition, multiplication, and vector operations.

[0038] As an embodiment of this specification, obtaining the first bit value of each read port of the UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module further includes: During a dispatch micro-operation, the first bit value of each read port of the UOPQ module and the second bit value of each dispatch port of the DISPATCH module are obtained.

[0039] In this embodiment, the reading and dispatching of micro-operations adopts a pipelined approach. A single micro-operation dispatch process refers to the process by which the DISPATCH module effectively dispatches one or more micro-operations to the corresponding downstream execution units within one clock cycle. Effective dispatch means that the dispatched micro-operation can be successfully received and executed by the execution unit. For example, the eight read ports SLOT 0-SLOT of the UOPQ module... 7. Sending 8 micro-operations (UOP(0)-UOP(7)) to the DISPATCH module for dispatch. Under ideal conditions, all 8 micro-operations can be sent to the DISPATCH module at once, and the DISPATCH module will dispatch all 8 micro-operations to the corresponding execution units. In most cases, due to the software and hardware resource limitations (i.e., dynamic constraints and static constraints) of the DISPATCH module or the execution unit, the DISPATCH module cannot effectively dispatch all 8 micro-operations to the downstream corresponding execution units. For example, in one dispatch process, the DISPATCH module can only dispatch 4 of the micro-operations (UOP(0)-UOP(3)) to the corresponding execution units. At the same time, the UOPQ module reads 4 new micro-operations (UOP(8)-UOP(11)) from the 4 micro-operations that have been effectively dispatched. Then, after waiting for one or several clock cycles, the DISPATCH module can dispatch the last 4 micro-operations (UOP(4)-UOP(7)) to the corresponding execution units through a second dispatch.

[0040] As an embodiment of this specification, the bit value corresponding to the corresponding bit in the dispatch port includes: the bit value corresponding to the corresponding bit in the dispatch port where the dispatched micro-operation is located, or the preset bit value in the dispatch port where the undispatched micro-operation is located.

[0041] In this embodiment, the UOPQ module reads 8 micro-operations through its read port. However, due to the dynamic or static constraints of the DISPATCH module or the execution unit, the DISPATCH module's dispatch port can only dispatch 4 of these micro-operations (e.g., UOP(0)-UOP(3)) to the corresponding downstream execution unit. Among them, the DISPATCH module's dispatch ports SLOT 0-SLOT 3 effectively dispatch micro-operations UOP(0)-UOP(3). Since dispatch ports SLOT 4-SLOT 7 do not dispatch micro-operations, they have preset bit values, such as all 0s or all 1s. Thus, the second bit value of the DISPATCH module's dispatch port is different from the first bit value of the micro-operation in the UOPQ module's read port.

[0042] When the number of micro-operations read from the UOPQ module's port significantly exceeds the number of micro-operations effectively dispatched by the DISPATCH module over multiple consecutive clock cycles, it indicates that the DISPATCH module or a downstream execution unit of a certain type is subject to static or dynamic constraints, causing congestion in micro-operation dispatch. If dispatch is continuously slow, these micro-operations read from the UOPQ module cannot be dispatched to the corresponding execution unit in the backend in a timely manner, resulting in toggling and jumping within the UOPQ module. Simultaneously, this increases invalid jumps in the DISPATCH pipeline, leading to higher processor power consumption. In the embodiments of this specification, when the difference between the bit value of the UOPQ module's read port and the bit value of the DISPATCH module's dispatch port reaches a preset condition, reducing the number of UOPQ module read ports can alleviate congestion in micro-operation dispatch, reduce invalid toggling within the UOPQ module, and thus lower power consumption.

[0043] Similarly, when the number of read ports of the UOPQ module is reduced through the aforementioned steps, and the number of micro-operations of the UOPQ module read ports is significantly less than the number of effective micro-operations dispatched by the DISPATCH module in multiple consecutive clock cycles, the number of read ports of the UOPQ module can be increased to improve the utilization of device hardware resources and improve processor efficiency.

[0044] As one embodiment of this specification, the differences include: The number of bits whose first bit value at the read port differs from the second bit value at the dispatch port; or... The proportion of bits whose first bit value of the read port is different from the second bit value of the dispatch port to the total number of bits in all dispatch ports.

[0045] In this embodiment, for simplicity, the bit value corresponding to the micro-operation is described as 8 bits. The bit value of the micro-operation read by any read port of the UOPQ module, such as SLOT 0 port, is 00000001. The corresponding bit value of the micro-operation effectively dispatched by the corresponding dispatch port of the DISPATCH module is 00000010. After comparing the two bit values ​​bit by bit, the lower two bits of SLOT 0 port change from "01" to "10", and the number of flipped bits is 2. If there are other read ports, such as SLOT 1 port, and the micro-operation is effectively dispatched by the DISPATCH module, then the bit value of SLOT 1 read port is the same as the corresponding bit value of the corresponding dispatch port after bit by bit comparison. For simplicity, only the bit value of SLOT 0 port is flipped, then the number of different bits is 2.

[0046] In another embodiment, the total number of bits across all dispatch ports of the DISPATCH module is 128 bits, and the number of bits whose bit values ​​at the read port differ from those at the dispatch port is 64. Therefore, the invalid flip rate is 64 / 128 = 50%. Alternatively, if the total number of bits across all read ports of the UOPQ module is 128, and the number of bits whose bit values ​​differ from those of the dispatch port is 64, then the invalid flip rate is 64 / 128 = 50%.

[0047] There are several other methods to calculate the invalid flip rate. The purpose is to compare the differences between the bit values ​​of micro-operations effectively dispatched by the dispatch port and the bit values ​​of micro-operations read by the read port. This reflects the fact that a large number of micro-operations read by the read port were not effectively dispatched, resulting in more invalid flips and wasted energy in the UOPQ module. The number of bits that differ between the read port and the dispatch port represents the number of bit flips. When the dispatch module cannot effectively dispatch micro-operations, a higher number of bit flips indicates greater wasted energy consumption. By adjusting the number of read ports in the UOPQ module, the number of micro-operations read by the read ports can be reduced, decreasing invalid flips and thus reducing energy consumption.

[0048] As an embodiment of this specification, adjusting the number of UOPQ module read ports based on the difference between the first bit value and the second bit value further includes: when the difference is greater than a first threshold, adjusting the number of UOPQ module read ports in a state of excessive power consumption.

[0049] In this implementation, when there are many bits that are different between the first bit value of the read port and the second bit value of the dispatch port, such as exceeding the first threshold, the state is determined to be a power consumption overrun state, indicating that the power consumption loss of both the UOPQ module and the Dispatch module exceeds expectations.

[0050] In this step, by comparing the number of bits with different bit values ​​with the first threshold, some interference items can be filtered out, such as excluding some cases with few bit flips, without having to make subsequent continuous judgments, thus saving system resources.

[0051] As an embodiment of this specification, adjusting the number of read ports of the UOPQ module to address a power consumption exceeding limit further includes: If the power consumption exceeds the limit and the clock cycle lasts for more than the second threshold, then the number of read ports of the UOPQ module is reduced. If the clock cycle of the excessive power consumption state is less than the third threshold, then the number of read ports of the UOPQ module is increased.

[0052] In this embodiment, taking the number of bits that differ between the first bit value of the read port and the second bit value of the dispatch port as an example, when the number of bits that differ between the read port and the dispatch port is 40, 30, and 40 respectively in multiple clock cycles of three consecutive dispatch micro-operations, it exceeds the first threshold (the first threshold for the number of differences is, for example, 30). This indicates that the DISPATCH module has been unable to effectively dispatch all eight micro-operations read by the eight read ports of the UOPQ module for multiple clock cycles, indicating that there are many invalid flips in the UOPQ module, and the number of clock cycles exceeds the second threshold (or the number of dispatches with the number of differences exceeding the first threshold is greater than the second threshold can be used as the judgment criterion). It is necessary to reduce the number of read ports of the UOPQ module, so that the number of new micro-operations read by the read ports in subsequent clock cycles is reduced, the number of movement micro-operations will also be reduced, thereby reducing the power consumption of invalid flips.

[0053] Meanwhile, when the number of bits whose bit values ​​at the read port differ from those at the dispatch port is 30 and 40 respectively in two consecutive micro-operation dispatches, although the number of differences is greater than the first threshold (e.g., 30), the number of clock cycles in which the power consumption exceeds the limit is less than or equal to the third threshold (or the criterion can be that the number of dispatches in which the number of differences is greater than the first threshold is less than the third threshold), it indicates that the DISPATCH module has been effectively dispatching the vast majority of micro-operations read by the eight read ports of the UOPQ module for multiple clock cycles. The utilization rate of downstream execution units may be low, and there may be idle execution units. There are basically no excessive invalid flips within the UOPQ module, so more micro-operations can be read and dispatched to downstream execution units, thereby improving the efficiency of the processor. For example, the number of read ports of the UOPQ module can be increased (e.g., restored to the maximum number of read ports), so that the number of new micro-operations read by the read ports in subsequent clock cycles will increase.

[0054] As an embodiment of this specification, before adjusting the number of read ports of the UOPQ module based on the difference between the first bit value and the second bit value, the following steps are also included: Match the first bit value and the second bit value according to the thread corresponding to the first bit value and the second bit value.

[0055] In this embodiment, the microarchitecture may include various structures. One structure includes multiple instruction dispatch units corresponding to multiple threads, while another structure involves multiple threads sharing the same instruction dispatch unit. When multiple threads share the same instruction dispatch unit, the thread is marked in the micro-operation to be dispatched to distinguish which thread the micro-operation belongs to. In this embodiment, the micro-operation bit value obtained from the UOPQ module read port and the micro-operation bit value obtained from the DISPATCH module dispatch port both have their own thread markers. Based on the thread markers, it can be distinguished whether the bit value from the read port and the bit value from the dispatch port belong to the same thread, and subsequent difference judgments are performed on the bit values ​​of the same thread.

[0056] As an embodiment of this specification, adjusting the number of read ports of the UOPQ module includes at least one of the following methods: The number of UOPQ module read ports is adjusted by using the adjustment enable signal of the UOPQ module read port; or, By controlling the number of micro-operations written to the UOPQ module read port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ module read port, the number of UOPQ module read ports is adjusted; or, By controlling the number of micro-operations written to the UOPQ module read port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ module read port, and after the micro-operations in the UOPQ module have been dispatched, the number of UOPQ module read ports is adjusted; or, The number of read ports of the UOPQ module is adjusted by triggering the FLUSH operation at the front end of the pipeline through the REDIRECT action.

[0057] In this embodiment, the number of UOPQ module read ports can be controlled by the enable signal of the UOPQ module read port. This method can achieve UOPQ module read port adjustment without delay. Figure 3 The diagram shown is a schematic of the UOPQ module read port adjustment circuit according to an embodiment of this specification. This diagram illustrates the internal adjustment circuit of the UOPQ module read port, such as... Figure 3 The 5-level registers and write / read control logic shown receive the decision-maker instruction from the UOPQ module to read the port adjustment (decrease or increase), and complete the port width switching within one cycle (or switch directly without delay). When the adjustment is reversed, the operation is abandoned to prioritize the adjustment speed.

[0058] The UOPQ module includes five register levels: register 1, register 2, register 3, register 4, and register 5. The register groups are register 2 and register 3. Register 1 is the enable signal, indicating whether to reduce the number of UOPQ module read ports. The default value is 0 (no reduction in the number of ports). If the enable signal is 1, the number of UOPQ module read ports is directly reduced (i.e., the current number of UOPQ module read ports is reduced), directly reducing the power consumption waste introduced by invalid UOPQ toggling.

[0059] Register 2 is a buffer for reading micro-operations from the UOPQ module's read port, including 8 SLOTs (valid0-7 marks whether it is valid, uops0-7 stores the micro-operation), which directly output UOP to subsequent modules, as shown in Table 9 in the figure.

[0060] Register 3 is the main storage array for the UOPQ module read port, as shown in Table 8 in the figure, which stores more UOPs to be buffered.

[0061] Registers 4 and 5 are RdPtr (read pointer) and RdPtrPlus (read pointer + 1) respectively, which control the range of SLOT read from register 3. Register 5 can solve the timing problem of the circuit.

[0062] Read controls 6, 7, and 12 select valid UOPs in register 3 based on RdPtr / RdPtrPlus and output them to subsequent modules. Specifically, based on registers 4 and 5, the number of valid SLOTs read from the UOPQ module read port is updated, and new UOP data information is selected through the updated read controls 6 and 7, as shown in read control 12.

[0063] Upon receiving the read control enable information, read control 10 outputs the valid UOPs in Table 9 to the UOPQ module according to the adjusted valid read ports.

[0064] Write control 13 implements the write logic of UOP. Type 1 is: if the number of valid micro-operations in register 3 is less than the number of valid read ports of UOPQ, the micro-operations in register 3 and the newly written micro-operations are both written to register 2. Type 2 is: when the number of valid micro-operations in register 3 is greater than the number of valid read ports of UOPQ, the valid micro-operations in register 3 are selected and written to register 2. At the same time, the valid micro-operation data in write operation 13 will be written to register 3.

[0065] When controlling the number of UOPQ module read ports to decrease the number of UOPQ module read ports, the number of UOPQ module read ports is reduced through the following steps: Step 1: Trigger a reduction in the number of read ports for the UOPQ module by setting the En signal.

[0066] In this step, when the peripheral circuit determines that the number of UOPQ module read ports needs to be adjusted, it inputs Enable=1 into register 1 (indicating a reduction in the number of UOPQ module read ports). At this time, the circuit starts polling to check the status of valid0-7 in register 2 (to determine whether the micro-operations of the current UOPQ module read ports can be cleared).

[0067] Step 2: Enter the emptying state and stop receiving new micro-operations.

[0068] In this step, the circuit control write control 13 pauses writing new micro-operations to register 2 (regardless of the number of valid micro-operations in register 3, register 2 will no longer be filled); at the same time, read control 6, 7, and 12 continue to read the UOPs in register 2 according to the original bandwidth logic until all valid0-7 in register 2 are set to 0, that is, all 8 UOPs in register 2 are dispatched, and the emptying is completed.

[0069] Step 3: Fill in the UOPQ module with the UOP corresponding to the read port after adjustment.

[0070] In this step, after register 2 is emptied, the circuit switches write control 13 to type 1 or type 2. If the number of valid micro-operations in register 3 is greater than 2, the new micro-operation is first written to register 3; then, only 2 valid UOPs are taken from register 3 and filled into the first 2 SLOTs of register 2 (valid0-1 is set to 1, uops0-1 stores UOP), while the last 6 SLOTs (valid2-7) remain invalid. If the number of valid data in register 3 is less than 2, and there is a valid micro-operation in register 3, 1 valid micro-operation in register 3 and the newly written micro-operation are combined to form 2 valid UOPs, which are then filled into the first 2 SLOTs of register 2 (valid0-1 is set to 1, uops0-1 stores UOP), while the last 6 SLOTs (valid2-7) remain invalid. If the number of valid data in register 3 is 0, 2 valid UOPs are taken from the newly written micro-operation and filled into the first 2 SLOTs of register 2 (valid0-1 is set to 1, uops0-1 stores UOP), while the last 6 SLOTs (valid2-7) remain invalid.

[0071] Step 4: Turn off the unused register clock.

[0072] In this step, the RdPtr / RdPtrPlus range of registers 4 and 5 is reduced to 0-1, and UOP is read only from the first two SLOTs of register 2 to achieve read control update; the circuit turns off the register clock corresponding to SLOT2-7 in register 2, and the registers of these SLOTs stop working and no longer consume power; at this time, register 2 only outputs the UOP of the first two SLOTs to the DISPATCH module to achieve the purpose of reducing the number of UOPQ ports.

[0073] When the number of read ports of the UOPQ module is controlled to increase, the number of read ports of the UOPQ module is increased by the following steps: Step 1: Trigger increase, set the En signal to 0.

[0074] In this step, when the peripheral circuit determines that an additional port needs to be added, it inputs Enable=0 into register 1. The circuit then polls the valid0-1 status of register 2 and waits for it to be cleared.

[0075] Step 2: Empty the micro-operations of the current two ports.

[0076] In this step, write control 13 pauses writing data to register 2, and simultaneously pauses reading data from register 3 and writing it into register 2. Read control continues reading the first two slots of register 2 until valid0-1 is set to 0 (both UOPs are dispatched).

[0077] Step 3: Prefetch 8 UOPs and fill 8 SLOTs in register 2.

[0078] In this step, write control 13 switches back to type 1. If the number of valid micro-operations in register 3 is less than 8, all micro-operations in register 3 and the newly read micro-operations are written to register 2. Eight valid UOPs are taken from register 3 or the input terminal and filled into the eight SLOTs of register 2 (valid0-7 are all set to 1). Alternatively, write control 13 switches back to type 2. Write control 13 writes the latest data into register 3, and at the same time, writes the micro-operations in register 3 into the eight SLOTs of register 2 (valid0-7 are all set to 1).

[0079] Step 4: Turn on the clock and restore the 8 bandwidths to work.

[0080] In this step, RdPtr / RdPtrPlus is restored to the range of 0-7, covering all SLOTs to achieve read control update; the register clock of SLOT2-7 in register 2 is turned on, and all 8 SLOTs resume operation; register 2 outputs the UOP of the 8 SLOTs to the DISPATCH module to increase the number of UOPQ ports.

[0081] Through the UOPQ module read port adjustment circuit in the embodiments of this specification, and the read / write control logic for reducing and increasing UOPQ module read ports, register 3 can be used to buffer new input micro-operations. The read control logic reads the micro-operations in register 2 and sends them to the downstream processing logic of the pipeline (e.g., the DISPATCH module) for dispatching micro-operations. When the micro-operations in register 2 are emptied, the write control logic can obtain the adjusted number of valid read micro-operations and write them into register 2, or obtain the adjusted number of valid read micro-operations from register 3 and write them into register 2, or obtain the adjusted number of valid read micro-operations from register 3 and the write control logic and write them into register 2. By controlling the number of UOPQ module read ports, invalid flips in the UOPQ module are reduced.

[0082] In other embodiments, it may also be as follows: Figure 4The diagram illustrates a blocking write operation in an embodiment of this specification. By controlling the number of micro-operations written to the UOPQ module read ports by the pipeline front end, and waiting for all micro-operations from the pipeline front end to be written to the UOPQ module read ports, the number of UOPQ module read ports is adjusted. When an instruction to adjust the number of UOPQ module read ports is generated (e.g., reducing the number of read ports from 8 to 2), a STALL signal is sent to the instruction decoding channel and the OC (Op Cache) fetch channel; after waiting for N cycles (e.g., N=3 for a 3-stage pipeline), it is ensured that there are no micro-operations in the pipeline that have not been written to UOPQ; then the number of UOPQ module read ports is adjusted, the read pointer is switched to SLOT0-1, and the clocks of SLOT2-7 are turned off; after confirming that the UOPQ module read port configuration is complete, the STALL signal is set low, allowing the instruction decoding channel and the OC fetch channel to write new micro-operations to the UOPQ module read ports; the UOPQ module reads micro-operations using 2 read ports. For example, the decision maker determines that the number of read ports of the UOPQ module needs to be reduced to 2, and sends a STALL signal to the instruction decoding channel; waits for 3 cycles, and all 3 in-flight micro-operations (micro-operations that have not been dispatched) in the pipeline are written to the corresponding read ports of the UOPQ module and dispatched; adjusts the read pointer to SLOT0-1, and turns off the clock of the remaining 6 SLOTs; sets the STALL signal low, and the new micro-operation is written to the UOPQ module through 2 read ports to avoid conflicts caused by data residue during switching.

[0083] In other embodiments, such as Figure 4 As shown, the number of micro-operations written to the UOPQ module read port by the pipeline front end can also be controlled. After all the micro-operations at the pipeline front end are written to the UOPQ module read port, and after the micro-operations in the UOPQ module have been dispatched, the number of UOPQ module read ports can be adjusted. The difference from the previous embodiment is that in this embodiment, after all the micro-operations at the pipeline front end are written to the UOPQ module read port, it is also necessary to wait for the micro-operations in the UOPQ module to be dispatched before adjusting the number of UOPQ module read ports. For simplicity, the similarities between this embodiment and the previous embodiment will not be repeated.

[0084] In other embodiments, it may also be as follows: Figure 5The diagram illustrates an embodiment of this specification where the pipeline front-end is cleared to adjust the number of UOPQ module read ports. A FLUSH operation is triggered at the pipeline front-end by a REDIRECT action. At this time, the micro-operations in the UOPQ module read ports are also cleared. The number of UOPQ module read ports is adjusted using the enable signal of the UOPQ module read ports. By using the REDIRECT action to FLUSH the pipeline front-end (branch prediction unit, instruction fetch unit, etc.), all unfinished micro-operations in the UOPQ module read ports are cleared, and the execution start point is redirected, readjusting the number of UOPQ module read ports and simplifying the adjustment circuit design. For example, the decision-maker determines that the number of read ports of the UOPQ module needs to be increased to 8, triggering the REDIRECT action; the REDIRECT signal triggers FLUSH, clearing the residual micro-operations or instructions in the instruction fetch unit, instruction decoding channel, and OC instruction fetch channel at the front end of the pipeline; while clearing the read ports of the UOPQ module, the number of read ports of the UOPQ module is adjusted, the read pointer covers slots 0-7, and all SLOT clocks are turned on; REDIRECT specifies a new execution start point (the current program counter PC), and the pipeline fetches and decodes instructions from this address again; the new micro-operations are written to the read ports of the UOPQ module with a bandwidth of 8 ports, the system returns to normal, and the adjustment of the number of read ports is completed.

[0085] For example, in the current state: the UOPQ module read port count is 2 (i.e., SLOT is 2), the historical FLUSH recovery time is 3 cycles, the current IPC is 95% of the baseline value, and the maximum token capacity of the execution unit is 8. The decision-maker determines that the UOPQ module read port bandwidth needs to be restored for 8 micro-operations, triggering the REDIRECT action; FLUSH clears the 3 pending instructions in the instruction fetch unit and the 2 undisassembled micro-operations in the instruction decoding channel; the UOPQ module read port is adjusted to 8, the read pointer covers all 8 SLOTs, and the clock is fully enabled; the pipeline fetches instructions again from the current PC value, 8 new micro-operations are written to the UOPQ module read port, the DISPATCH module dispatches according to the bandwidth of 8 micro-operations, and the performance is fully restored after 3 cycles, with no data interference during the adjustment process.

[0086] As one embodiment of this specification, reducing the number of read ports of the UOPQ module further includes: Monitor the number of micro-operations effectively dispatched by the DISPATCH module; If the number of micro-operations effectively dispatched by the DISPATCH module remains below a preset threshold, the number of read ports of the UOPQ module will be increased.

[0087] In this embodiment, when the number of UOPQ module read ports is reduced, the number of micro-operations effectively dispatched by the DISPATCH module remains below a preset threshold. To prevent the processor performance from degrading, the number of UOPQ module read ports can be increased to restore the full micro-operation read capability of the UOPQ module read ports.

[0088] In another embodiment, if the difference between the number of micro-operations that can be dispatched by the DISPATCH module and the number of micro-operations that are effectively dispatched by the DISPATCH module continuously exceeds a certain preset threshold after the number of micro-operations effectively read by the UOPQ module read port is reduced, the number of UOPQ module read ports can be increased, for example, it can be increased to the maximum.

[0089] The duration can be a predetermined time period, multiple consecutive clock cycles, or multiple consecutive valid dispatches; the micro-operations effectively dispatched by the DISPATCH module refer to the micro-operations dispatched by the DISPATCH module to the downstream execution units of the pipeline that can be executed, and the micro-operations that the DISPATCH module can dispatch refer to the micro-operations that the downstream execution units of the pipeline can still receive and process according to the DISPATCH module, that is, the available resources of the execution units.

[0090] By monitoring the number of micro-operations effectively dispatched by the DISPATCH module, the number of read ports of the UOPQ module can be increased in real time, thereby improving the processor's processing efficiency and avoiding the problem of excessive adjustment of the number of read ports of the UOPQ module, which would reduce the processor's processing efficiency.

[0091] As one embodiment of this specification, after adjusting the number of effective read micro-operations of the UOPQ port, the following is also included: Monitor changes in processor performance (Instructions Per Cycle, IPC); As the DISPATCH module completes the dispatch of micro-operations, if the IPC remains below a preset threshold, the number of micro-operations effectively read by the UOPQ port will be increased.

[0092] In this step, before adjusting the number of UOPQ module read ports, the IPC value under stable processor conditions is recorded in the history. For example, an average IPC of 5 over 10 consecutive cycles serves as a preset threshold for judging performance degradation. If the IPC falls below the preset threshold within a certain duration or clock cycle, performance degradation can be considered, triggering an error correction mechanism to increase the number of UOPQ module read ports, for example, by increasing it to the maximum.

[0093] In another embodiment, after reducing the number of UOPQ module read ports, the number of execution units (TOKENs) at the back end of the processor pipeline is monitored; as the DISPATCH module completes the dispatch of micro-operations, the number of TOKENs corresponding to the micro-operation type continues to be higher than a preset threshold, and there are still micro-operation types of this type in the DISPATCH module, then the number of UOPQ module read ports is adjusted to be increased.

[0094] By monitoring changes in processor performance IPC through this embodiment, the number of UOPQ module read ports can be increased in real time, thereby improving the processor's processing efficiency and avoiding the problem of excessive adjustment of UOPQ module read ports, which would reduce the processor's processing efficiency.

[0095] like Figure 6 The diagram shown is a schematic representation of the adaptive bandwidth scheduling device according to an embodiment of this specification. This diagram illustrates a device for operating the methods described in the above embodiments. Each functional module can be implemented through circuits, or the functions of each module can be accomplished through a combination of instruction sets and circuits. Specifically, the device includes: The acquisition unit 601 is configured to acquire the first bit value of each read port of the operation queue UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module, wherein the first bit value of each read port of the UOPQ module includes the bit value corresponding to at least a portion of the bits in the read port, and the second bit value of each dispatch port of the DISPATCH module includes the bit value corresponding to the corresponding bit in the dispatch port. Decision maker 602 is configured to adjust the number of read ports of the UOPQ module based on the difference between the first bit value and the second bit value.

[0096] like Figure 7 The diagram shown is a structural schematic of the instruction dispatch unit according to an embodiment of this specification. The internal structure of the instruction dispatch unit is illustrated in this diagram. The functional components can be implemented using hardware circuitry or a combination of instruction sets and hardware circuitry. Not all functional modules are required in this diagram; only a portion of the functional modules are needed to achieve the desired purpose. Specifically, the instruction dispatch unit includes: UOPQ module 701, intermediate processing module 702, dispatch module 703, information collection module 704, TOKEN module 705, flip recognition module 706, historical record module 707, decision module 708, performance monitoring module 709, execution unit 710.

[0097] The UOPQ module 701 is configured to adjust the number of read ports according to the control of the decision module 708 and transmit the micro-operation to the intermediate processing module 702. The intermediate processing module 702 is configured to send the micro-operations read by the UOPQ module 701 to the dispatch module 703.

[0098] The dispatch module 703 is configured to dispatch micro-operations to the corresponding execution unit 710; The type identification module 704 is configured to identify the type of micro-operation read by the UOPQ module 701 and the type of micro-operation effectively dispatched by the dispatch module 703; The TOKEN module 705 is configured to acquire available resources of the execution unit 710; The flip recognition module 706 is configured to acquire the first bit value of each read port of the micro-operation queue UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module. The history record module 707 is configured to record the first bit value and the second bit value during the multiple dispatch micro-operations obtained by the type recognition module 704 and the flip recognition module 706, and also includes the static constraints and dynamic constraints of the intermediate processing module 702. The decision module 708 is configured to adjust the number of read ports of the UOPQ module 701 based on the difference between the first bit value and the second bit value recorded by the history module 707.

[0099] The performance monitoring module 709 is configured to monitor the system performance of the processor and feed back the changes in system performance to the decision module 708, so that the decision module 708 can adjust the number of read ports of the UOPQ module 701 according to the changes in system performance.

[0100] The following describes the process steps of an instruction dispatch unit with the above structure to reduce invalid toggling by reducing the number of UOPQ module read ports: Step 1: Collect data from multiple modules.

[0101] In this step, the type recognition module 704, the TOKEN module 705, the flip recognition module 706, the history module 707, and the performance monitoring module 709 work synchronously to output full-dimensional data to the decision module 708 to ensure accurate decision-making.

[0102] Specifically, the flip recognition module 706 collects in real time the bit values ​​of the micro-operations at the read port of the UOPQ module 701 and the bit values ​​of the micro-operations that are effectively dispatched at the dispatch port of the dispatch module 703.

[0103] The TOKEN module 705 iterates through fixed-point (EX), floating-point (FP), and memory access (LSU) execution units to count the available TOKENs for each execution unit.

[0104] The type identification module 704 parses the micro-operations of the DISPATCH GROUP of the dispatch module 603, classifies them by function (fixed-point / floating-point / memory access, which may include subdivided types such as addition / multiplication / vector), and marks the micro-operation type.

[0105] The history record module 707 retrieves the second bit value of each dispatch port of the dispatch module 703, the first bit value of each read port of the UOPQ module 701, the number of TOKENs of each execution unit, the micro-operation type, and the port adjustment record from multiple historical cycles.

[0106] The performance monitoring module 709 collects real-time values ​​of the current IPC (instructions per cycle) and the number of DISPATCH micro-operations, compares them with historical baseline values ​​(stable values ​​before adjustment), and outputs the performance fluctuation status.

[0107] All module output data is synchronously transmitted to decision module 708 via the internal bus.

[0108] Step 2: The decision module 708 calculates the difference between the first bit value of each read port of the UOPQ module 701 and the second bit value of each dispatch port of the dispatch module 703.

[0109] In this step, the decision module 708, in conjunction with the bit values ​​collected by the flip recognition module 706 (or the historical record module 707), compares the first bit value of each read port with the second bit value of the dispatch port bit by bit, calculates the number of bits with different bit values, and obtains the difference. If the difference is greater than a first threshold, it proceeds to the subsequent judgment step 3; if it is less than the first threshold, it returns to step 1 to continue collecting data.

[0110] Step 3: The decision module 708 determines whether the number of clock cycles during which the difference persists is greater than the second threshold, or whether the number of clock cycles during which the difference persists is less than the third threshold.

[0111] In this step, if the difference is greater than the first threshold and the number of consecutive clock cycles is greater than the second threshold, it indicates that the effective dispatch micro-operations of the DISPATCH module are continuously congested, and there are many invalid flips in the bit values ​​of the read micro-operations. The number of read ports of the UOPQ module 701 needs to be reduced. If the difference is greater than the first threshold and the number of consecutive clock cycles is less than the third threshold, it indicates that the effective dispatch micro-operations of the DISPATCH module are continuously unimpeded, and the number of read ports of the UOPQ module 701 can be increased.

[0112] Step 4: The decision module 708 adjusts the number of read ports of the UOPQ module 701 based on the difference judgment result.

[0113] In this step, the decision module 708 sends an instruction to the UOPQ module 701 to reduce the number of read ports, and the intermediate processing module 702 synchronously adapts to the new port width.

[0114] Specifically, the decision module 708 outputs an instruction and simultaneously sends a port adjustment notification to the intermediate processing module 702 (informing that the new reading range is SLOT 0-3).

[0115] Refer to the aforementioned attached diagram. Set the Enable signal of register 1 in UOPQ module 701 to 1 (start port reduction); poll the valid status of register 2 and enter the emptying stage; adjust the range to SLOT 0-3 through registers 4 / 5 and stop reading SLOT 4-7; the clock gating module turns off the register clock of SLOT 4-7 to reduce the power consumption of invalid toggling; after emptying is completed, prefetch 4 UOPs from the main memory array (register 3) and fill SLOT 0-3.

[0116] After receiving the port adjustment notification, the intermediate processing module 702 adjusts the working range of the Mapping and Arbiter, processing only the UOPs of SLOT 0-3 output by the UOPQ module 701 to avoid data flow outside the range and reduce circuit transitions. The number of ports of the UOPQ module 701 becomes 4, outputting only 4 UOPs / cycle, matching the effective dispatch of the dispatch module 703 and eliminating bandwidth differences.

[0117] Step 5: The decision module 708 restores the reading port of the UOPQ module 701 based on the monitoring results of the performance monitoring module 709.

[0118] In this step, the performance monitoring module 709 acquires processor performance data in real time, including IPC and the number of micro-operations effectively dispatched by the dispatch module 703. The decision module 708 decides whether to increase the read port resources of the UOPQ module 701 based on the processor performance data.

[0119] Specifically, when the available resources (TOKEN) of the execution unit are continuously greater than a preset threshold, the decision module 708 can adjust the number of read ports of the UOPQ module 701. For specific methods, please refer to the aforementioned embodiments.

[0120] The methods and apparatus described in the embodiments of this specification can reduce the power consumption of the processor without changing its performance.

[0121] In one embodiment of this specification, a processor including the above-described instruction dispatch unit can execute the aforementioned adaptive bandwidth scheduling method to reduce the processor's power consumption.

[0122] like Figure 8The diagram illustrates a computer device according to an embodiment of this specification. The computer device in this embodiment may include the processor described in this specification and utilize the aforementioned adaptive bandwidth scheduling method. This method can also be run on the computer device in this embodiment to execute the methods described in this specification. The computer device 802 may include one or more processors 804, such as one or more central processing units (CPUs), each of which can implement one or more hardware threads. The computer device 802 may also include any memory 806 for storing information of any kind, such as code, settings, data, etc. Non-limitingly, for example, memory 806 may include any type of RAM, any type of ROM, flash memory, hard disk, optical disk, etc. More generally, any memory can use any technology to store information. Furthermore, any memory can provide volatile or non-volatile retention of information. Furthermore, any memory can represent a fixed or removable component of the computer device 802. In one case, when the processor 804 executes associated instructions stored in any memory or combination of memories, the computer device 802 can perform any operation of the associated instructions. The computer device 802 also includes one or more drive mechanisms 808 for interacting with any memory, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.

[0123] Computer device 802 may also include an input / output module 810 (I / O) for receiving various inputs (via input device 812) and providing various outputs (via output device 814). A specific output mechanism may include a presentation device 816 and an associated graphical user interface (GUI) 818. In other embodiments, the input / output module 810 (I / O), input device 812, and output device 814 may be omitted, and the device may function solely as a computer device within a network. Computer device 802 may also include one or more network interfaces 820 for exchanging data with other devices via one or more communication links 822. One or more communication buses 824 couple the components described above together.

[0124] Communication link 822 can be implemented in any way, such as via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 822 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.

[0125] This specification also provides computer-readable instructions, wherein when a processor executes the instructions, the program therein causes the processor to perform the methods described above.

[0126] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the methods described above.

[0127] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method.

[0128] It should be understood that in the various embodiments of this specification, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this specification.

[0129] It should also be understood that, in the embodiments of this specification, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this specification generally indicates that the preceding and following related objects have an "or" relationship.

[0130] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this specification.

[0131] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0132] In the several embodiments provided in this specification, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or units, or they may be electrical, mechanical, or other forms of connection.

[0133] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments described in this specification, depending on actual needs.

[0134] Furthermore, the functional units in the various embodiments of this specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0135] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this specification, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this specification. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0136] This specification uses specific embodiments to illustrate the principles and implementation methods of this specification. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this specification. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this specification. Therefore, the content of this specification should not be construed as a limitation of this specification.

Claims

1. An adaptive bandwidth scheduling method, characterized in that, include: The first bit value of each read port of the micro-operation queue UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module are obtained. The first bit value of each read port of the UOPQ module includes the bit value corresponding to at least a portion of the bits in the read port, and the second bit value of each dispatch port of the DISPATCH module includes the bit value corresponding to the corresponding bit in the dispatch port. The number of read ports of the UOPQ module is adjusted based on the difference between the first bit value and the second bit value.

2. The method according to claim 1, characterized in that, Obtaining the first bit value of each read port of the UOPQ microoperation queue module and the second bit value of each dispatch port of the DISPATCH instruction dispatch module further includes: During a dispatch micro-operation, the first bit value of each read port of the UOPQ module and the second bit value of each dispatch port of the DISPATCH module are obtained.

3. The method according to claim 2, characterized in that, The bit value corresponding to the corresponding bit in the dispatch port includes: the bit value corresponding to the corresponding bit in the dispatch port where the dispatched micro-operation is located, or the preset bit value in the dispatch port where the undispatched micro-operation is located.

4. The method according to claim 3, characterized in that, At least some of the bit values ​​corresponding to the read port include: The bit values ​​corresponding to multiple consecutive bits in the read port; or, The bit values ​​corresponding to multiple non-contiguous bits in the read port.

5. The method according to claim 1, characterized in that, The differences include: The number of bits whose first bit value at the read port differs from the second bit value at the dispatch port; or... The proportion of bits whose first bit value of the read port is different from the second bit value of the dispatch port to the total number of bits in all dispatch ports.

6. The method according to claim 5, characterized in that, Adjusting the number of read ports of the UOPQ module based on the difference between the first bit value and the second bit value further includes: When the number of differences exceeds the first threshold, the number of read ports of the UOPQ module is adjusted to a state of excessive power consumption.

7. The method according to claim 6, characterized in that, Adjusting the number of read ports of the UOPQ module to address excessive power consumption further includes: If the power consumption exceeds the limit and the clock cycle lasts for more than the second threshold, then the number of read ports of the UOPQ module is reduced. If the clock cycle of the excessive power consumption state is less than the third threshold, then the number of read ports of the UOPQ module is increased.

8. The method according to claim 7, characterized in that, When the excessive power consumption condition persists for more than a second threshold for a given number of clock cycles, further reducing the number of read ports of the UOPQ module includes: If, during multiple micro-operations, the power consumption exceeds the second threshold for several consecutive clock cycles, the number of read ports of the UOPQ module will be reduced.

9. The method according to claim 1, characterized in that, Before adjusting the number of read ports of the UOPQ module based on the difference between the first bit value and the second bit value, the following steps are also included: Match the first bit value and the second bit value according to the thread corresponding to the first bit value and the second bit value.

10. The method according to claim 1, characterized in that, Adjusting the number of read ports of the UOPQ module includes at least one of the following methods: The number of UOPQ module read ports is adjusted by using the adjustment enable signal of the UOPQ module read port; or, By controlling the number of micro-operations written to the UOPQ module read port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ module read port, the number of UOPQ module read ports is adjusted; or, By controlling the number of micro-operations written to the UOPQ module read port at the pipeline front end, and waiting for all micro-operations at the pipeline front end to be written to the UOPQ module read port, and after the micro-operations in the UOPQ module have been dispatched, the number of UOPQ module read ports is adjusted; or, The number of read ports of the UOPQ module is adjusted by triggering the FLUSH operation at the front end of the pipeline through the REDIRECT action.

11. The method according to claim 7, characterized in that, After reducing the number of read ports of the UOPQ module, the following is also included: Monitor the number of micro-operations effectively dispatched by the DISPATCH module; If the number of micro-operations effectively dispatched by the DISPATCH module remains below a preset threshold, the number of read ports of the UOPQ module will be increased.

12. The method according to claim 7, characterized in that, After reducing the number of read ports of the UOPQ module, the following is also included: Monitor changes in processor performance IPC; As the DISPATCH module completes the dispatch of micro-operations, if the IPC remains below a preset threshold, the number of ports read by the UOPQ module will be increased.

13. The method according to claim 7, characterized in that, After reducing the number of read ports of the UOPQ module, the following is also included: Monitor the number of execution unit tokens at the back end of the processor pipeline; As the DISPATCH module completes the dispatch of micro-operations, if the number of TOKENs corresponding to a certain type of micro-operation continues to exceed a preset threshold, and if there are still micro-operations of that type in the DISPATCH module, then the number of read ports of the UOPQ module is increased.

14. An adaptive bandwidth scheduling device, characterized in that... include: The acquisition unit is configured to acquire the first bit value of each read port of the operation queue UOPQ module and the second bit value of each dispatch port of the instruction dispatch DISPATCH module, wherein the first bit value of each read port of the UOPQ module includes the bit value corresponding to at least a portion of the bits in the read port, and the second bit value of each dispatch port of the DISPATCH module includes the bit value corresponding to the corresponding bit in the dispatch port. The decision-maker is configured to adjust the number of read ports of the UOPQ module based on the difference between the first bit value and the second bit value.

15. An instruction dispatching unit, characterized in that, The instruction dispatching unit executes the adaptive bandwidth scheduling method according to any one of claims 1-13.

16. A processor, characterized in that... Includes the instruction dispatching unit as described in claim 15.

17. A computer device comprising a memory and a computer program stored in the memory and executable on a processor, characterized in that, Includes the processor as described in claim 16 above.