Processor instruction dispatching method and apparatus

CN122816705APending Publication Date: 2026-09-25JINDI SPACE TIME (ZHUHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610964308.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

上述方式实现相对简单,但容易出现某些发射队列长期拥塞、另一些发射队列空闲的情况,从而降低整体并行度

Benefits of technology

[0011]相较于现有技术,本申请实施例中,通过获取待派遣的多个微操作,并结合微操作类型和年龄顺序确定各微操作对应的候选发射队列集合,使不同类型的微操作能够在满足类型约束的候选范围内参与派遣选择。通过获取总体发射队列集合中各发射队列的剩余容量信息和有效占用信息,并确定容量受限发射队列,使初始队列分配过程能够避开或减少选择负载较高的发射队列,从而降低部分发射队列持续拥塞、其他发射队列利用不足的概率。进而,本申请实施例根据多个微操作的寄存器读写信息以及各微操作对应的初始发射队列,校验寄存器依赖关系与发射队列之间的关联关系,并确定各微操作对应的依赖发射队列,使派遣过程能够感知生产者微操作与消费者微操作之间的寄存器依赖关系。在检验通过后,根据各微操作对应的依赖发射队列将多个微操作派遣至对应的发射队列,有利于减少寄存器相关微操作在后续唤醒、旁路和队列调度过程中的路径开销。本申请实施例能够在派遣阶段同时兼顾微操作类型约束、微操作年龄顺序、发射队列容量状态以及寄存器依赖关系,提高多个发射队列之间的负载均衡程度,降低热点发射队列拥塞风险,并减少寄存器相关微操作的依赖相关延迟,从而提升处理器派遣阶段和后续发射阶段的整体效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816705A_ABST
    Figure CN122816705A_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the chip and semiconductor technical field, and provides a processor instruction dispatching method and device, which comprises the following steps: obtaining a plurality of micro-operations to be dispatched; determining a candidate emission queue set of each micro-operation according to a micro-operation type and an age order; obtaining queue state information of a total emission queue set, and determining a capacity-limited emission queue; performing initial queue allocation based on the candidate emission queue set and the capacity-limited emission queue; checking an association relationship between a register dependency relationship and an emission queue according to register read-write information, and determining a dependent emission queue; and after the checking, dispatching the plurality of micro-operations to corresponding emission queues. Therefore, under the premise of meeting the age order and type constraints, the embodiment of the application can improve queue utilization balance, reduce dependency correlation delay, and improve the overall efficiency of the processor dispatching stage and the subsequent emission stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chip and semiconductor technology, and more specifically to a processor instruction dispatch method and apparatus. Background Technology

[0002] In a processor, instructions typically go through a pipelined process of fetch, decode, rename, dispatch, issue, and execution. For out-of-order execution processors, micro-operations generated in the front end can be issued and executed out of order while satisfying data dependencies and resource constraints, thereby improving instruction-level parallelism. The dispatch phase, located after the rename phase and before the issue phase, is typically used to write the micro-operations to be executed into the corresponding issue queue. In related technologies, processors usually use fixed priority, round-robin selection, or micro-operation type-based allocation methods to determine the issue queue assigned to a micro-operation during micro-operation dispatch. These methods are relatively simple to implement, but they are prone to situations where some issue queues are congested for extended periods while others are idle, thus reducing overall parallelism. Summary of the Invention

[0003] This application provides a processor instruction dispatching method and apparatus that is compatible with instruction dispatching schemes that take into account issue queue capacity status and register dependencies. Under the premise of ensuring micro-operation age order and type constraints, it achieves a more balanced queue utilization and lower dependency-related latency.

[0004] In a first aspect, embodiments of this application provide a processor instruction dispatch method, the method comprising: Retrieve multiple micro-operations to be dispatched; The candidate launch queue set corresponding to each micro-operation is determined based on the micro-operation type and age order of the multiple micro-operations. Obtain the overall set of launch queues available for dispatch in the processor and the corresponding queue status information, wherein the queue status information includes the remaining capacity information and effective occupancy information of each launch queue in the overall set of launch queues; Based on the queue status information, determine the capacity-limited launch queues in the overall launch queue set; Based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, the multiple micro-operations are initially queued to obtain the initial launch queue corresponding to each micro-operation. Based on the register read / write information of the multiple micro-operations and the initial launch queue corresponding to each micro-operation, the association between register dependencies and launch queues is verified to determine the dependent launch queue corresponding to each micro-operation. After the verification is passed, the multiple micro-operations are dispatched to the corresponding launch queues according to the dependent launch queues of each micro-operation.

[0005] Secondly, embodiments of this application provide a processor instruction dispatching apparatus having functions corresponding to the processor instruction dispatching method provided in the first aspect above. These functions can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. In one embodiment, the processor instruction dispatching apparatus includes: The input / output module is configured to acquire multiple micro-operations to be dispatched; acquire the overall set of launch queues available for dispatch in the processor and the corresponding queue status information, wherein the queue status information includes the remaining capacity information and effective occupancy information of each launch queue in the overall set of launch queues; The processing module is configured to determine the candidate launch queue set corresponding to each micro-operation based on the micro-operation type and age order of the plurality of micro-operations; The processing module is further configured to determine the capacity-limited launch queue in the overall launch queue set based on the queue status information; The processing module is further configured to perform initial queue allocation on the multiple micro-operations based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, so as to obtain the initial launch queue corresponding to each micro-operation. The processing module is further configured to verify the association between register dependencies and launch queues based on the register read / write information of the multiple micro-operations and the initial launch queues corresponding to each micro-operation, so as to determine the dependent launch queues corresponding to each micro-operation. The processing module is further configured to, after the verification is passed, dispatch the multiple micro-operations to the corresponding launch queues through the input / output module according to the dependent launch queues corresponding to each micro-operation.

[0006] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the processor instruction dispatching method as described in the first aspect.

[0007] Fourthly, embodiments of this application provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the processor instruction dispatch method described in the first aspect.

[0008] Fifthly, embodiments of this application provide a chip, the chip including a processor, the processor being used to execute the processor instruction dispatch method provided in the first aspect of embodiments of this application. In one possible design, the chip includes dispatch logic circuitry, an issue queue, a capacity state maintenance circuitry, and a register dependency processing circuitry for implementing the processor instruction dispatch method.

[0009] In a sixth aspect, embodiments of this application provide a chip system including a processor for implementing the functions involved in the first aspect above, such as generating or processing information involved in the processor instruction dispatch method provided in the first aspect above.

[0010] In one possible design, the aforementioned chip system further includes a memory connected to the processor via a circuit structure. This memory stores program instructions and data necessary for the terminal. The chip system can be composed of a single chip or may include chips and other discrete devices. Further optionally, the chip also includes a communication interface to which the processor connects. The communication interface receives data and / or information that needs to be processed. The processor obtains the data and / or information from the communication interface, processes the data and / or information, and outputs the processing result through the communication interface. This communication interface can be an input / output interface.

[0011] Compared to existing technologies, this application embodiment obtains multiple micro-operations to be dispatched and determines the candidate launch queue set corresponding to each micro-operation by combining the micro-operation type and age order. This allows different types of micro-operations to participate in dispatch selection within the candidate range that meets type constraints. By obtaining the remaining capacity information and effective occupancy information of each launch queue in the overall launch queue set and determining the capacity-constrained launch queues, the initial queue allocation process can avoid or reduce the selection of launch queues with high loads, thereby reducing the probability of some launch queues being continuously congested and other launch queues being underutilized. Furthermore, this application embodiment verifies the association between register dependencies and launch queues based on the register read / write information of multiple micro-operations and the initial launch queues corresponding to each micro-operation, and determines the dependent launch queues corresponding to each micro-operation. This allows the dispatch process to perceive the register dependencies between producer micro-operations and consumer micro-operations. After the verification is passed, multiple micro-operations are dispatched to the corresponding launch queues according to the dependent launch queues corresponding to each micro-operation, which helps reduce the path overhead of register-related micro-operations in subsequent wake-up, bypassing, and queue scheduling processes. The embodiments of this application can simultaneously consider micro-operation type constraints, micro-operation age order, launch queue capacity status, and register dependencies during the dispatch phase, thereby improving the load balancing among multiple launch queues, reducing the risk of congestion in hot launch queues, and reducing the dependency-related latency of register-related micro-operations, thus improving the overall efficiency of the processor dispatch phase and subsequent launch phases. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating a processor instruction dispatching method in an embodiment of this application. Figure 2 This is another flowchart illustrating the processor instruction dispatch method according to an embodiment of this application; Figure 3 This is a schematic diagram illustrating the principle of the capacity-aware initial selection method according to an embodiment of this application; Figure 4 This is a schematic diagram illustrating the principle of the final queuing channel selection and blocking method in an embodiment of this application; Figure 5 This is a schematic diagram illustrating another principle of the final queuing channel selection and blocking method in an embodiment of this application; Figure 6 This is a schematic diagram of the processor instruction dispatching device according to an embodiment of this application; Figure 7 This is a schematic diagram of a server structure in one embodiment of this application. Detailed Implementation

[0013] This application provides a processor instruction dispatch method and apparatus, applicable to instruction processing flows under out-of-order execution processors. The solution provided in this application involves technologies such as processor architecture, out-of-order execution, instruction dispatch, issue queue scheduling, and register dependency handling, which are specifically illustrated through the following embodiments.

[0014] The processor architecture consists of the processor's instruction set, pipeline, execution resources, register file, cache structure, and control logic, as well as its hardware and software interfaces and microarchitectural implementation methods. For high-performance processors, out-of-order execution techniques are typically used to improve instruction-level parallelism, allowing multiple micro-operations to be executed without strictly adhering to the original program order, while still satisfying data dependencies, resource availability, and program semantic constraints.

[0015] Out-of-order execution processors typically include pipelined stages such as fetch, decode, rename, dispatch, issue, execution, and commit. After fetching and decoding, the front-end can convert macro instructions into one or more micro-operations; the rename stage can allocate physical registers for micro-operations or maintain the mapping relationship between logical registers and physical registers; the dispatch stage can write the micro-operations to be executed into the issue queue; the issue stage can select micro-operations that meet the issue conditions from the issue queue and send them to the execution unit; the commit stage submits the execution results in program order.

[0016] The issue queue is a crucial scheduling structure in out-of-order execution processors, used to store micro-operations that have been dispatched but not yet issued for execution. Depending on the execution resource type, execution latency, or functional unit configuration, a processor can have multiple issue queues, such as single-cycle issue queues and multi-cycle issue queues. Different types of micro-operations can have different issue queue type constraints. For example, branch-type micro-operations can enter a single-cycle issue queue, multi-cycle micro-operations can enter a multi-cycle issue queue, and single-cycle micro-operations can enter a single-cycle issue queue and / or a multi-cycle issue queue as needed.

[0017] In related technologies, processors typically employ fixed priority, round-robin selection, or allocation methods based solely on micro-operation type to determine the launch queue to which a micro-operation enters during micro-operation dispatch. While these methods are relatively simple to implement, the effective occupancy status, remaining capacity, and available enqueue channels of each launch queue dynamically change during processor operation. If the dispatch selection fails to adapt to the real-time status of the launch queues, some launch queues may experience continuous backlog, while other launch queues may still have idle resources, thereby reducing launch queue utilization and the overall parallelism of the processor.

[0018] Furthermore, register dependencies may exist between micro-operations. For example, the source logic register of a consumer micro-operation depends on the destination logic register of a producer micro-operation. If the dispatch phase does not incorporate register read / write relationships, producer and consumer micro-operations may be assigned to positions unfavorable for subsequent wake-up, bypassing, or internal queue scheduling, thereby increasing the waiting time and control path overhead of dependent micro-operations. Especially in processors with multiple launch queues, simply relying on fixed rules for queue allocation makes it difficult to simultaneously consider type constraints, age order, queue capacity status, and register dependencies.

[0019] Compared to existing technologies, this application embodiment acquires multiple micro-operations to be dispatched and determines a candidate launch queue set corresponding to each micro-operation based on its micro-operation type and age order, enabling micro-operations to participate in dispatch selection while satisfying type constraints and age order. By acquiring the remaining capacity and effective occupancy information of each launch queue in the overall launch queue set and determining the capacity-limited launch queue based on queue status information, the initial queue allocation process can perceive the capacity status of the launch queues, reducing the probability of an overfilled launch queue continuing to receive a large number of micro-operations, thereby improving load balancing among launch queues. Furthermore, this application embodiment verifies the association between register dependencies and launch queues based on the register read / write information of multiple micro-operations and the initial launch queues corresponding to each micro-operation, to determine the dependent launch queues corresponding to each micro-operation. The dispatch phase can identify the register dependencies between producer micro-operations and consumer micro-operations, and dispatches them according to the dependent launch queues corresponding to each micro-operation after the verification is passed, making the queue allocation results of related micro-operations more conducive to subsequent dependency wake-up, bypassing, and internal queue scheduling.

[0020] In some embodiments, the processor instruction dispatching method provided in this application can be implemented based on a processor instruction dispatching device in a processor core. This device includes an input / output module and a processing module. These modules can be integrated and deployed within the same processor core, or they can be distributed according to the processor microarchitecture in the form of pipelines, logic modules, or hardware circuits. In an optional embodiment, the aforementioned functional modules are located after the renaming phase and before the issue phase, and are used to receive multiple micro-operations to be dispatched and dispatch the micro-operations to the corresponding issue queues. In another embodiment, the aforementioned device can also be divided into the following functional modules: The launch queue is used to buffer micro-operations to be launched, and launches the micro-operation to the corresponding execution device when the source operand of the micro-operation is ready and the execution resources are available.

[0021] The dispatch buffer can cache multiple micro-operations to be dispatched in age order. The type recognition module can identify the micro-operation type of each micro-operation and determine its candidate launch queue set based on the micro-operation type. The capacity-aware initial selection module can obtain the remaining capacity information and effective occupancy information of each launch queue, determine the capacity-constrained launch queue, and generate an initial launch queue based on the candidate launch queue set and the capacity-constrained launch queue. The register dependency processing module can determine the association between register dependencies and launch queues based on the destination logic register, source logic register, and initial launch queue of the micro-operation. The queue allocation adjustment module can adjust the initial queue allocation result based on the dependent launch queue. The final enqueue module can dispatch the micro-operation to the corresponding launch queue based on the adjusted queue allocation result and the available enqueue resources of each launch queue.

[0022] In some implementations, the capacity-aware initial selection module can generate a capacity-aware mask based on the effective occupancy information of the transmit queues. This capacity-aware mask is used to identify capacity-constrained transmit queues. When determining the initial transmit queues, transmit queues not matched by the capacity-aware mask can be preferentially selected; if a transmit queue not matched by the capacity-aware mask does not meet the dispatch requirements, a fallback selectable transmit queue matched by the capacity-aware mask can be made to improve queue load balancing while avoiding unnecessary reduction in dispatch bandwidth.

[0023] In some implementations, the register dependency processing module can generate register dependency mapping information based on the correspondence between the destination logic register of a micro-operation and the initial issue queue, and query this register dependency mapping information based on the source logic register of subsequent micro-operations. When the query is successful, the issue queue indicated by the successful register dependency mapping information can be determined as the dependent issue queue of the corresponding micro-operation. Furthermore, this register dependency mapping information can be one-time mapping information, cleared after being hit by a consumer micro-operation, thereby prioritizing the mapping for earlier consumers of producer micro-operations and reducing the impact of historical dependencies on the dispatch selection of subsequent irrelevant micro-operations.

[0024] Furthermore, within the same dispatch cycle, the register dependency processing module can also receive destination register bypass information for older micro-operations within the same dispatch cycle. For younger consumer micro-operations, both the one-time mapping information and the destination register bypass information can be queried simultaneously based on their source registers. When both are matched, the destination register bypass information has higher priority than the historical mapping information, enabling the producer-consumer relationship within the same dispatch cycle to be identified first.

[0025] In the aforementioned co-occurrence dependency scenario, the issue queue allocation result obtained by the producer micro-operation during the capacity-aware initial selection remains fixed and does not participate in register dependency-aware swapping; the consumer micro-operation, on the other hand, uses the issue queue allocation result of the producer micro-operation as the dependent issue queue to participate in subsequent swapping. This avoids instability in the dependent queue caused by mutual swapping between the producer and consumer within the same dispatch cycle and reduces the control complexity of the swapping network.

[0026] Because producer and consumer micro-operations with register dependencies have a data priority relationship, consumer micro-operations typically need to wait for the producer micro-operation to produce a result before they can be launched. Therefore, they cannot be launched simultaneously in the same timeframe. Under the constraints of type and capacity, prioritizing the dispatch of the first consumer micro-operation of a producer micro-operation to the launch queue containing the producer micro-operation does not reduce the chance of simultaneous launches of independent micro-operations. On the contrary, it helps to shorten dependent wake-up paths, result bypass paths, and internal queue scheduling paths.

[0027] When multiple consumer micro-operations that depend on the same producer micro-operation exist within the same dispatch window, the register dependency processing module can prioritize matching the earliest consumer micro-operation with the dispatch queue where the producer micro-operation is located, according to age order. The remaining consumer micro-operations can be distributed to a limited extent through a fixed switching network, provided that age order, type constraints, and capacity constraints are met, in order to avoid the formation of new dependency accumulation in a single dispatch queue.

[0028] In some implementations, the queue allocation adjustment module can adjust the initial queue allocation results through a fixed switching network. Specifically, multiple micro-operations can be divided into at least one adjustment group according to age order, and within each adjustment group, the launch queue allocation results of at least two micro-operations can be adjusted according to a preset priority. This adjustment applies to the launch queue allocation results of the micro-operations without changing the age order of the micro-operations in the dispatch buffer, thus improving the queue allocation results of register-dependent micro-operations while maintaining age order constraints.

[0029] In some implementations, the final enqueue module can perform enqueue resource checks on the target launch queue according to the age order of multiple micro-operations. When the target launch queue has sufficient available enqueue resources, the micro-operation can be written into the corresponding launch queue. When the target launch queue lacks sufficient available enqueue resources, the micro-operation and the micro-operations following it in age order can be blocked to maintain the age order constraint during the dispatch process.

[0030] It should be noted that the processor involved in the embodiments of this application may be a central processing unit, a graphics processing unit, a digital signal processor, an artificial intelligence processor, a network processor, a microcontroller, a processor core in a system-on-a-chip, or other processing units with instruction dispatch and issue queue structures. The processor may be integrated into a chip, a chip system, a computing device, a server, a terminal device, an embedded device, an automotive device, or an edge computing device. The dispatching device in the embodiments of this application may be implemented through hardware logic circuits, or through configurable logic, microcode control logic, or a combination of hardware and software; the embodiments of this application do not limit this approach.

[0031] It should be noted that the chip involved in the embodiments of this application may include one or more processor cores, an issue queue, a register file, an execution unit, a cache, and dispatch logic circuitry for implementing the above-described processor instruction dispatch method. The chip may be a general-purpose processor chip, a dedicated processor chip, or a system-on-a-chip including a processor core. By implementing the above-described dispatch logic in the chip, capacity-aware and register-dependency-aware dispatch decisions can be completed within the processor pipeline, thereby improving processor execution efficiency.

[0032] Figure 1 This is a flowchart illustrating a processor instruction dispatch method provided in an embodiment of this application. The method can be executed by a processor instruction dispatching device and can be applied to an out-of-order execution processor with multiple issue queues. It is used to allocate multiple micro-operations to be dispatched to corresponding issue queues after the renaming phase and before the issue phase. For example, the processor may include multiple single-cycle issue queues and / or multiple multi-cycle issue queues. The processor instruction dispatching device can determine the target issue queue corresponding to each micro-operation based on the micro-operation type, micro-operation age order, issue queue capacity status, and register dependencies between micro-operations, thereby improving issue queue utilization and reducing dependency-related latency. (Refer to...) Figure 1 The method includes steps 101-107: Step 101: Obtain multiple micro-operations to be dispatched.

[0033] In this embodiment, a micro-operation refers to a basic scheduling unit flowing through the processor pipeline after the processor front-end decodes, splits, or merges macro instructions. Exemplarily, micro-operations can be processed in pipeline stages such as renaming, dispatching, issuing, and executing. One architecture instruction can correspond to one micro-operation, or it can be split into multiple micro-operations based on instruction complexity, processor decoding method, or microarchitecture implementation. In some cases, multiple simple operations can also be merged into one micro-operation.

[0034] Multiple micro-operations to be dispatched can be micro-operations output by the dispatch buffer in age order. The age order represents the sequential relationship of the multiple micro-operations in program order, decoding order, or the order in which they enter the dispatch buffer. Older micro-operations typically participate in the dispatch decision before younger micro-operations to maintain the order constraint in the processor dispatch process. Without ambiguity, the micro-operations in this embodiment can also be referred to as microoperations or mops.

[0035] In one possible implementation, the processor can set up a dispatch buffer to cache micro-operations that have undergone preceding pipelined stages such as decoding and renaming. Multiple micro-operations in the dispatch buffer can be arranged in age order, which can correspond to the order in which the micro-operations entered the dispatch buffer, or to the program order or decoded output order of the macro instructions to which the micro-operations belong.

[0036] Within each dispatch cycle, the dispatching device can read multiple micro-operations located at the head or first few positions in the queue from the dispatch buffer, as the multiple micro-operations to be dispatched in the current week. For example, the dispatching device can pop up to several micro-operations from the dispatch buffer each cycle to enter the subsequent dispatch process. The read micro-operations can be arranged in ascending order, for example, sequentially labeled as the first micro-operation, the second micro-operation, the third micro-operation, and the fourth micro-operation. The first micro-operation is older than the second micro-operation, the second micro-operation is older than the third micro-operation, and the third micro-operation is older than the fourth micro-operation.

[0037] In some embodiments, the multiple micro-operations acquired in the current cycle may include micro-operations of different types. For example, the first micro-operation may be a branch-type micro-operation, the second micro-operation may be a single-cycle micro-operation, the third micro-operation may be a multi-cycle micro-operation, and the fourth micro-operation may be a single-cycle micro-operation. After acquiring the above multiple micro-operations, the dispatching device may continue to perform type identification, candidate launch queue determination, capacity-aware initial queue allocation, and register dependency-aware adjustment on each micro-operation.

[0038] In other embodiments, if the number of available micro-operations in the dispatch buffer is less than the processor's maximum dispatch quantity per cycle, the dispatching device can acquire one or more existing micro-operations in the dispatch buffer as micro-operations to be dispatched. If a micro-operation fails to be dispatched successfully during subsequent dispatching due to insufficient remaining capacity in the target launch queue or insufficient available enqueue channels, that micro-operation, along with the younger micro-operations that follow it in age order, can remain in the dispatch buffer and participate in the dispatching decision again as micro-operations to be dispatched in the next dispatching cycle.

[0039] In the above manner, the multiple micro-operations to be dispatched obtained in step 101 can not only reflect the program order output by the processor front end, but also provide an input basis for subsequent dispatch selection based on micro-operation type, issue queue capacity status and register dependency.

[0040] Figure 2 This is another flowchart illustrating the processor instruction dispatch method. (Example) Figure 2As shown, the processor instruction dispatching device may include or be connected to a dispatch buffer, a type identification module, a capacity-aware initial selection module, a register dependency-aware exchange module, a final enqueue module, and multiple launch queues. The dispatch buffer outputs micro-operations to be dispatched in age order; the type identification module identifies the type of the micro-operation to be dispatched, such as branch-type micro-operations, multi-cycle micro-operations, and single-cycle micro-operations; the capacity-aware initial selection module selects an initial launch queue for the micro-operation to be dispatched based on the micro-operation type, capacity-aware mask, and rotation start point; the register dependency-aware exchange module adjusts the initial launch queue selection result based on register dependencies; and the final enqueue module writes the micro-operation into the corresponding launch queue based on the adjusted queue allocation result and the available enqueue channels of each launch queue.

[0041] In this embodiment, the dispatch buffer can output multiple micro-operations to be dispatched in age order. The type recognition module can classify each micro-operation to be dispatched and provide the classification result as a type constraint to the capacity-aware initial selection module. The capacity-aware initial selection module can receive the number of valid items, remaining capacity, and queuing channel status of multiple launch queues, and determine the initial launch queue corresponding to each micro-operation based on the capacity-aware mask and the round-robin pointer. For example, when a launch queue is marked as a capacity-limited launch queue by the capacity-aware mask, the capacity-aware initial selection module can prioritize skipping the capacity-limited launch queue if other optional launch queues exist, so as to reduce the number of new micro-operations received by a relatively full queue.

[0042] Furthermore, the register dependency-aware switching module can make limited adjustments to the initial launch queue allocation results obtained by the capacity-aware initial selection module based on the dependency mapping information from the destination logic register to the launch queue. This ensures that consumer micro-operations with register dependencies are preferentially allocated to launch queues that match their dependent launch queues. The final enqueue module can then write the adjusted micro-operations into the corresponding launch queues based on the availability of the low-order and high-order enqueue channels for each target launch queue. For micro-operations that fail to be dispatched, they can be retained in the dispatch buffer and continue to participate in dispatch judgment in subsequent dispatch cycles.

[0043] Figure 2 The diagram also shows single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0 to represent an exemplary overall launch queue set. Gray slots in each queue represent occupied valid items, while empty slots represent currently available space. When a launch queue is marked with a mask, it indicates that the launch queue was identified as a capacity-constrained launch queue during the capacity-aware phase, and the subsequent initial queue allocation process can preferentially avoid this queue. It should be noted that... Figure 2The number of launch queues, launch queue types, and module divisions shown are merely examples and do not constitute a limitation on the scope of protection of the embodiments of this application. In an optional embodiment, the processor may include different numbers of single-cycle launch queues, multi-cycle launch queues, or other launch queues with different types of constraints, capacity states, and queuing channel configurations.

[0044] Step 102: Determine the candidate launch queue set corresponding to each micro-operation based on the micro-operation type and age order of the multiple micro-operations.

[0045] In this embodiment, the micro-operation type can be used to characterize the micro-operation's requirements for execution resources, execution latency, or launch queue type. Different types of micro-operations can correspond to different launch queue type constraints. For example, the micro-operation type can include branch-type micro-operations, multi-cycle micro-operations, and single-cycle micro-operations. Branch-type micro-operations can be micro-operations related to branch jumps, conditional judgments, or control flow; multi-cycle micro-operations can be micro-operations that require processing by multi-cycle execution resources; and single-cycle micro-operations can be micro-operations that can be processed by single-cycle execution resources.

[0046] In this embodiment, the age order of micro-operations can be used to characterize the sequential relationship between multiple micro-operations. This sequential relationship can be determined based on the program order of the macro instructions to which the micro-operation belongs, the decoding output order, the order in which they enter the dispatch buffer, or their arrangement order within the dispatch buffer. Older micro-operations are typically located earlier in the dispatch window, while younger micro-operations are typically located later in the dispatch window. By combining the age order to identify the micro-operation type, the subsequent candidate launch queue determination and initial queue allocation process can be kept consistent with the dispatch order, preventing younger micro-operations from inappropriately skipping older micro-operations to complete dispatch.

[0047] In this embodiment, the candidate launch queue set corresponding to a micro-operation refers to a set of launch queues within the overall launch queue set available for dispatch in the processor that satisfy the type constraints, execution resource constraints, or microarchitectural constraints of the micro-operation. In other words, the candidate launch queue set is the range of selectable queues after filtering the overall launch queue set for a specific micro-operation. Subsequent capacity-aware selection, round-robin selection, register-dependent adjustment, or final enqueue check can be performed within the range defined by this candidate launch queue set.

[0048] As an optional embodiment, the overall launch queue set includes single-cycle launch queues and multi-cycle launch queues. Based on this, step 102, determining the candidate launch queue set corresponding to each micro-operation according to the micro-operation type and age order of the plurality of micro-operations, includes: identifying the micro-operation type of each micro-operation according to the age order of the plurality of micro-operations; and determining the range of launch queues usable by each micro-operation from the overall launch queue set based on the launch queue type constraints corresponding to the micro-operation type. Specifically, branch-type micro-operations among the plurality of micro-operations can use the single-cycle launch queue, multi-cycle micro-operations among the plurality of micro-operations can use the multi-cycle launch queue, and single-cycle micro-operations among the plurality of micro-operations can use both the single-cycle launch queue and / or the multi-cycle launch queue.

[0049] For example, suppose the processor includes three single-cycle launch queues and one multi-cycle launch queue. The three single-cycle launch queues are denoted as Single-cycle Launch Queue 0, Single-cycle Launch Queue 1, and Single-cycle Launch Queue 2, respectively, and the multi-cycle launch queue is denoted as Multi-cycle Launch Queue 0. These four launch queues constitute the overall launch queue set. It should be noted that the above combination of queue number and type is merely an example. In other embodiments, the overall launch queue set may also include other numbers of single-cycle launch queues, multi-cycle launch queues, or launch queues with other types of constraints.

[0050] In the above embodiments, within the current dispatch cycle, the dispatch buffer can output multiple micro-operations to be dispatched in age order. For example, the multiple micro-operations can be sequentially designated as the first micro-operation, the second micro-operation, the third micro-operation, and the fourth micro-operation, wherein the age of the first micro-operation is earlier than that of the second micro-operation, the age of the second micro-operation is earlier than that of the third micro-operation, and the age of the third micro-operation is earlier than that of the fourth micro-operation. The dispatch device can sequentially identify the micro-operation type of each micro-operation in the above age order and determine the corresponding candidate launch queue set based on the identification results.

[0051] For example, when the first micro-operation is a branching micro-operation, since branching micro-operations can only use single-cycle launch queues, single-cycle launch queue 0, single-cycle launch queue 1, and single-cycle launch queue 2 can be determined as the candidate launch queue set corresponding to the first micro-operation. That is to say, the first micro-operation can only select its initial launch queue from the above-mentioned single-cycle launch queues.

[0052] When the second micro-operation is a multi-cycle micro-operation, since multi-cycle micro-operations can only use multi-cycle launch queues, multi-cycle launch queue 0 can be determined as the candidate launch queue set corresponding to the second micro-operation. In this case, the candidate range for the second micro-operation is relatively narrow, and its subsequent initial queue allocation usually points to multi-cycle launch queue 0, or is selected from multiple multi-cycle launch queues when multiple multi-cycle launch queues exist.

[0053] When the third micro-operation is a single-cycle micro-operation, since single-cycle micro-operations can be processed by single-cycle execution resources, and in some processor implementations can also enter a multi-cycle launch queue based on resource utilization, single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0 can be determined as the candidate launch queue set corresponding to the third micro-operation. In other words, the candidate launch queue set for a single-cycle micro-operation can simultaneously include both single-cycle and multi-cycle launch queues, allowing for more flexible dispatch selection based on queue capacity status.

[0054] When the fourth micro-operation is also a single-cycle micro-operation, its candidate launch queue set can be determined in the same way. If the fourth micro-operation is a branch-type micro-operation or a multi-cycle micro-operation, its candidate launch queue set can be determined according to the launch queue type constraints corresponding to the branch-type micro-operation or the multi-cycle micro-operation, respectively.

[0055] Using the above method, step 102 can first determine the processing order of each micro-operation in the current dispatch window based on age order, and then determine the candidate dispatch queue set from the overall dispatch queue set based on the type constraints of each micro-operation. Therefore, subsequent steps can further combine capacity-constrained dispatch queues, round-robin start points, register dependencies, and available enqueue resources within the candidate dispatch queue set to determine the dispatch queue for each micro-operation.

[0056] Optionally, the weight of the launch queue type constraint for the branch-type micro-operation and the multi-cycle micro-operation is higher than the weight of the launch queue type constraint for the single-cycle micro-operation. That is, the range of launch queues that can be used by branch-type micro-operations and multi-cycle micro-operations is relatively narrower, and the type constraint is stronger. The range of launch queues that can be used by single-cycle micro-operations is relatively wider, and the type constraint is weaker.

[0057] For example, when the processor includes single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0, branch-type micro-operations can only enter single-cycle launch queue 0, single-cycle launch queue 1, or single-cycle launch queue 2; multi-cycle micro-operations can only enter multi-cycle launch queue 0; while single-cycle micro-operations can enter either the aforementioned single-cycle launch queues or, depending on implementation needs, multi-cycle launch queue 0. During initial queue allocation, initial launch queues can be prioritized for branch-type and multi-cycle micro-operations with stronger type constraints. Then, based on the remaining launch queue resources, initial launch queues can be determined for single-cycle micro-operations with weaker type constraints. This reduces the probability that micro-operations with stronger type constraints cannot be dispatched due to insufficient available queues.

[0058] Based on this, in an optional implementation of step 105, initial queue allocation is performed on the plurality of micro-operations based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, including: determining the branch-type micro-operation and the multi-cycle micro-operation from the plurality of micro-operations according to the age order of the plurality of micro-operations; for the multi-cycle micro-operation, determining the corresponding multi-cycle launch queue based on the candidate launch queue set corresponding to the multi-cycle micro-operation, and using the multi-cycle launch queue as the initial launch queue of the multi-cycle micro-operation; for the branch-type micro-operation, determining the initial launch queue of the branch-type micro-operation from the single-cycle launch queues available for the branch-type micro-operation based on the candidate launch queue set corresponding to the branch-type micro-operation, the capacity-limited launch queue, and the rotation order; after the initial queue allocation of the branch-type micro-operation and the multi-cycle micro-operation is completed, initial queue allocation is performed on the single-cycle micro-operation based on the remaining launch queue resources occupied by the allocated micro-operations.

[0059] For example, within the current dispatch cycle, the dispatch buffer outputs multiple micro-operations in age order, such as the first micro-operation, the second micro-operation, the third micro-operation, and the fourth micro-operation in sequence. The dispatch device can first identify branch-type micro-operations and multi-cycle micro-operations in age order. For example, if the first micro-operation is a branch-type micro-operation and the third micro-operation is a multi-cycle micro-operation, the dispatch device can first process the initial queue allocation of the first and third micro-operations.

[0060] For a multi-cycle micro-operation like the third micro-operation, since its candidate launch queue set only includes multi-cycle launch queue 0, the dispatching device can determine multi-cycle launch queue 0 as the initial launch queue for the multi-cycle micro-operation. In some embodiments, if multi-cycle launch queue 0 has multiple queuing channels, a lower-order queuing channel can be preferentially allocated to the multi-cycle micro-operation for subsequent queuing resource management.

[0061] For branching micro-operations like the first micro-operation, since its candidate launch queue set includes single-cycle launch queue 0, single-cycle launch queue 1, and single-cycle launch queue 2, the dispatching device can determine its initial launch queue from the aforementioned single-cycle launch queues by combining the capacity-limited launch queues and the rotation order. For example, when single-cycle launch queue 0 is determined to be a capacity-limited launch queue, the dispatching device can preferentially skip single-cycle launch queue 0 during the rotation scan and select the initial launch queue for the first micro-operation from single-cycle launch queue 1 or a subsequent single-cycle launch queue not marked as capacity-limited.

[0062] After the initial queue allocation for branch-type micro-operations and multi-cycle micro-operations is completed, the dispatching device can statistically analyze or update the occupancy status of the allocated micro-operations for each launch queue's enqueue resources, and perform initial queue allocation for single-cycle micro-operations based on the remaining launch queue resources. For example, if the second and fourth micro-operations are single-cycle micro-operations, the dispatching device can prioritize selecting initial launch queues for the second and fourth micro-operations from single-cycle launch queues with ample remaining capacity that are not identified as capacity-limited, or from available multi-cycle launch queues, based on the queue resources already occupied by the first and third micro-operations. This two-tiered selection method, which processes micro-operations with strong type constraints first and then those with weak type constraints, can improve the overall dispatch success rate of the current dispatch cycle while ensuring type constraints are met.

[0063] Further optionally, if the remaining capacity of the multi-cycle launch queue is insufficient to meet the dispatch requirements of the multi-cycle micro-operations, the multi-cycle micro-operations and the micro-operations that follow the multi-cycle micro-operations in age order are blocked.

[0064] For example, during the current dispatch cycle, the dispatch buffer outputs the first, second, third, and fourth micro-operations in age order. The third micro-operation is a multi-cycle micro-operation, and its candidate launch queue set only includes multi-cycle launch queue 0. If the dispatch device determines during initial queue allocation or final enqueue check that the remaining capacity of multi-cycle launch queue 0 is insufficient, or that multi-cycle launch queue 0 currently has no available enqueue channel, then the third micro-operation cannot be dispatched to multi-cycle launch queue 0 in the current cycle.

[0065] In this scenario, to maintain the age order of micro-operations, the dispatching mechanism can block the third micro-operation and also block the fourth micro-operation that follows the third micro-operation in age order. Even if the fourth micro-operation is a single-cycle micro-operation and an available single-cycle launch queue exists in its candidate launch queue set, the fourth micro-operation will not skip the third micro-operation to complete dispatch. Instead, it will remain in the dispatch buffer along with the third micro-operation, waiting to participate again in the capacity-aware initial selection, register dependency-aware adjustment, and final enqueue check in the next dispatch cycle.

[0066] If the first and second micro-operations are older than the third micro-operation in age order, and their target launch queues have available enqueue resources, then the first and second micro-operations can still be dispatched in the current cycle. This prevents younger micro-operations from overtaking older micro-operations when launch queue capacity is insufficient, thus maintaining the age order constraint during the dispatch phase and reducing the complexity of subsequent scheduling state maintenance.

[0067] Step 103: Obtain the overall set of launch queues available for dispatch in the processor and the corresponding queue status information. The queue status information includes the remaining capacity information and effective occupancy information of each launch queue in the overall launch queue set.

[0068] In this embodiment, the overall launch queue set can refer to the collection of multiple launch queues in the processor currently available for receiving micro-operations to be dispatched. Exemplarily, the overall launch queue set may include one or more single-cycle launch queues and one or more multi-cycle launch queues. Single-cycle launch queues can be used to receive micro-operations suitable for processing by single-cycle execution resources, while multi-cycle launch queues can be used to receive micro-operations suitable for processing by multi-cycle execution resources. In some implementations, some micro-operations may also be selected from multiple different types of launch queues based on type constraints and processor implementation.

[0069] In this embodiment, queue status information refers to information characterizing the current operating status of each launch queue in the overall launch queue set. Queue status information may include, but is not limited to, remaining capacity information, effective occupancy information, and available enqueue channels for each launch queue. The processor instruction dispatching device can determine, based on the queue status information, whether a launch queue is nearly full, whether it still has available space, and whether it can receive new micro-operations in the current dispatch cycle, thereby providing a basis for subsequent capacity-aware initial queue allocation.

[0070] The remaining capacity information can be used to characterize the amount of space in the transmit queue that can still receive micro-operations. For example, if a transmit queue has multiple queue slots, and some of these slots already hold valid micro-operations, the number of remaining unoccupied queue slots can be considered the remaining capacity of the transmit queue. The remaining capacity information can also be expressed as the number of remaining free items, the number of free slots, or a count of available space. For instance, when the remaining capacity of a transmit queue is 2, it means that the transmit queue can currently receive at least two more micro-operations. When the remaining capacity is 0, it means that the transmit queue cannot currently receive any new micro-operations.

[0071] Effective occupancy information can be used to characterize the number of valid micro-operations currently stored in a launch queue. For example, if a launch queue has multiple valid entries, the effective occupancy count can indicate the current congestion or load level of the queue. A higher effective occupancy count generally indicates that the launch queue is closer to full capacity. A lower effective occupancy count generally indicates that the launch queue still has more available space. The processor instruction dispatcher can identify launch queues that are more congested than others by comparing the effective occupancy counts among different launch queues.

[0072] Specifically, in step 103, within each dispatch cycle or preset detection cycle, the capacity counter, valid item counter, and queuing channel status information of each launch queue in the overall launch queue set can be read to obtain the remaining capacity information and valid occupancy information of each launch queue. For example, when the processor includes single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0, the processor instruction dispatch device can read the number of valid items, the number of remaining idle items, and the available status of the queuing channel in the current dispatch cycle for each of the aforementioned launch queues.

[0073] In some embodiments, the remaining capacity information can be determined by the difference between the total number of slots in the transmit queue and the number of active occupancy slots, or it can be directly provided by the idle counter corresponding to the transmit queue. For example, if a transmit queue has a total of 8 slots and currently has 6 active occupancy slots, then the remaining capacity of the transmit queue can be 2. If the remaining capacity of a transmit queue is 1, it means that the transmit queue can receive at most one micro-operation in the current cycle. If the remaining capacity is 0, it means that the transmit queue cannot currently receive any new micro-operations.

[0074] In some embodiments, effective occupancy information can be represented by the number of valid items in the launch queue. Valid items in the launch queue can be micro-operations that have been dispatched into the launch queue but have not yet been launched or removed from the launch queue. The processor instruction dispatching device can determine the load difference between queues based on the effective occupancy count of each launch queue. For example, when the effective occupancy count of single-cycle launch queue 0 is significantly higher than that of single-cycle launch queue 1 or single-cycle launch queue 2, it can be considered that single-cycle launch queue 0 is currently heavily loaded, and further determination of whether it belongs to a capacity-limited launch queue can be made using a preset threshold.

[0075] In some embodiments, available enqueue channel information can be used to characterize whether a launch queue can receive new micro-operations through enqueue channels during the current dispatch cycle. For example, each launch queue can receive up to two micro-operations per clock cycle, written through a low-order enqueue channel and a high-order enqueue channel, respectively. When a low-order enqueue channel is available, it can be used to receive older micro-operations in the same target queue; when a high-order enqueue channel is available, it can be used to receive younger micro-operations in the same target queue. The processor instruction dispatching device can combine remaining capacity information and available enqueue channel information to determine whether a launch queue meets the dispatching requirements for micro-operations in the current cycle.

[0076] In this way, step 103 can obtain the real-time status of each transmission queue before the initial queue allocation, so that the determination of subsequent capacity-limited transmission queues, the generation of capacity-aware masks and the selection of initial transmission queues can all be based on the current queue load and available resources, thereby reducing the probability of overly full transmission queues continuing to receive new micro-operations and improving the resource utilization balance among multiple transmission queues.

[0077] Step 104: Determine the capacity-limited launch queues in the overall launch queue set based on the queue status information.

[0078] As an optional embodiment, in step 104, determining the capacity-limited launch queue in the overall launch queue set based on the queue status information includes: determining the difference between the effective occupancy number of any launch queue and the effective occupancy number of at least one other launch queue based on the effective occupancy information of each launch queue in the overall launch queue set; if the difference is greater than a preset threshold, determining the any launch queue as a capacity-limited launch queue, and generating a capacity-aware mask for identifying the capacity-limited launch queue.

[0079] In this embodiment, a capacity-constrained transmit queue can refer to a transmit queue with a higher effective occupancy, heavier load, or less remaining capacity compared to other candidate transmit queues. A capacity-constrained transmit queue does not necessarily mean that the transmit queue is completely unavailable, but rather that during subsequent initial queue allocation, the transmit queue can be preferentially skipped or have its selection priority reduced to avoid a full queue continuously receiving new micro-operations.

[0080] Capacity-aware masks can be used to identify capacity-limited launch queues. In practical applications, a capacity-aware mask may include flags corresponding to each launch queue in the overall launch queue set. For example, a capacity-aware mask may include multiple flags corresponding one-to-one with each launch queue. If a launch queue is identified as a capacity-limited launch queue, the flags in the capacity-aware mask corresponding to that launch queue can be set to active. If a launch queue is not identified as a capacity-limited launch queue, the flags in the capacity-aware mask corresponding to that launch queue can remain inactive. During subsequent initial launch queue selection, the dispatching device can prioritize skipping launched queues that have been hit by the capacity-aware mask and prioritize selecting launch queues that have not been hit by the capacity-aware mask.

[0081] For example, the processor includes single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0. The dispatching device can read the effective occupancy, remaining capacity, and queuing channel status of each of the above launch queues in each dispatching cycle or a preset cycle. For example, if the effective occupancy of single-cycle launch queue 0 is high, while the effective occupancy of single-cycle launch queue 1 and single-cycle launch queue 2 is low, then single-cycle launch queue 0 can be considered more congested than the other single-cycle launch queues.

[0082] Furthermore, the difference between the effective occupancy of single-cycle transmit queue 0 and the effective occupancy of single-cycle transmit queue 1 can be calculated, or the difference between the effective occupancy of single-cycle transmit queue 0 and the effective occupancy of single-cycle transmit queue 2 can be calculated. If this difference is greater than a preset threshold, it indicates that the load of single-cycle transmit queue 0 is significantly higher than at least one other transmit queue. In this case, single-cycle transmit queue 0 can be identified as a capacity-constrained transmit queue and marked in the capacity-aware mask.

[0083] For example, after single-cycle transmit queue 0 is marked by a capacity-aware mask, during subsequent initial queue allocation, if the candidate transmit queue set for a certain branch-type micro-operation includes single-cycle transmit queue 0, single-cycle transmit queue 1, and single-cycle transmit queue 2, the dispatching device can prioritize skipping single-cycle transmit queue 0 during round-robin scanning and select an initial transmit queue from queues such as single-cycle transmit queue 1 or single-cycle transmit queue 2 that have not been hit by the capacity-aware mask. This reduces the probability that a relatively full single-cycle transmit queue 0 will continue to receive new micro-operations, allowing new micro-operations to be preferentially allocated to relatively empty transmit queues.

[0084] For example, for a single-cycle micro-operation, since its candidate launch queue set can include single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0, the dispatching device can integrate the remaining capacity information and capacity-aware masks of each candidate launch queue, and prioritize the queues with more remaining capacity or available enqueue channels among the queues not hit by the mask. If the launch queues not hit by the mask are insufficient to handle the micro-operations to be dispatched in the current dispatch window, and the launch queues hit by the mask still have available enqueue channels, the launch queues hit by the mask can be selected back according to the implementation strategy to avoid unnecessarily reducing dispatch bandwidth due to excessive avoidance of capacity-limited launch queues.

[0085] In some implementations, for single-cycle micro-operations, high-order enqueue channels can be prioritized from the transmit queues not hit by the capacity-aware mask, followed by low-order enqueue channels. When the transmit queues or enqueue channels not hit by the capacity-aware mask are insufficient to handle single-cycle micro-operations in the current dispatch window, a fallback can be made to select transmit queues hit by the capacity-aware mask or their remaining enqueue channels. The above channel priority is only one optional implementation, which can improve the utilization rate of enqueue channels in the current dispatch cycle while prioritizing the avoidance of congested queues.

[0086] In other implementations, the capacity-aware mask can be controlled by configuration bits. When the processor is configured to disable the capacity-aware mask, the initial capacity-aware selection can degenerate into a direct dispatch method based on the round-robin start point. This allows for support of both capacity-aware dispatch mode and round-robin priority dispatch mode within the same hardware architecture, and facilitates configuration according to different performance, power consumption, or verification requirements.

[0087] In another example, if the remaining capacity of multi-cycle launch queue 0 is 0, then multi-cycle micro-operations may not be dispatched in the current cycle because the candidate launch queue set only includes multi-cycle launch queue 0. In this case, the remaining capacity information of multi-cycle launch queue 0 can be used for subsequent enqueue resource checks and blocking determination. If the current multi-cycle micro-operation cannot enter multi-cycle launch queue 0, then this multi-cycle micro-operation, as well as the younger micro-operations that follow it in age order, can be retained in the dispatch buffer, waiting to participate in the dispatch determination again in the next cycle.

[0088] Using the above method, step 104 can identify heavily loaded transmission queues based on the effective occupancy information and preset thresholds of each transmission queue, and mark them using a capacity-aware mask. Subsequent initial queue allocation can be based on the capacity-aware mask and the round-robin start point, thereby suppressing overly full transmission queues from continuously receiving new micro-operations, reducing the risk of congestion in hotspot queues, and improving the load balancing among multiple transmission queues.

[0089] See Figure 3 , Figure 3 This is a schematic diagram illustrating the principle of a capacity-aware initial selection method provided in an embodiment of this application. Figure 3 As shown, the processor can include multiple launch queues, such as single-cycle queues 0, 1, and 2, and multi-cycle queue 0. Before initial queue allocation, the queue status information of each launch queue is obtained. This queue status information can include the number of valid items, remaining capacity, and available enqueue channels in the current dispatch cycle for each launch queue. The capacity-aware initial selection module can generate a capacity-aware mask based on the relationship between the difference in the number of valid items of each launch queue and a preset threshold to mark launch queues with relatively high current load.

[0090] In this embodiment, if the number of valid items in a certain transmission queue exceeds a preset threshold compared to other candidate transmission queues, then that transmission queue can be identified as a capacity-limited transmission queue and marked as a priority skip queue by a capacity-aware mask. For example... Figure 3 As shown, when the number of valid items in single-cycle queue0 is large, a capacity-aware mask can be applied to single-cycle queue0 through a priority mask. This allows subsequent micro-operations to prioritize avoiding single-cycle queue0 when other alternative launch queues exist during the initial queue selection process. This reduces the number of new micro-operations entering already full launch queues and decreases the load difference between different launch queues.

[0091] Furthermore, the capacity-aware initial selection can be divided into two selection phases based on micro-operation type constraints. The first phase prioritizes micro-operations with strong type constraints, such as branching micro-operations and multi-cycle micro-operations. The second phase then processes single-cycle micro-operations with relatively weak type constraints. By first allocating a narrower candidate launch queue to the micro-operations with stronger constraints, and then allowing the micro-operations with weaker constraints to use the remaining enqueue resources, the probability of preceding micro-operations being unable to be dispatched due to subsequent micro-operations occupying critical queue resources can be reduced.

[0092] For branch-type micro-operations (such as the `br` instruction), the candidate launch queue set can include `single-cyclequeue0`, `single-cycle queue1`, and `single-cycle queue2`. The capacity-aware initial selection module can scan and select from this candidate launch queue set, combining the round-robin start point and the capacity-aware mask. When `single-cycle queue0` is marked by the priority mask, if there are other unmarked single-cycle launch queues, `single-cycle queue0` can be skipped, and selection can begin from `single-cycle queue1` or subsequent candidate launch queues not hit by the capacity-aware mask. If all single-cycle launch queues are hit by the capacity-aware mask, but there are still available enqueue channels, launch queue selection based on round-robin priority can continue according to a preset backoff strategy to avoid excessive restriction on dispatching leading to a decrease in processor dispatch bandwidth.

[0093] For multi-cycle micro-operations (such as multi-cycle instructions), the candidate launch queue set may only include multi-cycle queue0. Therefore, the capacity-aware initial selection module can fix the initial launch queue of the multi-cycle micro-operation as multi-cycle queue0 and prioritize occupying the available enqueue channels of multi-cycle queue0. If multi-cycle queue0 has no available enqueue channels in the current dispatch cycle, or its remaining capacity is insufficient to receive the multi-cycle micro-operation, then the multi-cycle micro-operation cannot be dispatched in the current dispatch cycle, and according to the age order of micro-operations, younger micro-operations following the multi-cycle micro-operation cannot skip the multi-cycle micro-operation to complete dispatch.

[0094] For single-cycle micro-operations (such as single-cycle instructions), the candidate launch queue set can include single-cycle queue0, single-cycle queue1, single-cycle queue2, and multi-cycle queue0, which allows receiving single-cycle micro-operations. Since single-cycle micro-operations perform initial selection after branch-type and multi-cycle micro-operations, queue selection for single-cycle micro-operations can be based on the remaining capacity and remaining enqueue channels after the first-stage allocation. Specifically, the remaining enqueue channels of launch queues not hit by the capacity-aware mask can be prioritized. When multiple unhit launch queues are available, launch queues with more remaining capacity or unoccupied high-order enqueue channels can be prioritized. When the unhit launch queues are insufficient to handle single-cycle micro-operations in the current dispatch window, the remaining enqueue channels of launch queues hit by the capacity-aware mask are then selected.

[0095] Through the above methods Figure 3 The capacity-aware initial selection process shown can coordinate micro-operation type constraints, micro-operation age order, round-robin fairness, and launch queue capacity status. Branch-type micro-operations and multi-cycle micro-operations can preferentially obtain launch queue resources that match their type; single-cycle micro-operations can be distributed among multiple candidate launch queues based on remaining capacity and capacity-aware masks, thereby reducing the continuous aggregation of micro-operations to a single launch queue, improving load balancing among launch queues, and reducing the risk of congestion caused by local queue congestion in subsequent launch phases.

[0096] Step 105: Based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, perform initial queue allocation for the multiple micro-operations to obtain the initial launch queue corresponding to each micro-operation.

[0097] In this embodiment, the candidate launch queue set refers to the range of selectable launch queues determined from the overall launch queue set for a specific micro-operation, based on the micro-operation type, execution resource requirements, execution latency requirements, and processor microarchitecture constraints. The candidate launch queue sets for different micro-operations can be the same or different. For example, the candidate launch queue set for a branch-type micro-operation may include one or more single-cycle launch queues; the candidate launch queue set for a multi-cycle micro-operation may include one or more multi-cycle launch queues; the candidate launch queue set for a single-cycle micro-operation may include one or more single-cycle launch queues, or, depending on implementation needs, one or more multi-cycle launch queues.

[0098] In this embodiment, the initial launch queue refers to the target queue initially determined for the micro-operation based on factors such as the candidate launch queue set corresponding to the micro-operation, the capacity-constrained launch queue, the capacity-aware mask, the round-robin start point, and the available resources of the launch queue, before register dependency-aware adjustment. In other words, the initial launch queue is the queue allocation result obtained by the micro-operation in the capacity-aware initial selection phase. In subsequent steps, the processor instruction dispatching device can further adjust the initial launch queue according to the register dependency relationship to obtain the final target launch queue used for dispatching.

[0099] As an optional embodiment, in step 105, based on the candidate launch queue set corresponding to each micro-operation and the capacity-constrained launch queue, an initial queue allocation is performed on the multiple micro-operations, including: based on the round-robin start point, from the candidate launch queue set corresponding to each micro-operation, priority is given to selecting launch queues that have not been determined as the capacity-constrained launch queues as the initial launch queues for the corresponding micro-operations; if the launch queues that have not been determined as capacity-constrained launch queues do not meet the dispatch requirements, it is allowed to select the initial launch queues for the corresponding micro-operations from the capacity-constrained launch queues, so as to balance launch queue load balancing and dispatch bandwidth.

[0100] For example, the processor includes single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0. If, based on the effective occupancy information of each launch queue, it is determined that the effective occupancy of single-cycle launch queue 0 is significantly higher than that of other candidate launch queues, and the difference is greater than a preset threshold, then single-cycle launch queue 0 can be identified as a capacity-limited launch queue and marked using a capacity-aware mask. In this case, when performing initial queue allocation for micro-operations, the dispatching device can combine the candidate launch queue set corresponding to the micro-operation with the capacity-aware mask to preferentially select launch queues that have not been marked as capacity-limited.

[0101] If the micro-operation to be dispatched is a branching micro-operation, its candidate launch queue set includes single-cycle launch queue 0, single-cycle launch queue 1, and single-cycle launch queue 2. When single-cycle launch queue 0 is marked as a capacity-limited launch queue by a capacity-aware mask, the dispatching device can scan the candidate launch queues from the start of the rotation and preferentially skip single-cycle launch queue 0, selecting an available queue from single-cycle launch queue 1 or single-cycle launch queue 2 as the initial launch queue for the branching micro-operation. This avoids the relatively full single-cycle launch queue 0 from continuing to receive new branching micro-operations when other available single-cycle launch queues exist, thereby alleviating congestion in hotspot queues.

[0102] If the micro-operation to be dispatched is a single-cycle micro-operation, its candidate launch queue set includes single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0. When single-cycle launch queue 0 is marked as a capacity-constrained launch queue, the dispatching device can preferentially select an initial launch queue from single-cycle launch queue 1, single-cycle launch queue 2, or multi-cycle launch queue 0. When multiple queues not hit by the capacity-aware mask are available, the launch queue with more remaining capacity, available enqueue channels, or unoccupied high-order enqueue channels can be preferentially selected. In this way, single-cycle micro-operations can utilize idle single-cycle launch queues or, if the processor implementation allows, the available resources of multi-cycle launch queues.

[0103] In the above embodiments, firstly, based on the rotation starting point, from the candidate launch queue set corresponding to each micro-operation, the launch queue that has not been determined as the capacity-limited launch queue is preferentially selected as the initial launch queue for the corresponding micro-operation.

[0104] Specifically, based on a capacity-aware mask, the launch queues in the candidate launch queue set corresponding to each micro-operation are filtered to obtain a first candidate launch queue that is not hit by the capacity-aware mask, and a second candidate launch queue that is hit by the capacity-aware mask. Then, starting from the rotation start point, the first candidate launch queue is scanned, and the scanned available launch queues are determined as the initial launch queues for the corresponding micro-operations.

[0105] For example, if the overall launch queue set includes single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0, and the capacity-aware mask identifies single-cycle launch queue 0 as a capacity-limited launch queue, then for a branch-type micro-operation whose candidate launch queue set includes single-cycle launch queue 0, single-cycle launch queue 1, and single-cycle launch queue 2, single-cycle launch queue 1 and single-cycle launch queue 2 can be identified as the first candidate launch queues, and single-cycle launch queue 0 as the second candidate launch queue. If the current rotation start point points to single-cycle launch queue 0, the dispatching device can skip single-cycle launch queue 0 that is hit by the capacity-aware mask during scanning and continue scanning single-cycle launch queue 1 and single-cycle launch queue 2. If single-cycle launch queue 1 has an available enqueue channel, then single-cycle launch queue 1 can be identified as the initial launch queue for this branch-type micro-operation. If single-cycle launch queue 1 is unavailable but single-cycle launch queue 2 is available, then single-cycle launch queue 2 can be identified as the initial launch queue for this branch-type micro-operation.

[0106] For a single-cycle micro-operation where the candidate launch queue set includes single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0, if the capacity-aware mask hits single-cycle launch queue 0, then the first candidate launch queue can include single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0, and the second candidate launch queue can include single-cycle launch queue 0. The dispatching device can scan the first candidate launch queues from the rotation start point and preferentially select queues with more remaining capacity or unoccupied available enqueue channels as the initial launch queue. For example, if single-cycle launch queue 2 has more remaining capacity and its high-order enqueue channel is unoccupied, then single-cycle launch queue 2 can be preferentially determined as the initial launch queue for this single-cycle micro-operation.

[0107] Optionally, if a launch queue not identified as a capacity-constrained launch queue does not meet the dispatch requirements, an initial launch queue for the corresponding micro-operation may be selected from the capacity-constrained launch queues. Specifically, if the first candidate launch queue has no available enqueue channel, and the second candidate launch queue has an available enqueue channel, then the initial launch queue for the corresponding micro-operation is determined from the second candidate launch queue based on the rotation start point.

[0108] For example, consider single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0. If single-cycle launch queue 0 is hit by the capacity-aware mask, it becomes the second candidate launch queue; single-cycle launch queue 1 and single-cycle launch queue 2 become the first candidate launch queues. For branch-type micro-operations, they can only enter single-cycle launch queues. If neither single-cycle launch queue 1 nor single-cycle launch queue 2 has an available enqueue channel, but single-cycle launch queue 0, although hit by the capacity-aware mask, still has an available enqueue channel, the dispatching device can backtrack from the rotation start point and select single-cycle launch queue 0 as the initial launch queue for the branch-type micro-operation. This avoids unnecessary reduction in dispatch bandwidth for the current dispatch cycle due to strict avoidance of capacity-limited launch queues.

[0109] For a single-cycle micro-operation, if none of the single-cycle launch queues 1, 2, and 0 in the first candidate launch queue have available enqueue channels, while single-cycle launch queue 0 in the second candidate launch queue still has available enqueue channels, the dispatching device can allow the single-cycle micro-operation to roll back and select single-cycle launch queue 0 as the initial launch queue. This rollback selection does not indicate the cancellation of the capacity-aware strategy, but rather is a supplementary strategy adopted to maintain dispatch bandwidth and reduce dispatch pauses when queues not matched by the mask cannot meet dispatch requirements.

[0110] In this way, step 105 can prioritize avoiding capacity-constrained launch queues within the candidate launch queue set and select a more suitable initial launch queue based on the round-robin starting point. Simultaneously, when non-capacity-constrained queues cannot meet dispatch requirements, it allows for fallback to use capacity-constrained launch queues that still have available enqueue resources. Therefore, a balance can be achieved between queue load balancing and dispatch bandwidth, reducing the probability of overly full queues continuously receiving new micro-operations while avoiding unnecessary blocking of dispatches when available resources exist.

[0111] In another optional embodiment, step 105, based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, performs initial queue allocation for the plurality of micro-operations, further including: performing hierarchical selection of the plurality of micro-operations according to the type constraint strength of the plurality of micro-operations; performing initial queue allocation for micro-operations with higher type constraint strength and performing initial queue allocation for micro-operations with lower type constraint strength.

[0112] For example, a processor includes single-cycle issue queue 0, single-cycle issue queue 1, single-cycle issue queue 2, and multi-cycle issue queue 0. Branch-type micro-operations can only enter the single-cycle issue queue, and multi-cycle micro-operations can only enter the multi-cycle issue queue; therefore, their candidate issue queue set is relatively narrow, and their type constraint strength is high. Single-cycle micro-operations can enter the single-cycle issue queue or, depending on implementation needs, the multi-cycle issue queue; therefore, their candidate issue queue set is relatively wide, and their type constraint strength is low. Based on this, the processor instruction dispatching device can first perform initial queue allocation for branch-type and multi-cycle micro-operations, and then perform initial queue allocation for single-cycle micro-operations based on the remaining available resources, thereby reducing the probability that strongly type-constrained micro-operations cannot be dispatched due to the pre-occupation of optional queue resources.

[0113] Further optionally, in the above embodiments, it is assumed that the micro-operations with higher type constraint strength include branch-type micro-operations and multi-cycle micro-operations. Based on this, the initial queue allocation for micro-operations with higher type constraint strength includes: allocating the multi-cycle micro-operations to the corresponding multi-cycle launch queues; and determining the initial launch queue corresponding to the branch-type micro-operation from the candidate launch queue set corresponding to the branch-type micro-operation based on the round-robin order and the capacity-limited launch queues.

[0114] For example, if multiple micro-operations within the current dispatch cycle, ordered by age, include a first micro-operation, a second micro-operation, a third micro-operation, and a fourth micro-operation, where the first micro-operation is a branch-type micro-operation and the third micro-operation is a multi-cycle micro-operation, then the processor instruction dispatching device can prioritize the initial queue allocation of the first and third micro-operations in the first stage. For the third micro-operation, since it can only enter multi-cycle launch queue 0, multi-cycle launch queue 0 can be determined as the initial launch queue for the third micro-operation. If multi-cycle launch queue 0 has multiple enqueue channels, its lower-order enqueue channel can be used preferentially. For the first micro-operation, since its candidate launch queue set includes single-cycle launch queue 0, single-cycle launch queue 1, and single-cycle launch queue 2, a scan can be performed among these single-cycle launch queues based on the round-robin order. If single-cycle launch queue 0 is identified as a capacity-limited launch queue by a capacity-aware mask, single-cycle launch queue 0 can be skipped preferentially if other available single-cycle launch queues exist, and the initial launch queue for the first micro-operation can be determined from single-cycle launch queue 1 or single-cycle launch queue 2.

[0115] Furthermore, in the above embodiments, it is assumed that the micro-operations with lower type constraint strength include single-cycle micro-operations. Based on this, the initial queue allocation for micro-operations with lower type constraint strength includes: determining the remaining available queuing resources in the overall launch queue set based on the occupancy of launch queue enqueue resources by micro-operations that have completed initial queue allocation; and determining the initial launch queue corresponding to the single-cycle micro-operation from the candidate launch queue set corresponding to the single-cycle micro-operation based on the remaining available queuing resources.

[0116] For example, after the initial queue allocation for branch-type micro-operations and multi-cycle micro-operations is completed in the first phase, the processor instruction dispatching device can update the remaining available enqueue resources in single-cycle issue queue 0, single-cycle issue queue 1, single-cycle issue queue 2, and multi-cycle issue queue 0. For instance, if the first micro-operation has already occupied one enqueue channel of single-cycle issue queue 1, and the third micro-operation has already occupied a low-order enqueue channel of multi-cycle issue queue 0, then when performing initial queue allocation for single-cycle micro-operations in the second phase, the selection can be based on the remaining resources after the occupancy. If the second and fourth micro-operations are single-cycle micro-operations, their initial issue queues can be selected preferentially from the remaining enqueue resources of issue queues not hit by the capacity-aware mask; when multiple issue queues not hit by the mask are available, issue queues with more remaining enqueue resources or unoccupied high-order enqueue channels can be selected preferentially; when the issue queues not hit by the mask are insufficient to accommodate the current dispatch window, it is allowed to fall back to the issue queues hit by the capacity-aware mask. Therefore, single-cycle micro-operations can complete the initial queue allocation using the remaining queue resources without affecting the priority dispatch of micro-operations with strong type constraints.

[0117] In another optional embodiment, when the processor is configured to disable capacity awareness, in step 105, based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, an initial queue allocation is performed on the multiple micro-operations, and the restriction of the capacity-limited launch queue on the initial queue allocation can also be ignored; based on the round-robin order, the initial launch queue corresponding to each micro-operation is determined from the candidate launch queue set corresponding to each micro-operation.

[0118] For example, the processor can disable capacity awareness by configuring registers, control bits, or microarchitectural parameters. When capacity awareness is disabled, even if the effective occupancy of single-cycle issue queue 0 is higher than other issue queues, the processor instruction dispatcher can skip single-cycle issue queue 0 without using a capacity-aware mask. Instead, it can directly select the initial issue queue from the candidate issue queue set corresponding to each micro-operation according to the round-robin order. For example, for branch-type micro-operations, if its candidate issue queue set includes single-cycle issue queue 0, single-cycle issue queue 1, and single-cycle issue queue 2, and the current round-robin start point points to single-cycle issue queue 0, then the round-robin selection can begin with single-cycle issue queue 0. For single-cycle micro-operations, selection can also be made from their candidate issue queue set according to the round-robin order. In this way, when capacity awareness is disabled, initial queue allocation can degenerate into a round-robin-first direct dispatch method, allowing the processor to select different dispatch strategies based on different operating modes or debugging needs.

[0119] Step 106: Based on the register read / write information of the multiple micro-operations and the initial launch queue corresponding to each micro-operation, verify the association between the register dependency relationship and the launch queue to determine the dependent launch queue corresponding to each micro-operation.

[0120] In this embodiment, the register read / write information of a micro-operation can refer to information used to characterize the micro-operation's reading or writing of registers. For example, the register read / write information may include the source logical register, destination logical register, source operand identifier, and destination operand identifier of the micro-operation. The source logical register may represent the register that needs to be read during the execution of the micro-operation, and the destination logical register may represent the register where the result is written after the micro-operation is completed. If the source logical register of a younger micro-operation is the same as the destination logical register of an older micro-operation, then a register dependency relationship can be considered between the younger and older micro-operations. The older micro-operation can be called a producer micro-operation, and the younger micro-operation can be called a consumer micro-operation.

[0121] The initial launch queue corresponding to a micro-operation can refer to the launch queue initially determined for that micro-operation during the capacity-aware initial queue allocation phase. The initial launch queue can be determined based on factors such as the micro-operation type, the candidate launch queue set, the capacity-constrained launch queue, the capacity-aware mask, the round-robin start point, and the available resources of the launch queue. This initial launch queue is not necessarily the final launch queue the micro-operation will enter; in the subsequent register dependency-aware adjustment phase, the initial launch queue can be swapped or modified based on register dependencies.

[0122] The dependent launch queue corresponding to a micro-operation can refer to the launch queue of a producer micro-operation that has a register dependency relationship with that micro-operation, or it can be a launch queue determined based on register dependency mapping information, register bypass information within the same dispatch cycle, etc., that is suitable for the micro-operation to be close to its dependency source. For example, if the source logic register of a consumer micro-operation hits the mapping information from a destination logic register to a launch queue, then the launch queue indicated by the mapping information can be used as the dependent launch queue of that consumer micro-operation. The dependent launch queue can be used to subsequently determine whether the initial launch queue of a micro-operation needs to be adjusted, so that producer micro-operations and consumer micro-operations with register dependencies are allocated to launch queues that are conducive to dependency wake-up, bypassing, and internal queue scheduling.

[0123] Producer micro-operations and consumer micro-operations are a set of relative concepts based on register data dependencies. A producer micro-operation is a micro-operation that produces a result for a specific register. In other words, if a micro-operation writes a result to a destination logical register after execution, then that micro-operation can be considered a producer micro-operation for the corresponding data in that destination logical register. A consumer micro-operation is a micro-operation that needs to read a register result as its source operand. In other words, if a micro-operation needs to read data from a source logical register, and that data was produced by an earlier producer micro-operation, then that micro-operation can be considered a consumer micro-operation for that producer micro-operation.

[0124] As an optional embodiment, in step 106, register dependency mapping information is generated based on the destination logic registers of the plurality of micro-operations and their corresponding initial launch queues; the register dependency mapping information is queried based on the source logic registers of the plurality of micro-operations; if the query is successful, the launch queue indicated by the successful register dependency mapping information is determined as the dependent launch queue of the corresponding micro-operation; wherein, when the source logic register of any consumer micro-operation is successful in the one-time mapping information, the mapping entry corresponding to the source logic register is cleared, so that the register dependency mapping information is preferentially applied to the first consumer micro-operation of the producer micro-operation, reducing the impact of historical dependencies on the dispatch selection of subsequent irrelevant micro-operations; if consumer micro-operations of different age orders are included in the same dispatch period, when querying the register dependency mapping information, the destination register bypass information of the older micro-operation in the same dispatch period is also obtained; when both the destination register bypass information and the register dependency mapping information are successful, the priority of the destination register bypass information is higher than that of the register dependency mapping information, and the launch queue allocation result of the producer micro-operation remains fixed, and the consumer micro-operation determines the dependent launch queue based on the launch queue allocation result of the producer micro-operation.

[0125] For example, the processor includes single-cycle launch queue 0, single-cycle launch queue 1, single-cycle launch queue 2, and multi-cycle launch queue 0. In the current dispatch cycle or a previous dispatch cycle, if the destination logic register of producer micro-operation A is rd, and the initial launch queue corresponding to producer micro-operation A after initial queue allocation is single-cycle launch queue 1, then the dispatching device can generate register dependency mapping information based on the destination logic register of producer micro-operation A and the initial launch queue. This register dependency mapping information can be represented as rd→single-cycle launch queue 1, used to indicate that the producer micro-operation written to the destination logic register rd is located in or will be located in single-cycle launch queue 1. Within the same dispatch cycle, if the launch queue allocation result of producer micro-operation A has been determined, then the launch queue allocation result remains fixed and can be used as the basis for determining the dependent launch queue for younger consumer micro-operations within the same dispatch cycle.

[0126] When a subsequent consumer micro-operation B participates in the dispatch, the dispatching device can query the register dependency mapping information based on the source logic register of consumer micro-operation B. If the source logic register of consumer micro-operation B is rd, then the source logic register matches the aforementioned mapping information of rd → single-cycle launch queue 1. At this time, the dispatching device can determine that consumer micro-operation B depends on producer micro-operation A written to rd, and determine single-cycle launch queue 1 as the dependent launch queue corresponding to consumer micro-operation B. Furthermore, if there are micro-operations earlier in age order than consumer micro-operation B within the same dispatch cycle, the dispatching device can also obtain the destination register bypass information of the older micro-operation while querying the register dependency mapping information; when both the destination register bypass information and the register dependency mapping information match, the launch queue allocation result of the producer micro-operation indicated by the destination register bypass information is preferentially used as the dependent launch queue of consumer micro-operation B, so as to improve the real-time performance and accuracy of dependency identification between micro-operations in the same cycle.

[0127] If consumer micro-operation B is initially assigned to single-cycle launch queue 2 during the capacity-aware initial queue allocation phase, while its dependent launch queue is single-cycle launch queue 1, it indicates that the initial launch queue of consumer micro-operation B is inconsistent with its dependent launch queue. Subsequent steps can further adjust the launch queue allocation results of consumer micro-operation B and other micro-operations based on a fixed switching network or preset adjustment rules, so that consumer micro-operation B is dispatched to a launch queue that matches its dependent launch queue as much as possible.

[0128] Optionally, the register dependency mapping information is temporary mapping information. If any of the plurality of micro-operations hits the register dependency mapping information based on a source logic register query, the hit register dependency mapping information is cleared, so that the register dependency mapping information applies to a consumer micro-operation corresponding to the producer micro-operation. By clearing the corresponding mapping entry after a hit, the lasting impact of historical dependencies on the dispatch selection of subsequent irrelevant micro-operations can be reduced, and the long-term retention of earlier generated register dependency mappings can be avoided from interfering with subsequent dispatch decisions.

[0129] For example, if the register dependency mapping information corresponding to the destination logic register rd of producer micro-operation A is rd→single-cycle launch queue 1, and the source logic register rd of consumer micro-operation B hits this mapping information, then the dispatching device can determine single-cycle launch queue 1 as the dependent launch queue of consumer micro-operation B. After the hit, the dispatching device can clear the mapping information rd→single-cycle launch queue 1, so that the mapping no longer affects the dispatch selection of subsequent micro-operations. Through this temporary or one-time mapping method, it is possible to prioritize ensuring that earlier consumers of producer micro-operations are close to the launch queue where the producer micro-operation is located, while avoiding the long-term influence of the same historical producer micro-operation on the queue allocation of multiple subsequent consumer micro-operations or unrelated micro-operations, thereby reducing the continuous interference of dependency mapping on subsequent dispatching. Furthermore, if, within the same dispatch cycle as consumer micro-operation B, there exists an older producer micro-operation C, and the source logic register of consumer micro-operation B simultaneously hits the destination register bypass information of micro-operation C and the register dependency mapping information of the aforementioned rd→single-cycle launch queue 1, then the dispatching device prioritizes determining the dependent launch queue of consumer micro-operation B based on the destination register bypass information corresponding to micro-operation C. Simultaneously, the already determined launch queue allocation results for micro-operations C or A that act as producers remain fixed, and consumer micro-operation B determines its dependent launch queue based on the launch queue allocation results of the corresponding producer micro-operation. By combining the aforementioned bypass priority with one-time mapping, both the immediate dependency identification between older producers and younger consumers in the same batch and the maintenance of queue correlation between producers and consumers across batches can be considered, thereby improving the accuracy and stability of dependency-aware dispatching.

[0130] In another optional embodiment, in step 106, for micro-operations within the same dispatch cycle, register dependencies within the same dispatch cycle are determined by matching the destination logic register of at least one micro-operation with the source logic register of the next sequential micro-operation in descending order of age. If both the register dependencies within the same dispatch cycle and the register dependency mapping information indicate dependencies, the dependency dispatch queue for the corresponding micro-operation is determined preferentially based on the register dependencies within the same dispatch cycle.

[0131] For example, within the current dispatch cycle, the dispatch buffer outputs four micro-operations in age order, denoted as micro-operation m0, micro-operation m1, micro-operation m2, and micro-operation m3. Micro-operation m0 is older than micro-operation m1, micro-operation m1 is older than micro-operation m2, and micro-operation m2 is older than micro-operation m3. After generating the initial dispatch queue for each micro-operation, the dispatch device can compare the destination logic register of the older micro-operation with the source logic register of the younger micro-operation within the same dispatch cycle.

[0132] For example, if the destination logic register of micro-operation m0 is rd, and the source logic register of micro-operation m2 is also rd, then it can be determined that micro-operation m2 depends on the older micro-operation m0 within the same dispatch cycle. If the initial launch queue corresponding to micro-operation m0 in the initial queue allocation phase is single-cycle launch queue 1, then single-cycle launch queue 1 can be determined as the dependent launch queue of micro-operation m2. In other words, for producer micro-operations and consumer micro-operations that occur simultaneously within the same dispatch cycle, the destination logic register information of the older micro-operation can be used to bypass matching the source logic register of the younger micro-operation to generate a dependent launch queue within the same cycle.

[0133] For example, if the source logic register of micro-operation m2 hits both the historical register dependency mapping information and the destination logic register of the older micro-operation m0 within the same dispatch cycle, the register dependency relationship within the same dispatch cycle can be used to determine the dependent launch queue of micro-operation m2. That is, if the older micro-operation m0 within the same dispatch cycle has already been assigned to single-cycle launch queue 1 in the current initial queue allocation, and the historical register dependency mapping information indicates another launch queue, the dispatching device can prioritize determining single-cycle launch queue 1 as the dependent launch queue of micro-operation m2.

[0134] In some embodiments, for producer micro-operations and consumer micro-operations within the same dispatch period, the initial launch queue of the producer micro-operation can remain fixed and not participate in dependency-aware exchange; the consumer micro-operation can use the initial launch queue obtained by the producer micro-operation during the initial queue allocation phase as a dependency launch queue to participate in subsequent adjustments. This dependency bypass priority mechanism within the same dispatch period ensures that newly generated register dependencies within the current dispatch window participate in queue allocation adjustments in a timely manner, reducing latency caused by waiting for historical mapping updates and improving the accuracy of register dependency-aware dispatch.

[0135] Step 107: After the verification is passed, the multiple micro-operations are dispatched to the corresponding launch queues according to the dependent launch queues corresponding to each micro-operation.

[0136] As an optional embodiment, in step 107, if the dependent launch queue corresponding to a micro-operation is inconsistent with the initial launch queue corresponding to the micro-operation, the launch queue allocation result of at least one micro-operation is adjusted based on a preset adjustment rule to obtain the target launch queue corresponding to each micro-operation. Next, producer micro-operations and consumer micro-operations with register dependencies cannot be launched simultaneously in the same batch due to data sequence relationships. Under the condition of satisfying type constraints, capacity constraints, and age order constraints, the consumer micro-operation is preferentially dispatched to the launch queue where the producer micro-operation is located, so as to shorten the dependent wake-up path and result bypass path without reducing the chance of independent micro-operations being launched in the same batch, increase the probability of other launch queues launching different micro-operations, and improve launch parallelism. When there are multiple consumer micro-operations dependent on the same producer micro-operation within the same dispatch window, the earliest consumer micro-operation is preferentially matched to the launch queue where the producer micro-operation is located according to age order. Furthermore, based on the target launch queue corresponding to each micro-operation and the available enqueue resources of each launch queue in the overall launch queue set, the multiple micro-operations are dispatched to their corresponding launch queues.

[0137] For example, after capacity-aware initial selection, the processor instruction dispatcher obtains initial issue queues for multiple micro-operations. For instance, micro-operations m0, m1, m2, and m3 in the current dispatch cycle are initially assigned to single-cycle issue queue 0, single-cycle issue queue 1, single-cycle issue queue 2, and multi-cycle issue queue 0, respectively. If, based on register dependency mapping information or register dependencies within the same dispatch cycle, it is determined that micro-operation m2 depends on a producer micro-operation located in single-cycle issue queue 1, then the dependent issue queue for micro-operation m2 is single-cycle issue queue 1. Since the initial issue queue for micro-operation m2 is single-cycle issue queue 2, which is inconsistent with its dependent issue queue single-cycle issue queue 1, the processor instruction dispatcher can adjust the issue queue allocation results of at least some micro-operations based on preset adjustment rules. For example, the issue queue allocation results of m1 and m2 can be swapped, making the target issue queue for m2 become single-cycle issue queue 1, and the target issue queue for m1 become single-cycle issue queue 2, while the target issue queues for m0 and m3 remain unchanged. This is because there is a data priority relationship between producer and consumer micro-operations with register dependencies, and they typically cannot be issued simultaneously in the same cycle. Therefore, under the constraints of type, capacity, and age order, prioritizing the adjustment of consumer micro-operations to the launch queue where producer micro-operations reside does not reduce the chance of simultaneous launch of independent micro-operations. On the contrary, it can shorten the dependent wake-up path and result bypass path, and give other launch queues a better chance to launch independent micro-operations, thereby improving the overall launch parallelism. Subsequently, the processor instruction dispatcher can dispatch each micro-operation to the corresponding launch queue according to the adjusted target launch queue and the available enqueue resources of each launch queue.

[0138] Optionally, the preset adjustment rule is implemented by a fixed switching network. The fixed switching network can be a hardware logic network composed of a finite number of comparison units and switching units, used to make limited adjustments to the micro-operation launch queue allocation results within the current dispatch window without employing a full cross-switching structure. This fixed switching network can group multiple micro-operations according to their age order, and within each group, perform adjacent, spanning, or cross-position switching judgments according to a preset priority. Since the fixed switching network only compares and switches within a limited range, it can reduce hardware complexity and improve queue locality for register-dependent micro-operations while maintaining the dispatch age order. Within the same dispatch window, if there are multiple consumer micro-operations that depend on the same producer micro-operation, the fixed switching network can prioritize matching the earliest consumer micro-operation with the launch queue of the producer micro-operation according to age order. Younger consumer micro-operations, under the premise of satisfying type constraints, capacity constraints, and age order constraints, will have their target launch queue determined by combining available switching opportunities, thus avoiding multiple consumers simultaneously competing for the same launch queue and affecting the overall dispatch effect.

[0139] Based on this, the allocation result of the launch queue for at least one micro-operation is adjusted according to a preset adjustment rule, including: dividing the multiple micro-operations into at least one adjustment group according to age order; and adjusting the allocation result of the launch queue for at least two micro-operations in each adjustment group according to a preset priority through the fixed switching network, so that consumer micro-operations with register dependencies are preferentially allocated to launch queues that match the corresponding dependent launch queues. In one embodiment, when multiple consumer micro-operations depend on the same producer micro-operation within the same adjustment group, the fixed switching network preferentially performs adjustment judgment on the oldest consumer micro-operation that has not yet matched a dependent launch queue, so that it preferentially obtains a target launch queue consistent with the launch queue of the producer micro-operation. Further, for any two micro-operations to be compared in the fixed switching network, if the dependent launch queue of the first micro-operation is equal to the initial launch queue of the second micro-operation, or the dependent launch queue of the second micro-operation is equal to the initial launch queue of the first micro-operation, and the current initial launch queue of the micro-operation with the dependent launch queue is different from its dependent launch queue, then the two micro-operations are determined as candidate micro-operations for switching. Under the constraints of micro-operation type, launch queue capacity, maximum number of entries per launch queue per cycle, and fixed producer micro-operation queue within the same dispatch cycle, the fixed switching network exchanges the launch queue allocation results of the two micro-operations. If the previous priority level has already been exchanged, the subsequent priority is determined based on the queue allocation result after the previous level exchange. If the exchange conditions are not met, the original queue allocation result is retained.

[0140] Here, the execution of the launch queue allocation result for the micro-operations is adjusted without changing the age order of the multiple micro-operations.

[0141] For example, in the current dispatch cycle, the dispatch buffer outputs eight micro-operations m0 to m7 in age order. The processor instruction dispatching device can divide m0 to m3 into a first adjustment group and m4 to m7 into a second adjustment group. The fixed switching network can perform limited switching decisions within each adjustment group. For example, for the first adjustment group, if m2 is a consumer micro-operation that depends on a producer micro-operation located in single-cycle launch queue 1, and m2 is assigned to single-cycle launch queue 2 in the initial queue allocation, then the fixed switching network can adjust the launch queue allocation results of at least two of these micro-operations without changing the age order of m0, m1, m2, and m3, so that m2 can obtain a target launch queue that matches its dependent launch queue as much as possible. If there is another younger consumer micro-operation in the same adjustment group that also depends on the producer micro-operation, then the fixed switching network prioritizes matching the older m2 with the launch queue of the producer micro-operation, and handles the queue adjustment of the younger consumer micro-operation in a subsequent priority decision. It should be noted that the adjustment swaps the micro-operation launch queue allocation results, rather than swapping the position of the micro-operations in the dispatch buffer. Therefore, it will not cause younger micro-operations to bypass older micro-operations when they are enqueued.

[0142] Further optionally, the preset priority includes at least one of adjacent position adjustment priority, cross-position adjustment priority, and cross-position adjustment priority. Wherein, if the adjustment group includes four micro-operations, and the order of the four micro-operations is arranged in age sequence, then the adjacent position adjustment priority is used to determine the adjustment of micro-operations between the first and second positions, and between the third and fourth positions; the cross-position adjustment priority is used to determine the adjustment of micro-operations between the first and third positions, and between the second and fourth positions; and the cross-position adjustment priority is used to determine the adjustment of micro-operations between the first and fourth positions, and between the second and third positions.

[0143] For example, if an adjustment group includes four micro-operations m0, m1, m2, and m3 arranged in age order, where m0 is the first position, m1 is the second position, m2 is the third position, and m3 is the fourth position, the fixed switching network can perform adjustment judgments sequentially according to preset priorities. First, under the adjacent position adjustment priority, it can be determined whether m0 and m1 need to exchange transmission queue allocation results, and whether m2 and m3 need to exchange transmission queue allocation results. Second, under the cross-position adjustment priority, it can be determined whether m0 and m2 need to exchange transmission queue allocation results, and whether m1 and m3 need to exchange transmission queue allocation results. Third, under the cross-position adjustment priority, it can be determined whether m0 and m3 need to exchange transmission queue allocation results, and whether m1 and m2 need to exchange transmission queue allocation results.

[0144] For example, the initial queue allocation is as follows: m0 enters single-cycle launch queue 0, m1 enters single-cycle launch queue 1, m2 enters single-cycle launch queue 2, and m3 enters multi-cycle launch queue 0. If m2's dependent launch queue is single-cycle launch queue 1, and m1's initial launch queue is also single-cycle launch queue 1, then when adjusting m1 and m2 in the fixed switching network, it can be determined that m2's dependent launch queue matches m1's initial launch queue, and m2's current initial launch queue is different from its dependent launch queue. Under the conditions of satisfying type constraints, capacity constraints, per-queue per-cycle reception quantity constraints, and producers not participating in the exchange within the same dispatch cycle, the launch queue allocation results of m1 and m2 can be exchanged, making m2's target launch queue become single-cycle launch queue 1, and m1's target launch queue become single-cycle launch queue 2. In this way, consumer micro-operation m2 can be preferentially adjusted to the launch queue where its corresponding producer micro-operation is located, shortening the dependent wake-up path and result bypass path when m2 is waiting for the producer's result without reducing the opportunity for independent micro-operations to be launched in the same cycle. If other consumer micro-operations that also depend on the same producer micro-operation exist within the same dispatch window, priority will be given to ensuring that consumer micro-operations with higher age order complete the above matching first. After the previous priority level completes the exchange, subsequent priorities can continue to be determined based on the queue allocation result after the exchange; if the exchange conditions are not met, the original queue allocation result will be retained.

[0145] In another optional embodiment, step 107 involves dispatching the multiple micro-operations to their corresponding launch queues based on the target launch queues corresponding to each micro-operation and the available enqueue resources of each launch queue in the overall launch queue set. This includes: checking the enqueue resources of the target launch queues corresponding to each micro-operation according to their age order; if the available enqueue resources of the same target launch queue can receive multiple micro-operations, writing the micro-operation with the earlier age order into the lower-order enqueue channel of the same target launch queue, and writing the micro-operation with the later age order into the higher-order enqueue channel of the same target launch queue; if the available enqueue resources of the target launch queue do not meet the enqueue requirements of the corresponding micro-operation, blocking the corresponding micro-operation and the micro-operations that are located after the corresponding micro-operation according to their age order.

[0146] For example, the final queue allocation result can still be checked for enqueued resources according to the age order of micro-operations. Each launch queue can receive a maximum of two micro-operations per clock cycle, written through a low-order enqueue channel and a high-order enqueue channel respectively. The low-order enqueue channel can be used for micro-operations with higher age in the same target launch queue, and the high-order enqueue channel can be used for micro-operations with lower age in the same target launch queue.

[0147] For example, if the target launch queue is a single-cycle launch queue 1, and the remaining capacity of single-cycle launch queue 1 is greater than or equal to 2, then the processor instruction dispatching device can write two micro-operations, both targeting single-cycle launch queue 1, into the low-order enqueue channel and the high-order enqueue channel of that queue, respectively, within the same dispatch cycle. The older micro-operation is written into the low-order enqueue channel, and the younger micro-operation is written into the high-order enqueue channel.

[0148] For example, if the target launch queue is a single-cycle launch queue 2, and the remaining capacity of single-cycle launch queue 2 is equal to 1, then for multiple micro-operations targeting single-cycle launch queue 2, only the older micro-operations are allowed to be written to single-cycle launch queue 2 through the low-order enqueue channel. The remaining younger micro-operations targeting single-cycle launch queue 2 can remain in the dispatch buffer. As another example, if the target launch queue is a single-cycle launch queue 0, and the remaining capacity of single-cycle launch queue 0 is 0, then micro-operations targeting single-cycle launch queue 0 cannot be written to this queue, and this micro-operation, along with the younger micro-operations following it in age order, can be blocked to maintain the age order of the dispatch process.

[0149] After successfully enqueuing a dispatched micro-operation, the processor instruction dispatcher can update the rotation pointer based on the last successfully dispatched micro-operation, so that the capacity-aware initial selection for the next dispatch cycle starts from the new rotation starting point. If dispatch is interrupted due to insufficient capacity in a certain issue queue, the undispatched micro-operations can retain their original age relationships and wait for the next dispatch cycle to re-participate in capacity-aware initial selection, register dependency-aware adjustment, and final enqueue check.

[0150] In this embodiment, the processor can simultaneously consider micro-operation type, age order, issue queue capacity status, and register dependencies during the processor instruction dispatch process, thereby reducing issue queue load imbalance and hotspot queue congestion, improving issue queue resource utilization, and reducing register dependency-related latency, thus improving the overall execution efficiency of the processor dispatch phase and subsequent issue phases.

[0151] See Figure 4 , Figure 4 This is a schematic diagram illustrating the principle of a final queuing channel selection and blocking method provided in an embodiment of this application. Figure 4 As shown, after completing the capacity-aware initial selection and register dependency-aware exchange, the target launch queue corresponding to each micro-operation can be obtained. This target launch queue can be the initial launch queue obtained from the capacity-aware initial selection, or it can be the dependent launch queue corrected by the register dependency-aware exchange. Further, based on the available enqueue channel information of each target launch queue within the current dispatch cycle, it is determined whether each micro-operation can actually be written to the corresponding launch queue.

[0152] In this embodiment, each launch queue may have one or more enqueue channels within a dispatch period, such as low-order enqueue channels and high-order enqueue channels. The enqueue resources of the target launch queue are checked sequentially according to the age order of the micro-operations. When a target launch queue has an available enqueue channel, the corresponding micro-operation is written to that launch queue, and the available enqueue channel status of that launch queue in the current dispatch period is updated. When a target launch queue does not have an available enqueue channel, or the remaining capacity of the launch queue is insufficient to receive new micro-operations, the dispatch of that micro-operation is blocked, and subsequent minor micro-operations may also be blocked to maintain the age order of micro-operations within the dispatch window.

[0153] Furthermore, for micro-operations corrected by register dependency-aware switching, the dispatching device uses the corrected target launch queue as the basis for final enqueue determination. In other words, register dependency-aware switching only changes the queue selection result corresponding to the micro-operation, without altering its age position in the dispatch buffer. The final enqueueing stage still consumes micro-operations in the dispatch buffer according to their age order, and determines whether to complete the dispatch based on the available channels and remaining capacity of the target launch queue. Thus, while maintaining program order constraints and dispatch order constraints, it is possible to ensure that consumer micro-operations with register dependencies enter the launch queue of the producer as much as possible.

[0154] In one example, if after Figure 3 After the register-dependent sensing swap shown, the target emission queue for micro-operation m2 is changed from single-cycle queue2 to single-cycle queue1. Figure 4 The final enqueueing phase, as shown, checks whether single-cycle queue1 still has an available enqueue channel within the current dispatch cycle. If single-cycle queue1 has an available enqueue channel, then m2 is dispatched to single-cycle queue1. If the enqueue channel of single-cycle queue1 is already occupied by an older micro-operation, or if the remaining capacity of single-cycle queue1 is insufficient, then m2 is not dispatched in the current cycle and remains in the dispatch buffer, waiting to participate in the dispatch decision in a subsequent cycle.

[0155] By employing the above methods, the final queuing phase integrates capacity constraints, queuing channel constraints, micro-operation age order, and register dependency-aware exchange results into the actual dispatch action, avoiding situations where only queue selection is completed but queuing resources cannot be satisfied. Simultaneously, it ensures that the number of micro-operations received per cycle in the same transmit queue does not exceed its hardware queuing capacity, thereby improving transmit queue load balancing and dependency locality while reducing the risk of conflicts between dispatch logic and transmit queue write ports.

[0156] See Figure 5 , Figure 5 This is another schematic diagram illustrating the principle of a final queuing channel selection and blocking method provided in an embodiment of this application. For example... Figure 5 As shown, after completing the capacity-aware initial selection and register-dependent-aware exchange, the dispatching device obtains the final queue allocation result for multiple micro-operations. This final queue allocation result is still subject to enqueue checks according to the age order of the micro-operations to ensure that younger micro-operations do not skip older micro-operations to complete the dispatch.

[0157] In this embodiment, each launch queue can receive a maximum of two micro-operations within a dispatch cycle, and these are written through the first enqueue channel enq0 and the second enqueue channel enq1, respectively. Here, enq0 is used for older micro-operations within the same target launch queue, and enq1 is used for younger micro-operations within the same target launch queue. The dispatching device can determine whether one or more micro-operations with the same target can be enqueued in the same cycle based on the remaining capacity free_cnt of the target launch queue and the enqueue channels already occupied in the current cycle.

[0158] For example, when both micro-operations target the single-cycle queue1, and the remaining capacity of single-cycle queue1, free_cnt, is greater than or equal to 2, the dispatching device can write the two micro-operations to enq0 and enq1 of single-cycle queue1 respectively within the same dispatch cycle. The older micro-operation is written through enq0, and the younger micro-operation is written through enq1. Thus, without disrupting the age order, multiple enqueue channels within the same launch queue can be fully utilized, improving dispatch bandwidth.

[0159] For example, when multiple micro-operations are all targeted to the single-cycle queue2, and the remaining capacity of single-cycle queue2, free_cnt, is equal to 1, the dispatching device only allows the oldest micro-operation to be written to single-cycle queue2 via enq0. The other, younger micro-operations also targeted to single-cycle queue2 cannot be written to this queue in the current cycle and remain in the dispatch buffer waiting to participate in dispatch again in subsequent cycles. In this case, even if the queue has an enq1 channel in hardware, enq1 will not be used to write new micro-operations because the remaining capacity is insufficient to receive two micro-operations simultaneously.

[0160] For example, if the target launch queue for a micro-operation is single-cycle queue0, and the remaining capacity of single-cycle queue0, free_cnt, is 0, it means that the target launch queue cannot currently accept new micro-operations. In this case, the micro-operation cannot be enqueued within the current dispatch cycle, and younger micro-operations that follow it in age order are also blocked and retained in the dispatch buffer. This blocking rule prevents younger micro-operations from bypassing older micro-operations to enter the launch queue, thus maintaining the constraint of the age order of micro-operations during the dispatch phase.

[0161] After final enqueueing, the dispatching device can update the rotation pointer based on the last successfully dispatched micro-operation in the current cycle, so that the capacity-aware initial selection for the next dispatch cycle starts from the updated rotation starting point. If dispatching is interrupted due to insufficient capacity in a target launch queue or insufficient enqueueing channels, the micro-operations that were not successfully dispatched will retain their original age relationships and re-participate in capacity-aware initial selection, register dependency-aware exchange, and final enqueueing check in the next dispatch cycle.

[0162] Through the above methods Figure 5 The final enqueue channel selection and blocking process shown combines the final queue allocation result with the actual write capacity of the launch queue. When the target queue has sufficient remaining capacity and enqueue channels, multiple micro-operations targeting the same target can be enqueued in the same cycle. When the target queue capacity is insufficient, related micro-operations and subsequent minor micro-operations can be blocked promptly, thus balancing dispatch bandwidth, queue capacity constraints, and the age order of micro-operations.

[0163] The processor instruction dispatching method of the present application has been described above. The processor instruction dispatching apparatus for executing the above processor instruction dispatching method will be described below.

[0164] See Figure 6 ,like Figure 6 The diagram shows a schematic of a processor instruction dispatching device. The processor instruction dispatching device in this embodiment can achieve the functions described above. Figure 1 The steps of the processor instruction dispatching method executed in the corresponding embodiments are described above. The functions implemented by the processor instruction dispatching device can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, and the modules can be software and / or hardware. The processor instruction dispatching device may include an input / output module 601 and a processing module 602. The functional implementation of the processing module 602 and the input / output module 601 can be found in [reference]. Figure 1 The operations performed in the corresponding embodiments are not described in detail here. For example, the processing module 602 can be used to control the sending, receiving, and acquiring operations of the input / output module 601. Specifically, the device implements the processor instruction dispatch process through the following module functions, wherein: The input / output module 601 is configured to acquire multiple micro-operations to be dispatched; acquire the overall set of launch queues available for dispatch in the processor and the corresponding queue status information, wherein the queue status information includes the remaining capacity information and effective occupancy information of each launch queue in the overall set of launch queues; The processing module 602 is configured to determine a set of candidate launch queues corresponding to each micro-operation based on the micro-operation type and age order of the plurality of micro-operations. The processing module 602 is further configured to determine the capacity-limited launch queue in the overall launch queue set based on the queue status information; The processing module 602 is further configured to perform initial queue allocation on the multiple micro-operations based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, so as to obtain the initial launch queue corresponding to each micro-operation. The processing module 602 is further configured to verify the association between register dependencies and launch queues based on the register read / write information of the plurality of micro-operations and the initial launch queues corresponding to each micro-operation, so as to determine the dependent launch queues corresponding to each micro-operation. The processing module is further configured to, after the verification is passed, dispatch the multiple micro-operations to the corresponding launch queues through the input / output module 601 according to the dependent launch queues corresponding to each micro-operation.

[0165] In this embodiment, the input / output module 601 can interact with modules such as the dispatch buffer, launch queue, capacity counter, and register-related information maintenance unit to obtain information on multiple micro-operations to be dispatched, the remaining capacity information, effective occupancy information, and available enqueue resource information of each launch queue, and write the micro-operations after queue allocation to the corresponding launch queue. The processing module 602 can be used to perform processing logic such as type identification, candidate launch queue set determination, capacity-limited launch queue determination, initial queue allocation, register dependency verification, and queue allocation adjustment.

[0166] In one possible implementation, the input / output module 601 can obtain multiple micro-operations arranged in age order from the dispatch buffer. The processing module 602 can determine the corresponding candidate launch queue set based on the micro-operation type of each micro-operation, and determine the capacity-constrained launch queue by combining the remaining capacity information and effective occupancy information of each launch queue in the overall launch queue set. The processing module 602 can also perform initial queue allocation on the multiple micro-operations based on the candidate launch queue set and the capacity-constrained launch queue to obtain the initial launch queue corresponding to each micro-operation.

[0167] In another possible implementation, processing module 602 can determine the association between register dependencies and launch queues based on register read / write information of multiple micro-operations and the initial launch queues corresponding to each micro-operation. For example, processing module 602 can generate register dependency mapping information based on the destination logic register of a micro-operation and its corresponding initial launch queue, and query this register dependency mapping information based on the source logic register of subsequent micro-operations to determine the dependent launch queue for the corresponding micro-operation. If the dependent launch queue is inconsistent with the initial launch queue, processing module 602 can adjust the launch queue allocation results of at least some micro-operations to obtain the target launch queue.

[0168] In another possible implementation, processing module 602 can adjust the launch queue allocation results via a fixed switching network. This fixed switching network can divide multiple micro-operations into at least one adjustment group according to age order, and within each adjustment group, adjust the launch queue allocation results of at least two micro-operations according to a preset priority. This adjustment can affect the launch queue allocation results of micro-operations without changing the age order of the multiple micro-operations. Input / output module 601, after processing module 602 determines the target launch queue, can dispatch micro-operations to the corresponding launch queue according to the available enqueue resources of each launch queue; if the available enqueue resources of a target launch queue are insufficient, the corresponding micro-operation and the micro-operations following it in age order can be retained in the dispatch buffer.

[0169] In a further embodiment, the processor instruction dispatching device may also introduce a fast non-ready flag. The fast non-ready flag can be used to apply a conservative delay to consumer micro-operations that have a register read-after-write relationship within the same rename group, or to consumer micro-operations that rely on long-latency producer micro-operations, to prevent consumer micro-operations from incorrectly participating in the launch immediately after dispatch based on bypass or readiness information.

[0170] When a pause occurs during the dispatch phase, or when the target launch queue capacity is insufficient or the enqueue channel is insufficient, the processor instruction dispatching device can clear the fast non-ready flag associated with the blocked micro-operation. This allows the undispatched micro-operation to re-participate in capacity-aware initial selection, register dependency-aware swapping, and final enqueue check in subsequent dispatch cycles without being affected by the expired fast non-ready state.

[0171] With the above-described device structure, the processor instruction dispatching device can simultaneously combine micro-operation type, age order, issue queue capacity status, and register dependencies during the dispatching phase to allocate queues. This helps reduce dependency-related delays caused by uneven issue queue load and cross-queue allocation of register-related micro-operations, thereby improving processor dispatching efficiency.

[0172] The processor instruction dispatching device in the embodiments of this application has been described above from the perspective of modular functional entities. For parts not described herein, please refer to the specific description in the method embodiments; they will not be elaborated upon here. The processor instruction dispatching device in the embodiments of this application will now be described from the perspective of hardware processing.

[0173] It should be noted that, Figure 6 The physical device corresponding to the input / output module 601 shown can be a transceiver, radio frequency circuit, communication module, and input / output (I / O) interface, etc., and the physical device corresponding to the processing module 602 can be a processor.

[0174] This application also provides a server; please refer to [link / reference]. Figure 7 , Figure 7 This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1100 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1122 (e.g., one or more processors) and memory 1132, and one or more storage media 1130 (e.g., one or more mass storage devices) for storing application programs 1142 or data 1144. The memory 1132 and storage media 1130 can be temporary or persistent storage. The program stored in the storage media 1130 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the server. Furthermore, the CPU 1122 may be configured to communicate with the storage media 1130 and execute the series of instruction operations in the storage media 1130 on the server 1100. Server 1100 may also include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input / output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. The steps performed by the server in the above embodiments can be based on this... Figure 7 The structure of server 1100 shown. For example, as in the above embodiment, by Figure 6 The steps performed by the processor instruction dispatching device shown can be based on this Figure 7 The server structure is shown. For example, the central processing unit 1122 performs the following operations by calling instructions from memory 1132: The input / output interface 1158 is used to acquire multiple micro-operations to be dispatched; the overall set of launch queues available for dispatch in the processor and the corresponding queue status information are acquired, the queue status information including the remaining capacity information and effective occupancy information of each launch queue in the overall set of launch queues; The central processing unit 1122 is configured to determine a candidate launch queue set corresponding to each micro-operation based on the micro-operation type and age order of the plurality of micro-operations; determine a capacity-limited launch queue in the overall launch queue set based on the queue status information; perform initial queue allocation on the plurality of micro-operations based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue to obtain an initial launch queue corresponding to each micro-operation; verify the association between register dependencies and launch queues based on the register read / write information of the plurality of micro-operations and the initial launch queues corresponding to each micro-operation to determine the dependent launch queues corresponding to each micro-operation; and dispatch the plurality of micro-operations to the corresponding launch queues through the input / output module 1158 after the verification is passed.

[0175] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0176] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0177] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings, direct couplings, or communication connections may be indirect couplings or communication connections through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0178] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0179] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0180] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, it generates, in whole or in part, the processes or functions described in the embodiments of this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0181] The technical solutions provided in the embodiments of this application have been described in detail above. Specific examples have been used in the embodiments of this application to illustrate the principles and implementation methods of the embodiments of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.

Claims

1. A processor instruction dispatching method, characterized in that, The method includes: Retrieve multiple micro-operations to be dispatched; The candidate launch queue set corresponding to each micro-operation is determined based on the micro-operation type and age order of the multiple micro-operations. Obtain the overall set of launch queues available for dispatch in the processor and the corresponding queue status information, wherein the queue status information includes the remaining capacity information and effective occupancy information of each launch queue in the overall set of launch queues; Based on the queue status information, determine the capacity-limited launch queues in the overall launch queue set; Based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, the multiple micro-operations are initially queued to obtain the initial launch queue corresponding to each micro-operation. Based on the register read / write information of the multiple micro-operations and the initial launch queue corresponding to each micro-operation, the association between register dependencies and launch queues is verified to determine the dependent launch queue corresponding to each micro-operation. After the verification is passed, the multiple micro-operations are dispatched to the corresponding launch queues according to the dependent launch queues of each micro-operation.

2. The processor instruction dispatching method according to claim 1, characterized in that, The overall launch queue set includes single-cycle launch queues and multi-cycle launch queues; determining the candidate launch queue set corresponding to each micro-operation based on the micro-operation type and age order of the multiple micro-operations includes: The micro-operation type of each micro-operation is identified according to the age order of the multiple micro-operations; Based on the launch queue type constraints corresponding to the micro-operation type, the range of launch queues that each micro-operation can use is determined from the overall launch queue set; Among these, branch-type micro-operations can use the single-cycle launch queue, multi-cycle micro-operations can use the multi-cycle launch queue, and single-cycle micro-operations can use the single-cycle launch queue and / or the multi-cycle launch queue.

3. The processor instruction dispatching method according to claim 2, characterized in that, The weights of the launch queue type constraints for the branch-type micro-operations and the multi-cycle micro-operations are higher than the weights of the launch queue type constraints for the single-cycle micro-operations. The initial queue allocation for the multiple micro-operations based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue includes: Based on the age order of the multiple micro-operations, the branch-type micro-operation and the multi-period micro-operation are determined from the multiple micro-operations; For the multi-cycle micro-operation, a corresponding multi-cycle launch queue is determined based on the candidate launch queue set corresponding to the multi-cycle micro-operation, and the multi-cycle launch queue is used as the initial launch queue for the multi-cycle micro-operation. For the branch-type micro-operation, based on the candidate launch queue set corresponding to the branch-type micro-operation, the capacity-limited launch queue, and the rotation order, the initial launch queue of the branch-type micro-operation is determined from the single-cycle launch queues that can be used for the branch-type micro-operation. After the branch-type micro-operations and the multi-cycle micro-operations have completed their initial queue allocation, the single-cycle micro-operations are initially allocated based on the remaining launch queue resources occupied by the allocated micro-operations.

4. The processor instruction dispatching method according to claim 1, characterized in that, The step of determining the capacity-limited launch queue in the overall launch queue set based on the queue status information includes: Based on the effective occupancy information of each launch queue in the overall launch queue set, determine the difference between the effective occupancy number of any launch queue and the effective occupancy number of at least one other launch queue. If the difference is greater than a preset threshold, any one of the transmission queues is identified as a capacity-limited transmission queue, and a capacity-aware mask is generated to identify the capacity-limited transmission queue.

5. The processor instruction dispatching method according to claim 1, characterized in that, The initial queue allocation for the multiple micro-operations based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue includes: Based on the rotation starting point, from the candidate launch queue set corresponding to each micro-operation, the launch queue that has not been determined as the capacity-limited launch queue is preferentially selected as the initial launch queue for the corresponding micro-operation. If a launch queue not identified as a capacity-constrained launch queue does not meet the dispatch requirements, the initial launch queue for the corresponding micro-operation can be selected from the capacity-constrained launch queue to balance launch queue load balancing and dispatch bandwidth.

6. The processor instruction dispatching method according to claim 1, characterized in that, The step of verifying the association between register dependencies and launch queues based on the register read / write information of the multiple micro-operations and the initial launch queues corresponding to each micro-operation, in order to determine the dependent launch queues corresponding to each micro-operation, includes: Based on the target logic registers of the multiple micro-operations and the corresponding initial emit queues, register dependency mapping information is generated; The register dependency mapping information is queried based on the source logic register of the multiple micro-operations; If the query is hit, the issue queue indicated by the hit register dependency mapping information is determined as the dependency issue queue of the corresponding micro-operation. Wherein, when the source logic register of any consumer micro-operation hits the one-time mapping information, the mapping entry corresponding to the source logic register is cleared, so that the register-dependent mapping information is preferentially applied to the first consumer micro-operation of the producer micro-operation, reducing the impact of historical dependencies on the dispatch selection of subsequent irrelevant micro-operations. If consumer micro-operations of different age order are included in the same dispatch cycle, the destination register bypass information of the older micro-operation in the same dispatch cycle is also obtained when querying the register dependency mapping information. When both the destination register bypass information and the register dependency mapping information are hit, the destination register bypass information has a higher priority than the register dependency mapping information, and the producer micro-operation's launch queue allocation result remains fixed. The consumer micro-operation determines the dependent launch queue based on the producer micro-operation's launch queue allocation result.

7. The processor instruction dispatching method according to claim 1, characterized in that, The step of dispatching the multiple micro-operations to their respective launch queues based on their dependent launch queues includes: When the dependent launch queue corresponding to a micro-operation is inconsistent with the initial launch queue corresponding to the micro-operation, the launch queue allocation result of at least one micro-operation is adjusted based on a preset adjustment rule to obtain the target launch queue corresponding to each micro-operation. Producer micro-operations and consumer micro-operations with register dependencies cannot be launched simultaneously in the same timeframe due to data sequence relationships. Under the condition of satisfying type constraints, capacity constraints, and age order constraints, the consumer micro-operation is preferentially dispatched to the launch queue where the producer micro-operation is located. When there are multiple consumer micro-operations that depend on the same producer micro-operation within the same dispatch window, the earliest consumer micro-operation is preferentially matched with the launch queue where the producer micro-operation is located according to age order. Based on the target launch queue corresponding to each micro-operation and the available enqueue resources of each launch queue in the overall launch queue set, the multiple micro-operations are dispatched to the corresponding launch queues.

8. The processor instruction dispatching method according to claim 7, characterized in that, The preset adjustment rules are implemented by a fixed switching network; the adjustment of the transmit queue allocation result for at least one micro-operation based on the preset adjustment rules includes: The multiple micro-operations are divided into at least one adjustment group according to age order; The fixed switching network adjusts the launch queue allocation results of at least two micro-operations in each adjustment group according to a preset priority, so that consumer micro-operations with register dependencies are preferentially allocated to launch queues that match the corresponding dependent launch queues. In a fixed switching network, for any two micro-operations to be compared, if the dependent launch queue of the first micro-operation is equal to the initial launch queue of the second micro-operation, or the dependent launch queue of the second micro-operation is equal to the initial launch queue of the first micro-operation, and the current initial launch queue of the micro-operation with the dependent launch queue is different from its dependent launch queue, then the two micro-operations are determined as candidate micro-operations for switching. Under the conditions of satisfying the constraints of micro-operation type, launch queue capacity, maximum number of entries per launch queue per tick, and fixed constraints of producer micro-operation queue within the same dispatch cycle, the fixed exchange network exchanges the launch queue allocation results of the two micro-operations; if the previous priority has been exchanged, the subsequent priority is determined based on the queue allocation result after the previous exchange; if the exchange conditions are not met, the original queue allocation result is retained.

9. The processor instruction dispatching method according to claim 7, characterized in that, The step of dispatching the multiple micro-operations to their corresponding launch queues based on the target launch queues corresponding to each micro-operation and the available enqueue resources of each launch queue in the overall launch queue set includes: According to the age order of the multiple micro-operations, the target launch queue corresponding to each micro-operation is checked for enqueue resources; When the available enqueue resources of the same target launch queue can receive multiple micro-operations, the micro-operations with the higher age order are written into the low-order enqueue channel of the same target launch queue, and the micro-operations with the lower age order are written into the high-order enqueue channel of the same target launch queue. If the available enqueue resources in the target launch queue do not meet the enqueue requirements of the corresponding micro-operation, the corresponding micro-operation and the micro-operations that follow the corresponding micro-operation in age order are blocked.

10. A processor instruction dispatching device, characterized in that, The device includes: The input / output module is configured to acquire multiple micro-operations to be dispatched; acquire the overall set of launch queues available for dispatch in the processor and the corresponding queue status information, wherein the queue status information includes the remaining capacity information and effective occupancy information of each launch queue in the overall set of launch queues; The processing module is configured to determine the candidate launch queue set corresponding to each micro-operation based on the micro-operation type and age order of the plurality of micro-operations; The processing module is further configured to determine the capacity-limited launch queue in the overall launch queue set based on the queue status information; The processing module is further configured to perform initial queue allocation on the multiple micro-operations based on the candidate launch queue set corresponding to each micro-operation and the capacity-limited launch queue, so as to obtain the initial launch queue corresponding to each micro-operation. The processing module is further configured to verify the association between register dependencies and launch queues based on the register read / write information of the multiple micro-operations and the initial launch queues corresponding to each micro-operation, so as to determine the dependent launch queues corresponding to each micro-operation. The processing module is further configured to, after the verification is passed, dispatch the multiple micro-operations to the corresponding launch queues through the input / output module according to the dependent launch queues corresponding to each micro-operation.