Apparatus, system, and method for efficiently picking micro-operations to be executed

By replacing excess complex micro-operations with simple ones, the method ensures all execution resources are utilized, addressing underutilization issues and enhancing processor efficiency.

JP2025520525APending Publication Date: 2025-07-03ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024573863
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-30
Filing Date
2023-06-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Processors face inefficiencies due to the underutilization of execution resources when the number of complex micro-operations exceeds the available binary multipliers and/or FPU resources, leading to incomplete and underutilized picks, which degrade performance.

Method used

A method that selects a first set of micro-operations for execution and replaces excess complex micro-operations with simple micro-operations, ensuring all execution resources are utilized by matching the number of complex and simple resources within the processor.

Benefits of technology

This approach enhances processor performance by ensuring all execution resources are fully utilized, preventing incomplete picks and improving efficiency by replacing complex micro-operations with simple ones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520525000001_ABST
    Figure 2025520525000001_ABST
Patent Text Reader

Abstract

The disclosed method for efficiently picking micro-operations to be executed includes selecting a first set of micro-operations that are ready for execution during a particular clock cycle. The method also includes selecting a second set of micro-operations that are ready for execution during a particular clock cycle. The method additionally includes replacing one or more of the complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, at least partially due to the number of complex micro-operations included in the first set of micro-operations exceeding the set of complex resources capable of executing the complex micro-operations. Various other apparatuses, systems, and methods are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Priority Description) This application claims the benefit of priority to U.S. Patent Application No. 17 / 855,528, filed Jun. 30, 2022, the disclosure of which is hereby incorporated by reference in its entirety.

Background Art

[0002] Processors often include a picker responsible for picking a group of micro-operations (commonly called micro-ops) that are fed to execution resources such as an arithmetic logic unit (ALU), a binary multiplier, and / or a floating point unit (FPU) for execution. In some examples, the ALU cannot perform and / or execute certain complex micro-operations (e.g., multiplication and / or division operations). In these examples, a binary multiplier and / or an FPU can perform and / or execute such complex micro-operations. However, these binary multipliers and / or FPUs may require and / or consume more space and / or real estate on such processors than the ALU. For this reason, manufacturers often choose and / or prefer processor architectures that include more ALUs than binary multipliers and / or FPUs.

[0003] Accordingly, the present disclosure recognizes and addresses the need for additional improved apparatuses, systems, and methods for efficiently picking micro-operations in view of the number of execution resources (e.g., ALUs, binary multipliers, and / or FPUs) included in a particular processor architecture.

[0004] The accompanying drawings illustrate some exemplary embodiments and are a part of this specification. Together with the following description, these drawings demonstrate and explain various principles of the present disclosure.

Brief Description of the Drawings

[0005]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

DETAILED DESCRIPTION OF THE INVENTION

[0006] Throughout the drawings, the same reference numerals and descriptions indicate elements that are similar but not necessarily identical. Exemplary embodiments described herein may be subject to various modifications and alternative forms, but specific embodiments are shown by way of example in the drawings and described in detail herein. However, the exemplary embodiments described herein are not limited to the specific forms disclosed. Rather, this disclosure encompasses all modifications, equivalents, and alternative forms included within the scope of the appended claims.

[0007] This disclosure describes various apparatuses, systems, and methods for efficiently picking micro-operations for execution. As will be described in more detail below, the various apparatuses, systems, and / or methods described herein can provide various benefits and / or advantages over certain conventional embodiments of processors, pipelines, and / or pickers.

[0008] In some cases, a binary multiplier and / or FPU performs and / or executes complex micro-operations (e.g., multiplication and / or division operations) that an ALU cannot perform and / or execute. However, since these binary multipliers and / or FPUs may require and / or consume more space and / or real estate on such a processor than an ALU, manufacturers often choose and / or prefer a processor architecture that includes more ALUs than binary multipliers and / or FPUs. Unfortunately, if the number of complex micro-operations selected by a picker within a given clock cycle exceeds the number of binary multipliers and / or FPUs in a processor, the extra complex micro-operations are dropped and / or removed from the pick.

[0009] For example, if a processor includes five ALUs and one binary multiplier, the picker within the processor can select six micro-operations per clock cycle. However, in this example, if the picker selects two or more complex micro-operations within a given clock cycle, the picker is forced to drop all complex micro-operations that exceed one, at least partially due to the processor including only one binary multiplier. This dropping can result in incomplete and / or underutilized picks and thus potentially degrade the performance and / or efficiency of the processor.

[0010] The various apparatuses, systems, and / or methods described herein can address and / or resolve such incomplete and / or underutilized picks and thus improve the performance and / or efficiency of the processor. For example, the various apparatuses, systems, and / or methods described herein can ensure that complex micro-operations (e.g., multiplication and / or division operations) dropped by a picker due to insufficient complex resources within a given clock cycle are replaced by simple micro-operations (e.g., addition, subtraction, and / or comparison operations). By doing so, these apparatuses, systems, and / or methods can avoid issuing incomplete and / or underutilized picks having empty slots and thus potentially improve the performance and / or efficiency of the processor in which the picker is implemented.

[0011] In one example, a method for performing such a task includes selecting a first set of micro-operations that are ready for execution during a particular clock cycle. The method also includes selecting a second set of micro-operations that are ready for execution during the particular clock cycle. The method additionally includes replacing one or more complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, at least partially due to the number of complex micro-operations included in the first set of micro-operations exceeding a set of complex resources capable of executing the complex micro-operations.

[0012] In one example, the method further includes feeding the first set of micro-operations through a set of issue ports to a set of complex resources and a set of simple resources when replacing one or more complex micro-operations with one or more simple micro-operations in the first set of micro-operations. In one example, the set of complex resources can include one or more binary multipliers and / or one or more FPUs. Additionally or alternatively, the set of simple resources can include one or more ALUs.

[0013] In one example, the method includes selecting the first set of micro-operations from a scheduler queue, at least partially due to the first set of micro-operations being older than all other micro-operations in the scheduler queue during a particular clock cycle. Additionally or alternatively, the method can include selecting one or more simple micro-operations from the scheduler queue for inclusion in the second set of micro-operations, at least partially due to the second set of micro-operations being older than all other simple micro-operations in the scheduler queue during the particular clock cycle.

[0014] In one example, the method includes identifying the number of complex micro-operations by counting the number of complex micro-operations included in a first set of micro-operations during subsequent clock cycles. In this example, the method calculates the difference between the number of complex micro-operations included in the first set of micro-operations and the number of complex resources capable of executing complex micro-operations within the processor, and then determines that one or more complex micro-operations included in the first set of micro-operations are sufficient to satisfy the difference between the number of complex micro-operations and the number of complex resources included in the first set of micro-operations, and that one or more complex micro-operations included in the first set of micro-operations are newer than all other complex micro-operations included in the first set of micro-operations, and further includes replacing one or more complex micro-operations included in the first set of micro-operations.

[0015] In one example, the first set of micro-operations can include a combination of complex micro-operations and simple micro-operations. In this example, the second set of micro-operations consists of only simple micro-operations.

[0016] In one example, the first set of micro-operations includes a number of micro-operations that matches the total number of complex and simple resources within the processor. In this example, the second set of micro-operations includes a number of simple micro-operations that does not exceed the difference between the number of micro-operations and the total number of complex resources within the processor.

[0017] In one example, each of the complex micro-operations may require multiple clock cycles for execution by the processor. In this example, each of the simple micro-operations may require a single clock cycle for execution by the processor.

[0018] In one example, each of the complex micro-operations can include at least one of a multiplication operation and / or a division operation. In this example, each of the simple micro-operations can include at least one of an addition operation, a subtraction operation, and / or a comparison operation.

[0019] In one example, the method can include identifying a set of issue ports that lead to a set of complex resources and a set of simple resources. In this example, the method further includes identifying, within the set of issue ports, one or more issue ports that lead to the set of complex resources. Additionally or alternatively, the method can include rearranging the order of the first set of micro-operations such that all complex micro-operations included in the first set of micro-operations are fed to one or more issue ports that lead to the set of complex resources.

[0020] In one example, a processor configured to perform an efficient pick of micro-operations for execution includes a first picker configured to select a first set of micro-operations that are ready for execution during a particular clock cycle. The processor also includes a second picker configured to select a second set of micro-operations that are ready for execution during a particular clock cycle. In this example, the first picker or the second picker is configured to replace one or more complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, at least partially due to the number of complex micro-operations included in the first set of micro-operations exceeding the set of complex resources capable of executing the complex micro-operations.

[0021] In one example, the first picker is further configured to feed the first set of micro-operations to the set of complex resources and the set of simple resources via the set of issue ports when replacing one or more complex micro-operations with one or more simple micro-operations in the first set of micro-operations. In one example, the set of complex resources can include one or more binary multipliers and / or one or more FPUs. Additionally or alternatively, the set of simple resources can include one or more ALUs.

[0022] In one example, the first picker is further configured to select the first set of micro-operations from the scheduler queue at least in part because the first set of micro-operations is older than all other micro-operations in the scheduler queue during a particular clock cycle. Additionally or alternatively, the second picker is further configured to select one or more simple micro-operations from the scheduler queue for inclusion in the second set of micro-operations at least in part because the second set of micro-operations is older than all other simple micro-operations in the scheduler queue during a particular clock cycle.

[0023] In one example, the first or second picker is further configured to identify the number of complex micro-operations by counting the number of complex micro-operations included in the first set of micro-operations during subsequent clock cycles. In this example, the first or second picker calculates the difference between the number of complex micro-operations included in the first set of micro-operations and the number of complex resources capable of executing complex micro-operations within the processor, and then determines that one or more complex micro-operations included in the first set of micro-operations are sufficient to satisfy the difference between the number of complex micro-operations and the number of complex resources included in the first set of micro-operations, and that one or more complex micro-operations included in the first set of micro-operations are newer than all other complex micro-operations included in the first set of micro-operations, and is further configured to replace one or more complex micro-operations included in the first set of micro-operations.

[0024] In one example, the first set of micro-operations can include a combination of complex and simple micro-operations. In this example, the second set of micro-operations consists of only simple micro-operations.

[0025] In one example, the first set of micro-operations includes a number of micro-operations that matches the total number of complex and simple resources within the processor. In this example, the second set of micro-operations includes a number of simple micro-operations that does not exceed the difference between the number of micro-operations and the total number of complex resources within the processor.

[0026] In some examples, a computing device that makes an efficient pick of micro-operations for execution includes a processor and a memory device communicatively coupled to the processor. In one example, the processor is configured to select a first set of micro-operations that are ready for execution during a particular clock cycle. In this example, the processor is configured to select a second set of micro-operations that are ready for execution during a particular clock cycle. Also, the processor is configured to replace one or more of the complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, at least in part because the number of complex micro-operations included in the first set of micro-operations exceeds the set of complex resources capable of executing complex micro-operations. In one example, the memory device is configured to store one or more computer-readable instructions, from which the processor can derive the first set of micro-operations and the second set of micro-operations.

[0027] The following provides a detailed description of an exemplary apparatus, system, and / or corresponding embodiments for making an efficient pick of micro-operations for execution with reference to FIGS. 1-7. A detailed description of an exemplary method for making an efficient pick of micro-operations for execution is provided in connection with FIG. 8.

[0028] FIG. 1 shows an exemplary processor 100 that facilitates efficient picking of micro-operations for execution. As shown in FIG. 1, the processor 100 includes and / or represents a scheduler queue 102 that maintains, stores, and / or buffers micro-operations 108(1) to (N). In some examples, the scheduler queue 102 maintains, stores, and / or buffers the micro-operations 108(1) to (N) in a time-based order and / or a first-in, first-out (FIFO) order. Additionally or alternatively, the processor 100 includes and / or represents a picker 104 and / or a picker 106 that is responsible for picking and / or selecting a group of the micro-operations 108(1) to (N) from the scheduler queue 102.

[0029] In some examples, the processor 100 includes and / or represents a set of one or more complex resources 114(1) to (N) and / or a set of one or more simple resources 116(1) to (N). In one example, the complex resources 114(1) to (N) can perform, compute, and / or execute complex micro-operations picked and / or selected by the picker 104. In this example, the simple resources 116(1) to (N) can perform, compute, and / or execute simple micro-operations picked and / or selected by the picker 104 and / or the picker 106.

[0030] In some examples, processor 100 can include and / or represent any type or form of hardware implementation device capable of interpreting and / or executing computer-readable instructions. In one example, processor 100 includes and / or represents one or more semiconductor devices implemented and / or deployed as part of a computing system. Examples of processor 100 include, but are not limited to, a central processing unit (CPU), a microprocessor, a microcontroller, a field-programmable gate array (FPGA) implementing a softcore processor, an application-specific integrated circuit (ASIC), a system on a chip (SoC), one or more portions of these, one or more variations or combinations of these, and / or any other suitable processor.

[0031] Processor 100 can implement any of a variety of different architectures and / or microarchitectures, and / or can be configured using them. For example, processor 100 can be implemented and / or configured as a reduced instruction set computer (RISC) architecture. In another example, processor 100 can be implemented and / or configured as a complex instruction set computer (CISC) architecture. Additional examples of such architectures and / or microarchitectures include, but are not limited to, 16-bit computer architecture, 32-bit computer architecture, 64-bit computer architecture, x86 computer architecture, advanced RISC machine (ARM) architecture, microprocessor without interlocked pipelined stage (MIPS) architecture, scalable processor architecture (SPARC), load store architecture, one or more portions of these, one or more combinations or variations of these, and / or any other suitable architecture or microarchitecture.

[0032] In some examples, the scheduler queue 102 can include and / or represent any type or form of queue and / or buffer implemented and / or configured within the processor 100. In one example, the scheduler queue 102 can include and / or represent a data structure and / or abstract data type. In another example, the scheduler queue 102 can include and / or represent the feature of the CPU that maintains, presents, and / or provides the micro-operations 108(1)-(N) picked for issuance for the complex resources 114(1)-(N) and / or the simple resources 116(1)-(N). Additionally or alternatively, the scheduler queue 102 can include and / or represent hardware, software, and / or firmware implemented as part of the processor 100.

[0033] In some examples, the picker 104 and / or the picker 106 can include and / or represent any type or form of process, module, and / or unit that picks and / or selects a group of the micro-operations 108(1)-(N) for execution by the complex resources 114(1)-(N) and / or the simple resources 116(1)-(N). In one example, the picker 104 is configured to pick and / or select a specific number of the micro-operations 108(1)-(N) as the pick 110, and the picker 106 is configured to pick and / or select a specific number of the micro-operations 108(1)-(N) as the pick 112. For example, the picker 104 is configured to pick and / or select simple and / or complex operations, and the picker 106 is configured to pick and / or select only simple micro-operations. In certain embodiments, the picker 104 and / or the picker 106 can include and / or represent hardware, software, and / or firmware implemented as part of the processor 100.

[0034] In some examples, pick 110 may include and / or represent a greater or lesser number of micro-operations than pick 112. Additionally or alternatively, pick 110 may include and / or represent any combination of complex and simple micro-operations, while pick 112 may include and / or represent only simple micro-operations. In certain scenarios, picks 110 and 112 may share certain overlapping micro-operations that are common to each other.

[0035] In some examples, pick 110 includes a number of micro-operations that matches the total number of complex and simple resources within processor 100. In such examples, pick 112 includes a number of simple micro-operations that does not exceed the difference between the number of micro-operations included in pick 110 and the number of complex resources within processor 100. For example, in some pick cycles, pick 110 can include the maximum number of complex micro-operations allowed by processor 100, with the remainder being simple micro-operations. However, in other pick cycles, pick 110 can include all simple micro-operations without any complex micro-operations.

[0036] In some examples, complex resources 114(1)-(N) and / or simple resources 116(1)-(N) can include and / or represent any type or form of digital circuitry that performs micro-operations on numbers, data, and / or values. In one example, complex resources 114(1)-(N) can include and / or represent a binary multiplier and / or FPU that can perform complex micro-operations (such as multiplication and / or division operations). Additionally or alternatively, each of complex resources 114(1)-(N) can include and / or represent any other type of resource (such as a complex ALU) that can perform such complex micro-operations. In this example, simple resources 116(1)-(N) can include and / or represent an ALU that can perform simple micro-operations (such as addition, subtraction, and / or comparison operations).

[0037] In some examples, the micro-operations 108(1) to (N) can include and / or represent any type or form of code and / or instructions that are implemented and / or executed by the complex resources 114(1) to (N) and / or the simple resources 116(1) to (N) of the processor 100. In one example, the micro-operations 108(1) to (N) can include and / or represent one or more complex and / or special micro-operations (such as multiplication and / or division operations, etc.) that require multiple clock cycles for execution by the complex resources 114(1) to (N). Additionally or alternatively, the micro-operations 108(1) to (N) can include and / or represent one or more simple and / or general micro-operations (such as addition, subtraction, and / or comparison operations, etc.) that require only a single clock cycle for execution by the simple resources 116(1) to (N). Also, the micro-operations 108(1) to (N) can involve and / or represent updates to registers, data transfers to / from registers or between registers, and / or data transfers to / from registers to / from an interface (e.g., a bus).

[0038] In some examples, the processor 100 can include and / or incorporate one or more additional components that are not explicitly shown and / or represented in FIG. 1. Examples of such additional components include, but are not limited to, registers, memory devices, circuits, transistors, resistors, capacitors, diodes, connections, traces, buses, semiconductor (e.g., silicon) devices and / or structures, combinations or variations of one or more of these, and / or any other suitable components that enable the processor 100 to efficiently pick micro-operations for execution. Additionally or alternatively, the processor 100 can exclude and / or omit one or more of the components, devices, and / or features shown and / or labeled in FIG. 1. For example, in a pipeline embodiment characterized by only a single complex resource, the processor 100 can exclude and / or omit the complex resource 114(N).

[0039] In some examples, picker 104 selects and / or picks a specific number of micro-operations 108(1) to (N) from scheduler queue 102 for inclusion in pick 110. In such examples, the micro-operations selected and / or picked by picker 104 are ready for execution by one or more of complex resources 114(1) to (N) and / or simple resources 116(1) to (N). In one example, picker 104 can select and / or pick the oldest N micro-operations of any type and / or kind.

[0040] In some examples, picker 106 selects and / or picks a specific number of micro-operations 108(1) to (N) from scheduler queue 102 for inclusion in pick 112. In such examples, the micro-operations selected and / or picked by picker 106 are ready for execution by one or more of simple resources 116(1) to (N). In one example, picker 106 can select and / or pick the oldest M micro-operations 108(1) to (N) that can be executed and / or performed by simple resources 116(1) to (N).

[0041] In some examples, the phrase “ready to execute,” when used in this context, can indicate and / or imply that those micro-operations have no dependencies and / or contingencies that could potentially change the state of one or more variables of such micro-operations. In one example, a micro-operation can be considered and / or determined to be ready to execute if the state of its variables is not due to changes prior to the execution of the micro-operation. For example, a multiplication operation that is ready to execute includes and / or represents one or more variables in an appropriate and / or proper state for execution. In other words, the variables included and / or represented in the multiplication operation are not changed (e.g., by another operation) prior to the execution of the multiplication operation. Stated another way, if a particular micro-operation is ready to execute, the variables included and / or represented in that particular micro-operation are not acted upon and / or changed by any other micro-operation until after the execution of that particular micro-operation.

[0042] In some examples, picker 104 and / or another component of processor 100 can count and / or identify the number of complex micro-operations included in pick 110. In such examples, picker 104 and / or another component of processor 100 can determine that the number of complex micro-operations included in pick 110 exceeds the number of complex resources 114(1) to (N) that can execute complex micro-operations within processor 100. In response to this determination, picker 104 and / or another component of processor 100 can replace one or more of the complex micro-operations included in pick 110 with one or more simple micro-operations included in pick 112. In other words, picker 104 and / or another component of processor 100 can substitute one or more simple micro-operations included in pick 112 for one or more of the complex micro-operations included in pick 110. In one example, when replacing one or more complex micro-operations with one or more simple micro-operations in pick 110, picker 104 can feed, push, and / or issue pick 110 through the pipeline of processor 100 for execution by complex resources 114(1) to (N) and / or simple resources 116(1) to (N).

[0043] Figure 2 shows an exemplary embodiment 200 of a processor that facilitates an efficient pick of micro-operations for execution. As shown in Figure 2, the scheduler queue 102 maintains, stores, and / or loads various simple and complex micro-operations in a chronological order array. In some examples, the left side of the scheduler queue 102 in Figure 2 corresponds to and / or represents the oldest of the micro-operations placed in the queue, and the right side of the scheduler queue 102 in Figure 2 corresponds to and / or represents the newest of the micro-operations placed in the queue. For example, as shown in Figure 2, the scheduler queue 102 includes and / or represents complex micro-operations 208(1) and 208(2), and simple micro-operations 210(1), 210(2), 210(3), 210(4), and 210(5). In this example, the complex micro-operation 208(1) is the oldest within the scheduler queue 102, and the simple micro-operation 210(5) is the newest within the scheduler queue 102.

[0044] In some examples, the scheduler queue 102 in Figure 2 can include and / or represent various other micro-operations that are newer than the simple micro-operation 210(5) but are omitted and / or excluded from Figure 2 for the sake of brevity and / or clarity. Additionally or alternatively, the scheduler queue 102 in Figure 2 can include and / or represent various other micro-operations that are not yet ready for execution and are thus omitted and / or excluded from Figure 2 for the sake of brevity and / or clarity.

[0045] In some examples, an exemplary embodiment 200 of a processor includes and / or represents one complex resource and five simple resources (not necessarily shown or labeled in FIG. 2). In one example, picker 104 first selects complex micro-operations 208(1)-(2) and simple micro-operations 210(1)-(4) as pick 110 because they are the oldest micro-operations ready for execution in scheduler queue 102. In this example, picker 106 first selects simple micro-operations 210(1)-(5) as pick 112 because they are the oldest simple micro-operations ready for execution in scheduler queue 102. However, since the exemplary embodiment 200 of the processor includes only one complex resource, picker 104 and / or another component of processor 100 can drop and / or remove complex micro-operation 208(2) from pick 110. As a result, complex micro-operation 208(2) remains in scheduler queue 102 and / or is returned to scheduler queue 102 when and / or after pick 110 is issued for execution by the complex resource and the simple resources. Thus, complex micro-operation 208(2) remains available for selection by picker 104 in a subsequent pick.

[0046] In some examples, the picker 104 and / or another component of the processor 100 fill the slots emptied by the complex micro-operation 208(2) in the pick 110 with simple micro-operations 210(5). By doing so, the picker 104 and / or another component of the processor 100 effectively replace the complex micro-operation 208(2) with the simple micro-operation 210(5) in the pick 110. Next, the picker 104 issues the pick 110 for execution against complex and simple resources. Upon the issuance and / or execution of the pick 110, the complex micro-operation 208(1) and / or the simple micro-operations 210(1)-(5) can be broadcast to their dependents and / or associated registers to update the corresponding data and / or values within the processor 100.

[0047] As a specific example, an exemplary embodiment 200 of a processor can involve and / or represent a picking scheme that lasts for two clock cycles and / or spans two clock cycles. In this example, a simple micro-operation can have a latency of one clock cycle for execution, and a complex micro-operation can have a latency of N clock cycles for execution. During the first clock cycle of the two-cycle picking scheme, picker 104 can first select complex micro-operations 208(1)-(2) and simple micro-operations 210(1)-(4) as pick 110, and picker 106 can first select simple micro-operations 210(1)-(5) as pick 112. In the next clock cycle of the two-cycle picking scheme, picker 104 can drop complex micro-operation 208(2) from pick 110 and / or replace it with simple micro-operation 210(5) from pick 112. Next, picker 104 can issue pick 110 to complex and simple resources for execution. Upon issuance and / or execution of pick 110, simple micro-operations 210(1)-(4) can be immediately broadcast to their dependent and / or associated registers to update the corresponding data and / or values within processor 100. However, since simple micro-operation 210(5) replaced complex micro-operation 208(2) in pick 110, simple micro-operation 210(5) can be broadcast to its dependent and / or associated registers to update the corresponding data and / or values within processor 100 during the next clock cycle. Additionally, since complex micro-operation 208(1) has a latency of N clock cycles, complex micro-operation 208(1) can be broadcast to its dependent and / or associated registers to update the corresponding data and / or values within processor 100 after N clock cycles.

[0048] FIG. 3 shows an exemplary pipeline 300 of the processor whose embodiment is shown in FIG. 2. As shown in FIG. 3, the pipeline 300 of the processor can include and / or be accompanied by modifying pick 110 based at least in part on pick 112 via substitution operation 312. In some examples, upon completion of this modification, pick 110 can issue to complex resource 114(1) and simple resources 116(1)-(5) of the processor via ports 302(1), 302(2), 302(3), 302(4), 302(5), 302(6), respectively. Thus, in this embodiment, the pipeline 300 can include and / or represent six issue slots. In these examples, upon receiving the micro-operation issued by pick 110, complex resource 114(1) and simple resources 116(1)-(5) can perform, calculate, and / or execute the micro-operation.

[0049] In one example, the replacement operation 312 can involve and / or represent the picker 104 replacing a complex micro-operation 208(2) within the pick 110 with a simple micro-operation 210(5). Thus, prior to the replacement operation 312, the pick 110 includes and / or represents the first version and / or configuration of the pick 110. Conversely, after the replacement operation 312, the pick 110 includes and / or represents an updated and / or modified version of the pick 110. In this example, upon completion of the replacement operation 312, the picker 104 can direct, dispatch, and / or issue the pick 110 by supplying a complex micro-operation 208(1) to a complex resource 114(1) via port 302(1), a simple micro-operation 210(1) to a simple resource 116(1) via port 302(2), a simple micro-operation 210(2) to a simple resource 116(2) via port 302(3), a simple micro-operation 210(3) to a simple resource 116(3) via port 302(4), a simple micro-operation 210(4) to a simple resource 116(4) via port 302(5), and / or a simple micro-operation 210(5) to a simple resource 116(5) via port 302(6).

[0050] FIG. 4 shows an exemplary embodiment 400 of a processor that facilitates an efficient pick of micro-operations for execution. As shown in FIG. 4, a scheduler queue 102 maintains, stores, and / or loads various simple and complex micro-operations in a chronological order. In some examples, as shown in FIG. 4, the scheduler queue 102 includes and / or represents addition operations 410(1) and 410(2), multiplication operations 402(1) and 402(2), subtraction operations 412(1) and 412(2), division operation 404(1), and / or comparison operation 408(1). In one example, the scheduler queue 102 is ordered by age such that the addition operation 410(1) is the oldest micro-operation in the scheduler queue 102 and the subtraction operation 412(2) is the newest micro-operation in the scheduler queue 102. Additionally, the scheduler queue 102 of FIG. 4 includes and / or represents various other micro-operations that are newer than the simple micro-operation 210(5) but are omitted and / or excluded from FIG. 4 for brevity and / or clarity.

[0051] In some examples, an exemplary embodiment 400 of a processor includes and / or represents two complex resources and four simple resources (not necessarily shown or labeled in FIG. 4). In one example, picker 104 first selects addition operations 410(1)-(2), multiplication operations 402(1)-(2), subtraction operation 412(1), and / or division operation 404(1) as pick 430. In this example, picker 106 first selects addition operations 410(1)-(2), subtraction operations 412(1)-(2), and comparison operation 408(1) as pick 432. However, since the exemplary embodiment 400 of the processor includes only two complex resources, picker 104 and / or another component of processor 100 can drop and / or remove division operation 404(1) from pick 430. As a result, division operation 404(1) remains in scheduler queue 102 and / or is returned to scheduler queue 102 when and / or after pick 430 is issued for execution by complex and simple resources. Thus, division operation 404(1) remains available for selection by picker 104 in a subsequent pick.

[0052] In some examples, picker 104 and / or another component of processor 100 fills the slot vacated by division operation 404(1) in pick 430 with comparison operation 408(1). By doing so, picker 104 and / or another component of processor 100 can effectively replace division operation 404(1) with comparison operation 408(1) in pick 430. Next, picker 104 issues pick 430 to complex and simple resources for execution. Upon issuance and / or execution of pick 430, addition operations 410(1)-(2), multiplication operations 402(1)-(2), subtraction operation 412(1), and / or comparison operation 408(1) are broadcast to their dependent and / or associated registers to update corresponding data and / or values within processor 100.

[0053] As a specific example, an exemplary embodiment 400 of a processor can involve and / or represent a picking scheme that lasts for 2 clock cycles and / or spans 2 clock cycles. In this example, each of the addition operations 410(1)-(2), subtraction operation 412(1), and comparison operation 408(1) can have a latency of 1 clock cycle, and each of the multiplication operations 402(1)-(2) can have a latency of N clock cycles. During the first clock cycle of the 2-cycle picking scheme, picker 104 first selects the addition operations 410(1)-(2), multiplication operations 402(1)-(2), subtraction operation 412(1), and / or division operation 404(1) as pick 430, and picker 106 first selects the addition operations 410(1)-(2), subtraction operations 412(1)-(2), and comparison operation 408(1) as pick 432. In the next clock cycle of the 2-cycle picking scheme, picker 104 drops the division operation 404(1) from pick 430 and / or replaces it with the comparison operation 408(1) from pick 432. Next, picker 104 issues pick 430 to complex and simple resources for execution. Upon issuance and / or execution of pick 430, the addition operations 410(1)-(2) and / or subtraction operation 412(1) can broadcast directly to their dependent and / or associated registers to update the corresponding data and / or values within processor 100. However, since the comparison operation 408(1) replaced the division operation 404(1) in pick 430, the comparison operation 408(1) broadcasts to its dependent and / or related registers to update the corresponding data and / or values within processor 100 during the next clock cycle. Additionally, since each of the multiplication operations 402(1)-(2) has a latency of N clock cycles, the multiplication operations 402(1)-(2) broadcast to their dependent and / or related registers to update the corresponding data and / or values within processor 100 after N clock cycles.

[0054] FIG. 5 shows an exemplary pipeline 500 of the processor whose embodiment is shown in FIG. 4. As shown in FIG. 5, the pipeline 500 of the processor can include and / or be accompanied by modifying the pick 430 via a swizzle operation 502. In some examples, when this modification is complete, the pick 430 is issued to the multipliers 514(1) and 514(2) of the processor, and / or the ALUs 516(1), 516(2), 516(3), 516(4) via ports 504(1), 504(2), 504(3), 504(4), 504(5), 504(6) respectively. Thus, in this embodiment, the pipeline 500 includes and / or represents six issue slots. In these examples, upon receiving the micro-operations issued by the pick 430, the multipliers 514(1)-(2) and / or the ALUs 516(1)-(4) perform, calculate and / or execute those micro-operations.

[0055] In one example, the swizzle operation 502 involves and / or represents the picker 104 rearranging and / or reordering the addition operations 410(1)-(2), multiplication operations 402(1)-(2), subtraction operation 412(1), and / or comparison operation 408(1) at pick 110. For example, each micro-operation included in pick 430 corresponds to and / or is assigned to one of six issue slots. In this example, as part of the swizzle operation 502, the picker 104 rearranges and / or reorders pick 430 with respect to the six issue slots such that each of the multiplication operations 402(1)-(2) is aligned with and / or directed towards the multipliers 514(1)-(2). Additionally or alternatively, the picker 104 rearranges and / or reorders pick 430 with respect to the six issue slots such that each of the addition operation 410(1), subtraction operation 412(1), addition operation 410(2), and / or comparison operation 408(1) is aligned with and / or directed towards the ALUs 516(1)-(4). After completion of the swizzle operation 502, the picker 104 issues and / or supplies the multiplication operations 402(1)-(2) to the multipliers 514(1)-(2) via ports 504(1)-(2), and / or the picker 104 issues and / or supplies the addition operation 410(1), subtraction operation 412(1), addition operation 410(2), and / or comparison operation 408(1) to the ALUs 516(1)-(4) via ports 504(3)-(6).

[0056] FIG. 6 shows an exemplary embodiment 600 of a processor that facilitates efficient picking of micro-operations for execution. As shown in FIG. 6, the exemplary embodiment 600 includes and / or represents various micro-operations that are maintained, stored, and / or loaded in a scheduler queue in a chronological order. For example, such micro-operations include and / or represent Mul1, Op1, Mul2, Op2, Op3, Op4, Op5, Mul3, Op6, Op7, Op8, and / or Op9 in the scheduler queue. In this example, Mul1 is the oldest micro-operation in the scheduler queue and Op9 is the newest micro-operation in the scheduler queue.

[0057] In some examples, various micro-operations (e.g., Mul1, Op1, Op2, Mul3, Op6, and / or Op9) loaded into the scheduler queue are ready for execution, while various other micro-operations (e.g., Mul2, Op3, Op4, Op5, Op7, and / or Op8) are not yet ready for execution. As a result of not being ready for execution, those other micro-operations are not available for selection by picker 104 and / or picker 106 until a subsequent clock cycle. In one example, picker 104 selects Mul1, Op1, Op2, and / or Mul3 as pick 610. In this example, picker 106 selects Op1, Op2, Op6, and / or Op9 as pick 612.

[0058] In embodiment 600, the processor pipeline includes and / or represents an execution unit having one complex resource (e.g., a multiplier) and / or three simple resources (e.g., ALUs). At pick 610, Mul1 and / or Mul3 constitute and / or represent a multiplication operation, and Op1 and Op2 constitute and / or represent simple micro-operations (e.g., addition, subtraction, and / or comparison operations). In contrast, all micro-operations included in pick 612 constitute and / or represent simple micro-operations.

[0059] In some examples, since the processor's pipeline contains only one complex resource, picker 104 and / or another component of the processor counts the number of multiplication operations within pick 610 and then determines that pick 610 contains one multiplication operation that exceeds the number of complex resources. As a result, picker 104 and / or other components of the processor can perform replacement operation 622 by dropping and / or removing Mul3 from pick 610 and then adding Op6 from pick 612 to pick 610. By doing so, picker 104 and / or other components of the processor can ensure that all issue slots available for execution within the processor's pipeline are utilized, thereby improving the efficiency and / or performance of the processor. When replacement operation 622 is complete, the final pick consisting of Mul1, Op1, Op2, and / or Op6 is supplied to the complex and / or simple resources for execution as issue 628.

[0060] In some examples, the picking scheme shown in embodiment 600 involves and / or represents four micro-operations picked per clock cycle and / or four micro-operations issued per clock cycle. In one example, the picking scheme shown in embodiment 600 enables and / or facilitates selecting one multiplication operation per clock cycle. In this example, the selection of picks 610 and 612, as well as the performance of replacement operation 622, collectively continue and / or consume two clock cycles. Continuing with this example, the introduction of replacement operation 622 uses and / or utilizes the 1 / 2 cycle that was present and / or available during the two clock cycles of the picking scheme.

[0061] FIG. 7 shows an exemplary embodiment 700 with a computing device 702. As shown in FIG. 7, the computing device 702 includes and / or represents a processor 100 communicatively coupled to a memory device 704. In some examples, the memory device 704 maintains and / or stores one or more computer-readable instructions 710. In such examples, the processor 100 of the computing device 702 derives specific micro-operations (including any of those described above in connection with FIGS. 1-6) from the computer-readable instructions 710 within the memory device 704. For example, the processor 100 of the computing device 702 can decode the computer-readable instructions 710 to create and / or generate specific micro-operations to be later implemented, computed, and / or executed by complex resources 114(1)-(N) and / or simple resources 116(1)-(N).

[0062] In some examples, the computing device 702 can include and / or represent any type or form of computer capable of performing computing tasks and / or communicating with other computers. Examples of the computing device 702 include, but are not limited to, client devices, laptops, tablets, desktops, servers, mobile phones, personal digital assistants (PDAs), multimedia players, embedded systems, wearable devices, game consoles, routers, switches, hubs, modems, bridges, repeaters, gateways (such as broadband network gateways (BNGs), etc.), multiplexers, network adapters, network interfaces, line cards, one or more portions of these, one or more variations or combinations of these, and / or any other suitable computing device.

[0063] In some examples, the memory device 704 includes and / or represents any type or form of volatile or non-volatile memory device or medium capable of storing data and / or computer-readable instructions. Examples of the memory device 704 include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), flash memory, Hard Disk Drive (HDD), Solid-State Drive (SSD), optical disk drive, cache, one or more variations or combinations thereof, and / or any other suitable memory device.

[0064] FIG. 8 is a flowchart of an exemplary method 800 for efficiently picking micro-operations for execution. In one example, the steps shown in FIG. 8 can be performed by one or more components of a processor incorporated in a computing device. Additionally or alternatively, the steps shown in FIG. 8 can incorporate and / or involve various sub-steps and / or variations consistent with the description provided above in connection with FIGS. 1-7.

[0065] As shown in FIG. 8, the method 800 includes and / or involves a step (810) of selecting a first set of micro-operations that are ready for execution during a particular clock cycle. Step 810 can be performed in various ways including any of the techniques described above in connection with FIGS. 1-7. For example, a main picker included in the processor selects a first set of micro-operations that are ready for execution by a set of complex resources or a set of simple resources during a particular clock cycle from a scheduler queue included in the processor.

[0066] Also, method 800 includes a step (820) of selecting a second set of micro-operations that are ready for execution during a particular clock cycle. Step 820 can be implemented in various ways including any of the techniques described above in connection with FIGS. 1-7. For example, an alternative picker included in the processor selects a second set of micro-operations from the scheduler queue that are ready for execution by a set of simple resources during a particular clock cycle. In one example, the second set of micro-operations is smaller than the first set of micro-operations (e.g., includes fewer micro-operations).

[0067] Method 800 further includes a step (830) of replacing one or more complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, at least in part because the number of complex micro-operations included in the first set of micro-operations exceeds a set of complex resources capable of executing the complex micro-operations. Step 830 can be implemented in various ways including any of the techniques described above in connection with FIGS. 1-7. For example, the main picker replaces one or more complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, at least in part because the number of complex micro-operations included in the first set of micro-operations exceeds a set of complex resources capable of executing the complex micro-operations.

[0068] As described above in connection with FIGS. 1-8, an exemplary processor includes a main picker that can select any type of ready micro-operation (including both single-cycle operations and multi-cycle operations) queued by a scheduler, and / or an alternative picker that can select only the ready single-cycle operations queued in the scheduler. In some examples, the execution units of the processor include M special resources (e.g., multipliers and / or FPUs) that can perform special micro-operations (e.g., multiplication, division, etc.) that require multiple clock cycles, and general resources (e.g., a general ALU) that can perform general micro-operations (e.g., addition, subtraction, comparison, etc.) that require only a single clock cycle. In such examples, the main picker has the ability to select N micro-operations (M < N) per clock cycle, and the alternative picker has the ability to select N - M micro-operations per clock cycle. In one example, the main picker selects the N oldest ready micro-operations of any type during a particular clock cycle, and the alternative picker selects the N - M oldest ready general single-cycle micro-operations during that clock cycle. In the next clock cycle, the main picker drops all special operations that exceed M from its pick and replaces them with one or more of the general micro-operations selected by the alternative picker. Next, the main picker and / or another component of the processor supplies, pushes, and / or issues the mix resulting from the special and general micro-operations found when finally picking down the pipeline through a particular issue port to the special and general resources. Next, the special and general resources eventually execute those operations (e.g., after two or three clock cycles) to perform one or more computing tasks.

[0069] The above disclosure has described various embodiments using specific block diagrams, flowcharts, and examples. However, each block diagram component, flowchart step, operation, and / or component described and / or illustrated herein can be implemented individually and / or collectively using a wide range of hardware, software, or firmware (or any combination thereof) configurations. Additionally, since many other architectures can be implemented to achieve the same functionality, any disclosure of components included within other components should be considered to be exemplary in nature.

[0070] The apparatuses, systems, and methods described herein can employ any number of software, firmware, and / or hardware configurations. For example, one or more of the exemplary embodiments and / or implementations disclosed herein can be encoded as a computer program (also referred to as computer software, software application, computer-readable instructions, and / or computer control logic) on a computer-readable medium. The term "computer-readable medium" generally refers to any form of device, carrier, or medium that can store or convey computer-readable instructions. Examples of computer-readable media include, but are not limited to, transmission-type media such as carrier waves, as well as magnetic storage media (e.g., hard disk drives and floppy (registered trademark) disks), optical storage media (e.g., Compact Disk (CD) and Digital Video Disk (DVD)), electronic storage media (e.g., solid state drives and flash media), and / or non-transitory media such as other delivery systems.

[0071] In addition, one or more of the modules, instructions, and / or micro-operations described herein can transform data, physical devices, and / or representations of physical devices from one form to another. Additionally or alternatively, one or more of the modules, instructions, and / or micro-operations described herein can transform a processor, volatile memory, non-volatile memory, and / or any other part of a physical computing device from one form to another by executing on the computing device, storing data on the computing device, and / or interacting with the computing device in another way.

[0072] The process parameters and sequences of steps described and / or illustrated herein are provided by way of example only and can be varied as desired. For example, although the steps illustrated and / or described herein are shown or described in a particular order, these steps need not necessarily be performed in the order illustrated or described. Also, the various exemplary methods described and / or illustrated herein can include omitting one or more of the steps described or illustrated herein or including additional steps in addition to those disclosed.

[0073] The foregoing description has been provided to enable those skilled in the art to best utilize the various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the present disclosure. The embodiments disclosed herein should be considered exemplary in all respects and not restrictive. The appended claims and their equivalents should be referred to in determining the scope of the present disclosure.

[0074] Unless otherwise indicated, the terms "connected to" and "coupled to" (and their derivatives), as used in this specification and the claims, should be construed to permit both direct and indirect (i.e., via other elements or components) connections. Additionally, the term "a" or "an", as used in this specification and the claims, should be construed to mean "at least one of". Finally, for ease of use, the terms "including" and "having" (and their derivatives), as used in this specification and the claims, are interchangeable with the word "comprising" and have the same meaning.

Claims

1. A method comprising: selecting a first set of micro-operations that are ready for execution during a particular clock cycle; selecting a second set of micro-operations that are ready for execution during the particular clock cycle; replacing one or more complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, at least partially due to the number of complex micro-operations included in the first set of micro-operations exceeding a set of complex resources capable of executing the complex micro-operations; A method.

2. Including supplying the first set of micro-operations to the set of complex resources and the set of simple resources via a set of issue ports when replacing the one or more complex micro-operations with the one or more simple micro-operations in the first set of micro-operations; The method of Claim 1.

3. The set of complex resources includes at least one of one or more binary multipliers or one or more floating point units; The set of simple resources includes one or more arithmetic logic units; The method of Claim 2.

4. Selecting the first set of micro-operations includes selecting the first set of micro-operations from a scheduler queue, at least partially due to the first set of micro-operations being older than all other micro-operations in the scheduler queue during the particular clock cycle; The method of Claim 1.

5. Selecting the second set of micro-operations includes selecting the one or more simple micro-operations from a scheduler queue for inclusion in the second set of micro-operations, at least partially due to the second set of micro-operations being older than all other simple micro-operations in the scheduler queue during the particular clock cycle; The method of Claim 1.

6. Identifying the number of complex micro-operations by counting the number of complex micro-operations included in the first set of micro-operations during subsequent clock cycles; Replacing the one or more complex micro-operations included in the first set of micro-operations with the one or more simple micro-operations is calculating the difference between the number of complex micro-operations included in the first set of micro-operations and the number of complex resources capable of executing the complex micro-operations in the processor; determining that the one or more complex micro-operations included in the first set of micro-operations are sufficient to satisfy the difference between the number of complex micro-operations included in the first set of micro-operations and the number of complex resources, and are newer than all other complex micro-operations included in the first set of micro-operations, The method of claim 1.

7. The first set of micro-operations includes a combination of complex micro-operations and simple micro-operations, The second set of micro-operations includes only simple micro-operations, The method of claim 1.

8. The first set of micro-operations includes a number of micro-operations that matches the total number of complex and simple resources in the processor, The second set of micro-operations includes a number of simple micro-operations that does not exceed the difference between the number of micro-operations and the total number of complex resources in the processor, The method of claim 1.

9. Each of the complex micro-operations requires a plurality of clock cycles for execution by the processor, Each of the simple micro-operations requires a single clock cycle for execution by the processor, The method of claim 1.

10. Each of the complex micro-operations includes at least one of a multiplication operation or a division operation, Each of the simple micro-operations includes at least one of an addition operation, a subtraction operation, or a comparison operation, The method of claim 1.

11. identifying a set of issue ports connected to the set of complex resources and the set of simple resources; identifying one or more issue ports connected to the set of complex resources within the set of issue ports; rearranging the order of the first set of micro-operations such that all of the complex micro-operations included in the first set of micro-operations are supplied to the one or more issue ports connected to the set of complex resources, The method of claim 1.

12. A processor, a first picker configured to select a first set of micro-operations that are ready for execution during a particular clock cycle, a second picker configured to select a second set of micro-operations that are ready for execution during the particular clock cycle, and the first picker or the second picker is configured to replace one or more complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, at least partially due to the number of complex micro-operations included in the first set of micro-operations exceeding a set of complex resources capable of executing the complex micro-operations. Processor.

13. The first picker is configured to supply the first set of micro-operations to the set of complex resources and the set of simple resources via a set of issue ports when replacing the one or more complex micro-operations with the one or more simple micro-operations. The processor of claim 12.

14. The set of complex resources includes at least one of one or more binary multipliers or one or more floating-point units, The set of simple resources includes one or more arithmetic logic units, The processor of claim 13.

15. The first picker is configured to select the first set of micro-operations from the scheduler queue at least partially due to the first set of micro-operations being older than all other micro-operations in the scheduler queue during the particular clock cycle. The processor of claim 12.

16. The second picker is configured to select the one or more simple micro-operations from the scheduler queue for inclusion in the second set of micro-operations at least partially due to the second set of micro-operations being older than all other simple micro-operations in the scheduler queue during the particular clock cycle. The processor of claim 12.

17. The first picker or the second picker is Identifying the number of complex micro-operations by counting the number of complex micro-operations included in the first set of micro-operations during subsequent clock cycles; Replacing the one or more complex micro-operations included in the first set of micro-operations with the one or more simple micro-operations; Replacing the one or more complex micro-operations included in the first set of micro-operations with the one or more simple micro-operations; Calculating the difference between the number of complex micro-operations included in the first set of micro-operations and the number of complex resources capable of executing the complex micro-operations in the processor; Determining that the one or more complex micro-operations included in the first set of micro-operations are sufficient to satisfy the difference between the number of complex micro-operations included in the first set of micro-operations and the number of complex resources, and are newer than all other complex micro-operations included in the first set of micro-operations; The processor of claim 12.

18. The first set of micro-operations includes a combination of complex micro-operations and simple micro-operations; The second set of micro-operations includes only simple micro-operations; The processor of claim 12.

19. The first set of micro-operations includes a number of micro-operations that matches the total number of complex and simple resources within the processor; The second set of micro-operations includes a number of simple micro-operations that does not exceed the difference between the number of micro-operations and the total number of complex resources within the processor; The processor of claim 12.

20. A computing device, comprising: A processor; A memory device communicatively coupled to the processor and configured to store one or more computer-readable instructions; The processor is configured to: Select a first set of micro-operations that are ready for execution during a particular clock cycle; Select a second set of micro-operations that are ready for execution during the particular clock cycle; Due at least in part to the number of complex micro-operations included in the first set of micro-operations exceeding the set of complex resources capable of performing the complex micro-operations, replacing one or more complex micro-operations included in the first set of micro-operations with one or more simple micro-operations included in the second set of micro-operations, is configured to perform, the processor is capable of deriving the first set of micro-operations and the second set of micro-operations from the one or more computer-readable instructions, computing device.