Execution unit sharing between processing corres in a system-on-chip (SOC) cluster

By sharing idle execution units among processor cores and utilizing the execution engine manager to monitor and activate the execution units of inactive processor cores, the problem of underutilized processor cores is solved, processor efficiency and performance are improved, and more efficient resource utilization is achieved.

CN121844295APending Publication Date: 2026-04-10QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In modern processors, the underutilization of processor cores and their execution units leads to low execution efficiency in real-world applications, and the unavailability of power-efficient instructions results in additional time consumption.

Method used

By sharing idle execution units among active processor cores, the execution engine manager monitors and activates the execution units of inactive processor cores, transfers instruction and result buffer addresses, and replaces load operations in the instruction queue to achieve result forwarding.

Benefits of technology

It improves the efficiency and performance of the processor core, effectively utilizes available power, reduces runtime, achieves peak single-core performance, and supports scalable vector expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844295A_ABST
    Figure CN121844295A_ABST
Patent Text Reader

Abstract

A method of execution unit (EU) sharing between processor cores is described. The method includes encountering structural hazards associated with issued instructions in an instruction queue of a scheduling stage inside the active processor core. The method also includes issuing a request for an idle execution unit of the inactive processor core. The method also includes transferring a transaction containing a source operand of the issued instruction and a word address of a result buffer as a destination operand to an allocated EU of the inactive processor core. The method further includes replacing the published instruction with a load operation in the instruction queue to forward a result of the published instruction from the result buffer based on the word address.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 473,119, filed September 22, 2023, entitled “Exduction Unit Sharing Between Processing Cores in a Cluster of System-on-a-Chip (SOC),” the entire disclosure of which is expressly incorporated herein by reference. Background Technology Technical Field

[0004] Various aspects of this disclosure relate to semiconductor devices, and more specifically to the sharing of execution units among processor cores in a cluster of system-on-a-chip (SoC). Background Technology

[0005] Modern processors are equipped with multiple cores, ranging from efficient ordered execution to superscalar / superscalar architectures. The number of cores in these modern processors has steadily increased from approximately eight (8) processor cores in mobile processors to ninety-six (96) processor cores in server computing platforms. Each processor core contains multiple integer processing units, floating-point processing units, and load-memory units as part of its backend execution engine. During operation, some processor cores are in a constant utilization state while executing real-world applications, and some execution units are in a constant utilization state within the core while executing the code for these real-world applications.

[0006] Real-world application execution suffers from reduced efficiency due to the underutilization of processor cores and their associated execution. Unfortunately, power-efficient instructions for processor core execution are unavailable. Instead, processor core efficiency is achieved through early completion of specified computations and idle states that reach clock and power gating. In practice, the extra time consumed by the active processor core to complete computations is inefficient when unused processor cores and partially used execution engines exist. This inefficiency is expected to be addressed by sharing unused execution engines with active processor cores. Summary of the Invention

[0007] A method of execution unit (EU) sharing between processor cores is described. The method includes encountering a structural hazard associated with an issued instruction in an instruction queue of a dispatch stage within an active processor core. The method also includes issuing a request for an idle execution unit of an inactive processor core. The method also includes transferring a transaction containing a source operand of the issued instruction and a word address of a result buffer as a destination operand to an assigned EU of the inactive processor core. The method also includes replacing the issued instruction with a load operation in the instruction queue to forward a result of the issued instruction from the result buffer based on the word address.

[0008] A method for an execution engine (EE) manager to support processor cores is described. The method includes monitoring a state of execution units (EUs) in a cluster of processor cores. The method also includes receiving a request for an idle execution unit (EU) in the cluster of processor cores, the method further includes transferring a control signal to activate an assigned EU of an inactive processor core. The method also includes transferring an EU acknowledgement and an EU identification (EU ID) to a dispatch stage of a requesting processor core. The method also includes transferring an issued instruction in an instruction queue of the dispatch stage within an active processor core to the assigned EU over an EE network-on-chip (NOC) (EE NOC).

[0009] This has outlined in broad terms the features and technical advantages of the present disclosure so as to provide an overview where a detailed description is taken below. Additional features and advantages of the present disclosure will be described below. Those skilled in the art will appreciate the disclosure readily, realizing additional aspects and advantages thereof upon reading the following detailed description and viewing the accompanying drawings. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure, as claimed. BRIEF DESCRIPTION OF DRAWINGS

[0010] For a more complete understanding of the present disclosure, reference is now made to the following description taken in connection with the accompanying drawings in which:

[0011] Figure 1 Example implementations of a host system on chip (SoC) configured for execution unit sharing between processor cores are illustrated in accordance with various aspects of the present disclosure.

[0012] Figure 2 are further illustrated in accordance with various aspects of the present disclosure. Figure 1a circuit diagram of a system-on-a-chip (SoC) including an execution engine manager to support execution unit sharing operations between processor cores.

[0013] Figure 3 is a timing diagram illustrating execution engine sharing between processor cores in accordance with various aspects of the present disclosure. Figure 2 a block diagram of a system-on-a-chip (SoC) including an execution engine manager to support idle execution unit sharing operations between processor cores.

[0014] Figure 4 is a timing diagram illustrating execution engine sharing between processor cores in accordance with various aspects of the present disclosure.

[0015] Figure 5 is a timing diagram illustrating execution engine sharing between processor cores in accordance with various aspects of the present disclosure. Figure 2 a block diagram of a system-on-a-chip (SoC) including an execution engine manager to support execution unit sharing operations between processor cores.

[0016] Figure 6 is a process flow diagram illustrating a method for execution engine (EE) sharing between processor cores in accordance with various aspects of the present disclosure.

[0017] Figure 7 is a process flow diagram illustrating a method for an execution engine (EE) manager to support processor cores in accordance with various aspects of the present disclosure.

[0018] Figure 8 is a block diagram illustrating an example wireless communication system in which configurations of the present disclosure can be advantageously employed.

[0019] Figure 9 is a block diagram of a design workstation used for circuit, layout, and logic design of a semiconductor component in accordance with one configuration. DETAILED DESCRIPTION

[0020] The detailed description set forth below, in connection with the appended drawings, is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein can be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form in order to avoid obscuring such concepts.

[0021] As described herein, use of the term “and / or” is intended to represent “inclusive or” and use of the term “or” is intended to represent “exclusive or.” As described herein, the term “exemplary” is used herein to mean “serving as an example, instance, or illustration,” and not necessarily “preferred” over other exemplary configurations. As described herein, the term “coupled” is used herein to mean “directly or indirectly connected,” and not necessarily “directly connected,” and is intended to encompass a connection between two or more devices that is capable of permitting the direct or indirect exchange of information between such devices. Additionally, a connection can be permanent or releasable. A connection can be made through a switch. As described herein, the term “proximate” is used herein to mean “adjacent, very near, next to, or near.” As described herein, the term “on” is used herein to mean “directly on” in some configurations and “indirectly on” in other configurations. It is understood that the term “layer” includes a film and is not interpreted to indicate a vertical or horizontal thickness unless otherwise noted. As described, the term “substrate” can refer to a substrate of a diced wafer or can refer to a substrate of an undiced wafer. Similarly, the terms “chip” and “die” can be used interchangeably.

[0022] Modern processors are equipped with multiple cores ranging from efficient in-order execution to super / superscalar architectures. The number of cores in these modern processors has steadily risen from about eight (8) processor cores in mobile processors to ninety-six (96) processor cores in server computing platforms. Each processor core contains multiple integer processing units, floating point processing units, and load store units as part of its back-end execution engine. During operation, some of the processor cores are in constant utilization while executing real-world applications, and some of the execution units are in constant utilization inside the cores while executing the code of these real-world applications.

[0023] Due to underutilization of the processor cores and their associated execution units, execution of real-world applications involves reduced efficiency. Unfortunately, power efficient instructions for processor core execution are not available. Instead, processor core efficiency is achieved by early completion of specified computations and reaching an idle state of clock gating and power gating. In implementation, when there are unutilized processor cores and partially used execution engines, the additional time consumed by the active processor cores to complete the computations is inefficient. It is desirable to address this inefficiency by sharing the unused execution engines with the active processor cores.

[0024] Various aspects of the present disclosure relate to a process for execution unit (EU) sharing between active processor cores. The EU sharing process includes encountering a structural hazard associated with an issued instruction in an instruction queue of a dispatch stage within an active processor core. The EU sharing process also includes issuing a request for an idle EU of an inactive processor core. The EU sharing process also includes transferring a transaction containing a source operand of the issued instruction and a word address of a result buffer as a destination operand to an assigned EU of the inactive processor core. The EU sharing process is complete by replacing the issued instruction with a load operation in the instruction queue to forward a result of the issued instruction from the result buffer based on the word address.

[0025] Various aspects of the present disclosure relate to a process for an execution engine (EE) manager to support EU sharing operations between processor cores. The EE manager process includes monitoring a state of EUs in a cluster of processor cores. The EE manager process also includes receiving a request for an idle EU in the cluster of processor cores. The EE manager process also includes transmitting a control signal to activate an assigned EU of an inactive processor core. The EE manager process transmits an EU acknowledgment and an EU identification (EU ID) to a requesting processor core.

[0026] Figure 1 An example implementation of a host system on chip (SoC) 100 configured for execution unit sharing between processor cores is illustrated in accordance with aspects of the present disclosure. The host SoC 100 includes processing blocks customized for specific functionality, such as a connectivity block 110. The connectivity block 110 can include sixth generation (6G) connectivity, fifth generation (5G) new radio (NR) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, secure digital (SD) connectivity, and the like. ® The connectivity block 110 is coupled to a memory controller 120 that is coupled to a memory 130. The memory controller 120 is configured to manage access to the memory 130 by the various processing units of the host SoC 100. The memory 130 can include a dynamic random access memory (DRAM), a static random access memory (SRAM), a cache memory, a flash memory, a magnetic disk drive, a solid state drive, and the like.

[0027] In this configuration, the host SoC 100 includes various processing units that support multi-threaded operations. For example, the host SoC 100 includes a first processor core 140, a second processor core 150, a third processor core 160, and a fourth processor core 170. The first processor core 140 is coupled to a first execution unit (EU) 141, a second EU 142, and a third EU 143. The second processor core 150 is coupled to a first EU 151, a second EU 152, and a third EU 153. The third processor core 160 is coupled to a first EU 161, a second EU 162, and a third EU 163. The fourth processor core 170 is coupled to a first EU 171, a second EU 172, and a third EU 173. Figure 1In the illustrated configuration, host SoC 100 includes a multi-core central processing unit (CPU) 102, a graphics processor unit (GPU) 104, a digital signal processor (DSP) 106, and a neural processor unit (NPU) / neural signal processor (NSP) 108. Host SoC 100 can also include a sensor processor 114, an image signal processor (ISP) 116, a navigation module 120, which can include a global positioning system, and a memory 118. The multi-core CPU 102, GPU 104, DSP 106, NPU / NSP 108, and multimedia engine 112 support various functions such as video, audio, graphics, gaming, artificial networks, etc. Each processor core of the multi-core CPU 102 can be a reduced instruction set computing (RISC) machine, an advanced RISC machine (ARM), a microprocessor, or some other type of processor. The NPU / NSP 108 can be based on an ARM instruction set.

[0028] The multi-core CPU 102 is equipped with a number of cores that can range from efficient in-order execution to super / superscalar architecture. The number of cores in the multi-core CPU 102 can range from eight (8) processor cores in mobile processor implementations to ninety-six (96) processor cores in server computing platform implementations of the host SoC 100. Each processor core of the multi-core CPU 102 contains a number of integer processing units, floating point processing units, and load store units as part of its back-end execution engine. During operation, some of the processor cores of the multi-core CPU 102 are in constant utilization when executing real-world applications, and some of the execution units of the multi-core CPU 102 are in constant utilization inside the cores when executing code for these real-world applications.

[0029] Due to underutilization of the processor cores and their associated execution units, executing real-world applications using the multi-core CPU 102 involves reduced efficiency. Unfortunately, power-efficient instructions for execution by the processor cores of the multi-core CPU 102 are not available. Instead, efficiency of the processor cores of the multi-core CPU 102 is achieved in terms of early completion of specified computations and reaching an idle state of clock gating and power gating. In implementation, when there are unutilized processor cores and partially used execution engines, the additional time consumed by the active processor cores of the multi-core CPU 102 to complete computations is inefficient. It is desirable to address this inefficiency by sharing the unused execution engines with the active processor cores of the multi-core CPU 102.

[0030] Figure 2 is further illustrated Figure 1circuit diagram of a system-on-a-chip (SoC) that includes an execution engine manager to support idle execution unit (EU) sharing operations between active processor cores. As Figure 2 illustrated, SoC 200 includes access in memory 118 (e.g., level one (LI) cache and / or last level cache (LLC)) over coherent interconnect 230. In various aspects of the disclosure, SoC 200 is configured with a hardware-based execution engine (EE) manager 300 in which idle EUs are shared between active processor cores. In this example, EE manager 300 hardware is monitoring the status of EUs in a cluster of processor cores (e.g., core 0, core 1,..., core N). In operation, EE manager 300 hardware receives a request for an idle EU in the cluster of processor cores of multicore CPU 102. In response, EE manager 300 hardware transmits a control signal to activate an assigned EU of an inactive processor core and also transmits an EU acknowledgment and an EU identification (EU ID) to a dispatch stage of the requesting processor core, e.g., as Figure 3 further illustrated.

[0031] Figure 3 is further illustrated Figure 2 a block diagram of a system-on-a-chip (SoC) that includes an execution engine (EE) manager 300 to support idle execution unit (EU) sharing operations between processor cores. Figure 3 A cluster 301 of processor cores (e.g., core 0 and core 1) of SoC 200 is illustrated, including EE manager 300, in accordance with various aspects of the disclosure. In implementation, processor core cluster 301 can contain any number of cores (e.g., two, four, six, eight, or X) with instruction set architecture (ISA) compatibility. In this example, processor core cluster 301 is shown with two cores (e.g., core 0 and core 1), each containing an extraction stage 302 (302-0, 302-1), a decode stage 310 (310-0, 310-1), a dispatch stage 320 (320-0, 320-1), and an execution engine stage 330 (330-0, 330-1) in their respective pipelines. As Figure 3 illustrated, a high performance core in processor core cluster 301 includes a high performance EU relative to a power efficient core of processor core cluster 301.

[0032] Various aspects of this disclosure relate to architectural solutions utilizing unused (e.g., idle) EUs from inactive cores in a processor core cluster 301. In these aspects of the disclosure, the Execution Engine (EE) Network on-chip (NOC) (EENOC) 340 (340-0, 340-1) and EE Manager 300 are implemented using the processor core cluster 301. In this configuration, the Execution Engine level 330 in each core includes an Integer Processing Unit (IPU), a Floating-Point Unit (FPU), an Arithmetic Logic Unit (ALU), and a Load-Memory Unit (LSU). In this example, the Execution Engine level 330 includes Execution Unit (EU) Identifiers (EUIDs), such as EU ID A, EU ID B, EU ID C, and EU ID D in core 0, and EU ID P, EU ID Q, EU ID R, and EU ID S in core 1.

[0033] like Figure 3 As shown, high-bandwidth, high-speed EE NOCs 340 are added at both ends of the execution engine level 330. In operation, the EE NOC 340 carries instructions, data, power, and clock signals to / from the assigned EUs. In various aspects of this disclosure, the EE manager 300 monitors the processor core cluster 301 and communicates with the EE NOC 340, decoding level 310, scheduling level 320, and execution engine level 330 of each core in the processor core cluster 301.

[0034] In active mode, the EE Manager 300 maintains the active catalog of idle EUs by monitoring the decoded instruction queue in the decoding stage 310 and dependency chain information from the scheduling stage 320, including the current and future utilization status of each EU in the execution engine stage 330. In reactive mode, the EE Manager 300 collects EU status on demand when the core scheduling stage 320 issues a request for additional EUs. As described, EU activity status includes four types: busy, unused, clock-gated, and power-gated. Figure 3 As further illustrated, memory 350 (350-0, 350-1) is provided with result buffer 352 (352-0, 352-1) to store the results calculated by the allocated EU, such as Figure 4 As further described.

[0035] Figure 4 This is an example of a timing diagram 400 illustrating the shared execution engine among processor cores (core A, core B) according to various aspects of this disclosure. For example... Figure 4As shown in timing diagram 400, when a structural hazard is encountered in the instruction queue of the dispatch stage 320 within the active core (core A), the dispatch stage 320 issues a request for a specified type of execution unit (EU) at time 410. In response, the EE manager 300 checks the availability of the requested EU in the directory of the EE manager 300. In operation, if the requested EU is not available, the EE manager 300 transmits a no-assignment acknowledgement (ACK) to the dispatch stage 320. In this example, the requested EU is available in the processor core cluster 301. Accordingly, the EE manager 300 transmits a control signal to the EE NOC 340 at time 420 to provide power to activate the EU. At time 425, the EE NOC 340 transmits a control signal to the core B clock based on the state of the assigned EU and sends multiplexer (MUX) coordinates to connect the input / output (IO) ports of the assigned EU. Additionally, the EE manager 300 transmits an assigned EU ACK to the dispatch stage 320 of the requesting active core A at time 430, which includes the EU ID of the assigned EU.

[0036] In response to receiving the ACK for the assigned EU at time 430, the dispatch stage 320 transmits a transaction packet to the EE NOC 340 at time 440, which contains the source operand of the issued instruction and the result buffer word address as the destination operand of the assigned EU. Additionally, the dispatch stage 320 replaces the forwarded instruction with a simple load operation to forward the destination register of the instruction for transferring the result from the specified word address of the result buffer 352 to the destination register as shown in Figure 3

[0037] In this example, the EE NOC 340 provides an interface to the assigned EU for unpacking the transaction packet received from the dispatch stage 320 at time 440 and loading the result buffer 352 with the appropriate operands at time 450 (see Figure 3 ) and engaging the EU with the join signal. In response, the assigned EU operates on the source operand and transmits the processed result to the destination buffer address in the result signal at time 460, which is stored by the EE NOC 340 in the result buffer 352 at the word address provided by the active core A. In this example, the dispatch stage 320 dispatches the replaced load instruction signal at time 480, which causes the forwarded instruction to reach the commit stage at time 490. Once the forwarded instruction reaches the commit stage in the pipeline of the active core A at time 470, the result is loaded into the destination register of the instruction using the commit signal at time 490.

[0038] Figure 5 ​is further illustrative of various aspects in accordance with the present disclosure Figure 2 A block diagram of a system-on-a-chip (SoC) including an execution engine manager for supporting execution unit (EU) sharing operations between processor cores. As shown Figure 5 , the processor core cluster 501 is similar to the processor core cluster 301 of Figure 3 , and like reference numerals are used to describe.

[0039] In Figure 5 , a network-on-chip / bus (NOC / BUS) 540 (540-0, 540-1, 540-2, 540-3, 540-4) is added to the fetch stage 302, decode stage 310, dispatch stage 320, and execution engine stage 330 of Figure 3 . Additionally, a cluster resource manager 500 is provided to coordinate and allocate pipeline stages from power gated cores. In various aspects of the present disclosure, when a pipeline stage of one core reaches a structural hazard point (e.g., when two or more instructions in a processor core pipeline request access to the same resource), a request is issued to the cluster resource manager 500 to allocate a spare pipeline stage from other processor cores. If the requested stage is contextless, the stage can be allocated when the stage is not in use by another core. If the requested stage is context-aware, the stage can be allocated when the stage is power gated in another core.

[0040] Sharing idle EUs with active processor cores beneficially utilizes available power and performance cores, such as scalable vector extensions and other binary agnostic extension EUs in a processor core cluster. Additionally, idle EU sharing provides higher performance because more EUs are available to active cores during execution. Idle EU sharing provides various performance benefits, such as run-time reduction. Various aspects of the present disclosure utilize idle EUs in a processor core in a reset / power gated stage to execute non-predicted paths from a branch predictor. Executing both paths and committing results from the taken path provides a flushless pipeline execution. Additionally, idle EU sharing enables configuration of peak single core performance. For example, a scalable vector extension can run with wider register lengths by using all available scalable vector extension units.

[0041] Inclusion of NOC / BUS 540 and cluster resource manager 500 at each pipeline stage incurs area overhead. In addition to the 3D vertical cache in that tier, additional infrastructure can be offloaded to the top tier die in a 3D integrated circuit (IC) package configuration. Thus, inclusion of NOC / BUS 540 and cluster resource manager 500 at each pipeline stage provides an opportunity to dynamically reconfigure single core capabilities. This fluid nature of configuring the stages of any core to each other enables efficient utilization of the pipeline in the cluster. Multiple front-ends can feed a single execution engine in a front-end stalled application, or multiple execution engines served by a single core's front-end, forming a dynamic pipeline that attempts to complete execution by the shared EU. For example, as shown in Figure 6 the process for idle EU sharing can be performed.

[0042] Figure 6 is a process flow diagram illustrating a method for execution engine (EE) sharing between processor cores in accordance with various aspects of the present disclosure. The method 600 begins at block 602, where a structural hazard associated with an issued instruction is encountered in an instruction queue of a dispatch stage within an active processor core. For example, as shown in timing diagram 400 of Figure 4 a structural hazard is encountered in an instruction queue of a dispatch stage 320 within an active core (core A). In various aspects of the present disclosure, when a pipeline stage of one core reaches a structural hazard point (e.g., when two or more instructions in a processor core pipeline request access to the same resource), a request is issued to the cluster resource manager 500 to allocate a spare pipeline stage from other processor cores. If the requested stage is contextless, the stage can be allocated when the stage is not in use by another core. If the requested stage is context-aware, the stage can be allocated when the stage is power gated in another core.

[0043] At block 604, a request is issued for an idle execution unit of an inactive processor core. For example, in timing diagram 400 of Figure 4 when a structural hazard is encountered in an instruction queue of a dispatch stage 320 within an active core (core A), the dispatch stage 320 issues a request for a specified type of execution unit (EU) at time 410.

[0044] At block 606, a transaction containing a source operand of the issued instruction and a word address of a result buffer as a destination operand is transferred to the allocated EU of the inactive processor core. For example, as shown in timing diagram 400 of Figure 4As shown, EE Manager 300 transmits a control signal to EE NOC 340 at time 420 to provide power to activate the EU. At time 425, EE NOC 340 transmits a control signal to Core B clock based on the state of the allocated EU and sends multiplexer (MUX) coordinates to connect the input / output (IO) ports of the allocated EU. Additionally, EE Manager 300 transmits an allocated EU ACK to scheduling stage 320 of requesting active Core A at time 430, which includes the EU ID of the allocated EU.

[0045] At block 608, the issued instruction is replaced in the instruction queue with a load operation to forward the result of the issued instruction from the result buffer based on the word address. For example, as shown in Figure 4 At time 460, the allocated EU operates on the source operand and transmits the processed result to the destination buffer address in the result signal, which is stored by EE NOC 340 in the result buffer 352 at the word address provided by active Core A. In this example, scheduling stage 320 schedules the replaced load instruction signal at time 480, which causes the forwarded instruction to arrive at the commit stage at time 490. Once the forwarded instruction arrives at the commit stage in the pipeline of active Core A at time 470, the result is loaded to the destination register of the instruction using the commit signal at time 490.

[0046] In some aspects, the method 600 can be performed by the host SoC 100 Figure 1 ), by way of example and not limitation, each of the elements of the method 600 can be performed by the host SoC 100 or one or more processors (e.g., the multi-core CPU 102 and / or the NPU 130) and / or other components included therein.

[0047] Figure 7 is a process flow diagram illustrating a method for an execution engine (EE) manager to support processor cores in accordance with various aspects of the present disclosure. The method 700 begins at block 702, where a state of execution units (EUs) in a cluster of processor cores is monitored. For example, as shown in Figure 3 In active mode, EE Manager 300 maintains an active catalog of idle EUs by monitoring the decoded instruction queue in decode stage 310 and dependency chain information from scheduling stage 320, including the current utilization state and future utilization state of each EU of execution engine stage 330. In reactive mode, EE Manager 300 collects EU states on demand when scheduling stage 320 of a core issues a request for an additional EU.

[0048] At block 704, a request for an idle execution unit (EU) in a cluster of processor cores is received. For example, inFigure 4 In timing diagram 400, when a structural hazard is encountered in the instruction queue of the scheduling stage 320 within the active core (core A), the scheduling stage 320 issues a request for a specified type of execution unit (EU) at time 410.

[0049] At block 706, a control signal is transmitted to activate the allocated EU of the inactive processor core. For example, as shown in Figure 4 , the EE manager 300 transmits a control signal to the EE NOC 340 at time 420 to provide power to activate the EU. At block 708, an EU acknowledgement and an EU identification (EU ID) are transmitted to the scheduling stage of the requesting processor core. For example, as shown in Figure 4 , at time 425, the EE NOC 340 transmits a control signal to the core B clock based on the status of the allocated EU and sends multiplexer (MUX) coordinates to connect the input / output (IO) ports of the allocated EU. Additionally, the EE manager 300 transmits an allocated EU ACK to the scheduling stage 320 of the requesting active core A at time 430, which includes the EU ID of the allocated EU.

[0050] At block 710, an EE network-on-chip (NOC) (EE NOC) transmits the issued instruction in the instruction queue of the scheduling stage within the active processor core to the allocated EU. For example, as shown in Figure 4 , the EE NOC 340 provides an interface to the allocated EU for unpacking transaction packets received from the scheduling stage 320 at time 440 and loading a result buffer 352 with appropriate operands at time 450 (see Figure 3 ) and engaging the EU through engagement signals. In response, the allocated EU operates on the source operands and transmits the processed result to the destination buffer address in the result signal at time 460, which is stored by the EE NOC 340 in the result buffer 352 at the word address provided by the active core A.

[0051] Figure 8 is a block diagram illustrating an example wireless communication system 800 in which aspects of the disclosure can be advantageously employed. For purposes of example, three remote units 820, 830, and 850 and two base stations 840 are shown. Figure 8 It should be appreciated that the wireless communication system can have more remote units and base stations. The remote units 820, 830, and 850 include IC devices 825A, 825B, and 825C that include the disclosed execution unit sharing operations. It should be appreciated that other devices can also include the disclosed execution unit sharing operations, such as base stations, switching devices, and network equipment. Figure 7Forward link signals 880 from the base station 840 to remote units 820, 830, and 850 are shown, as are reverse link signals 890 from the remote units 820, 830, and 850 to the base station 840.

[0052] In Figure 8 The remote units 820 are illustrated as mobile telephones, the remote unit 830 is illustrated as a portable computer, and the remote unit 850 is illustrated as a fixed location remote unit in a wireless local loop system. The remote units may, for example, be mobile phones, hand-held personal communication systems (PCS) units, portable data units such as personal data assistants, GPS enabled devices, navigation devices, set top boxes, music players, video players, entertainment units, fixed location data units such as meter reading equipment, or any other device that stores or retrieves data or computer instructions, or a combination thereof. Although Figure 8 Remote units in accordance with aspects of the present disclosure are illustrated, but the present disclosure is not limited to these exemplary illustrated units. Aspects of the present disclosure can be applicable to a number of devices that include the disclosed performing unit sharing operations.

[0053] Figure 9 is a block diagram of a design workstation that illustrates a circuit, layout, and logic design for a semiconductor component, such as an execution engine (EE) manager that performs unit sharing operations as disclosed above. The design workstation 900 includes a hard disk 901 that contains an operating system, support files, and design software, such as Cadence or OrCAD. The design workstation 900 also includes a display 902 to facilitate the design of a circuit 910 or integrated circuit (IC) component 912, such as an EE manager that performs unit sharing operations. A storage medium 904 is provided for tangibly storing the design of the circuit 910 or IC component 912 (e.g., an EE manager that performs unit sharing operations). The design of the circuit 910 or IC component 912 can be stored on the storage medium 904 in a file format, such as GDSII or GERBER. The storage medium 904 can be a CD-ROM, a DVD, a hard disk, flash memory, or other appropriate device. Moreover, the design workstation 900 includes a drive apparatus 903 for accepting input from or writing output to the storage medium 904.

[0054] The data recorded on the storage medium 904 can specify a logic circuit configuration, pattern data for photolithography masks, or mask pattern data for a serial write tool, such as an e-beam lithography machine. The data can also include logic verification data, such as timing diagrams or network circuits associated with logic simulation. The provision of data on the storage medium 904 facilitates the design of the circuit 910 or IC component 912 by reducing the number of processes used to design a semiconductor wafer.

[0055] Various implementation examples are described in the following numbered clauses: 1. A method of execution unit (EU) sharing among processor cores, the method comprising: encountering a structural hazard associated with an issued instruction in an instruction queue of a dispatch stage within an active processor core; issuing a request for an idle execution unit of an inactive processor core; transferring a transaction containing a source operand of the issued instruction and a word address of a result buffer as a destination operand to an assigned EU of the inactive processor core; and replacing the issued instruction with a load operation in the instruction queue to forward a result of the issued instruction from the result buffer based on the word address.

[0056] 2. The method of clause 1, further comprising receiving an EU acknowledgement and an EU identification (EU ID) at the dispatch stage of a requesting processor core.

[0057] 3. The method of any one of clauses 1 or 2, further comprising storing the result in a register according to the word address to commit the instruction.

[0058] 4. The method of any one of clauses 1-3, wherein transferring the transaction comprises issuing the issued instruction to the assigned EU of the inactive processor core for execution.

[0059] 5. The method of any one of clauses 1-4, wherein encountering comprises: detecting the issued instruction that requires access to a same hardware resource as a previously issued instruction; and replacing the issued instruction with the load operation in the instruction queue.

[0060] 6. The method of any one of clauses 1-5, further comprising transferring a control signal to activate the assigned EU prior to transferring the transaction containing the source operand.

[0061] 7. The method of clause 6, wherein the control signal comprises instructions, data, power, and clock signals to / from the assigned execution unit.

[0062] 8. The method of any one of clauses 1-7, further comprising connecting an input / output (IO) port of the assigned EU to an execution engine (EE) network-on-chip (NOC) (EE NOC).

[0063] 9. The method of any of clauses 1-8, further comprising receiving a no-assignment acknowledgement (ACK) when a free EU is not available.

[0064] 10. A method for an execution engine (EE) manager to support processing cores, the method comprising: monitoring a state of execution units (EUs) in a cluster of processing cores; receiving a request for a free execution unit (EU) in the cluster of processing cores; transmitting a control signal to activate an assigned EU of an inactive processing core; transmitting an EU acknowledgement and an EU identification (EU ID) to a dispatch stage of a requesting processing core; and transmitting, over an EE network-on-chip (NOC) (EE NOC), an issued instruction in an instruction queue of the dispatch stage internal to an active processing core to the assigned EU.

[0065] 11. The method of clause 10, upon receiving the request, the method further comprising: identifying the free EU of the inactive processing core from a directory of free EUs in the cluster of processing cores; and allocating the free EU of the inactive processing core as the assigned EU.

[0066] 12. The method of any of clauses 10 or 11, the method further transmitting a no-assignment acknowledgement (ACK) to the dispatch stage of the requesting processing core in the event that a free EU from the cluster of processing cores is not available.

[0067] 13. The method of any of clauses 10-12, the method further comprising: executing the issued instruction by the assigned EU to generate a result; and transmitting the result to a destination buffer address.

[0068] 14. The method of clause 13, the method further comprising deactivating the assigned EU after transmitting the result to the destination buffer address.

[0069] 15. The method of any of clauses 10-14, wherein transmitting the control signal comprises sending a multiplexer (MUX) coordinate to connect an input / output (I / O) port of the assigned EU of the inactive processing core.

[0070] 16. The method of any of clauses 10-15, further comprising storing results from the allocated EU in a result buffer after execution of the issued instruction until the issued instruction is committed.

[0071] 17. The method of any of clauses 10-16, wherein the control signals comprise power and clock signals for the allocated EU.

[0072] 18. The method of any of clauses 10-17, further comprising connecting input / output (IO) ports of the allocated EU to an execution engine (EE) network-on-chip (NOC) (EE NOC).

[0073] 19. The method of any of clauses 10-18, wherein transmitting the issued instruction further comprises: receiving a transaction containing a source operand of the issued instruction and a word address of a result buffer as a destination operand; and sending the transaction to the allocated EU of the inactive processor core.

[0074] 20. The method of any of clauses 10-19, wherein monitoring comprises: detecting the idle EU in the cluster of processor cores; and adding the idle EU to an idle EU directory in the cluster of processor cores.

[0075] For firmware and / or software-based implementations, the methods can be implemented with modules (e.g., procedures, functions, and so on) that execute functions described herein. A machine-readable medium, tangibly embodying instructions, can be used to implement the methods described herein. For example, software codes can be stored in memory and executed by a processor unit. Memory can be implemented within the processor unit or external to the processor unit. As used herein, the term "memory" refers to all types of long-term, short-term, volatile, nonvolatile, or other memory and is not to be limited to a particular type of memory or number of memories, or type of media upon which memory is stored.

[0076] If implemented in firmware and / or software, the functions can be stored as one or more instructions or code on a computer-readable medium. Examples include computer-readable media encoded with a data structure and computer-readable media encoded with a computer program. Computer-readable media includes physical computer storage media. A storage medium can be any available medium or means of storing data. As an example, and not by way of limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Disk and disc, as used herein, includes compact discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks and blu-ray discs where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. ® In addition to storage on computer-readable medium, instructions and / or data can be provided as signals on transmission media included in a communication apparatus. For example, a communication apparatus can include a transceiver having signals indicative of instructions and data. These instructions and data are configured to cause one or more processors to implement the functions outlined in the claims.

[0077] instructions and data are configured to cause one or more processors to implement the functions outlined in the claims.

[0078] While the present disclosure and its advantages have been described in detail, various changes, substitutions and alterations can be made hereto without departing from the technology of the present disclosure as defined by the appended claims. For example, relational terms such as "above" and "below" are used. Of course, if the substrate or electronic device is inverted, then above becomes below and vice versa. Additionally, if side-oriented, then above and below can refer to the sides of the substrate or electronic device. Moreover, the scope of the present application is not intended to be limited to the particular configurations of processes, machines, manufactures, compositions of matter, means, methods and steps described in the specification. As one of ordinary skill in the art will readily appreciate from the disclosure, processes, machines, manufacture, compositions of matter, means, methods or steps, presently existing or later to be developed that perform substantially the same function or achieve substantially the same result as the corresponding configurations described herein can be utilized according to the present disclosure. Accordingly, the appended claims are intended to include within their scope such processes, machines, manufacture, compositions of matter, means, methods or steps.

[0079] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0080] The various illustrative logical blocks, modules, and circuits described in connection with the disclosure herein can be implemented or performed with a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0081] The steps of a method or algorithm described in connection with the present disclosure can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM, flash memory, ROM, EPROM, EEPROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal.

[0082] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for sharing execution units (EUs) among processor cores, the method comprising: Structural hazards associated with issued instructions are encountered in the instruction queue at the scheduling level within the active processor core; Issue a request for an idle execution unit in an inactive processor core; The transaction containing the source operand of the issued instruction and the word address of the result buffer as the destination operand are transferred to the allocated EU of the inactive processor core. as well as The published instruction is replaced with a load operation in the instruction queue to forward the result of the published instruction from the result buffer based on the word address.

2. The method of claim 1, further comprising receiving an EU acknowledgment and an EU identifier (EU ID) at the scheduling level of the requesting processor core.

3. The method according to claim 1, further comprising storing the result in a register according to the word address to submit the instruction.

4. The method of claim 1, wherein transmitting the transaction comprises issuing the issued instructions to the allocated EU of the inactive processor core for execution.

5. The method of claim 1, wherein encountering includes: The issued instruction is detected; the issued instruction requires access to the same hardware resources as the previously issued instruction. as well as The issued instruction is replaced in the instruction queue with the load operation.

6. The method of claim 1, further comprising transmitting a control signal to activate the allocated EU before transmitting the transaction containing the source operand.

7. The method of claim 6, wherein the control signal includes instructions, data, power, and clock signals to / from the assigned execution unit.

8. The method of claim 1, further comprising connecting the input / output (I / O) ports of the allocated EU to the Execution Engine (EE) Network on Chip (NOC) (EE NOC).

9. The method of claim 1, further comprising receiving an unassigned acknowledgment (ACK) when an idle EU is unavailable.

10. A method for providing an Execution Engine (EE) manager with support for a processor core, the method comprising: Monitor the status of execution units (EUs) in the processor core cluster; Receive requests for idle execution units (EUs) in the processor core cluster; Transmit control signals to activate the allocated EUs of inactive processor cores; The EU acknowledgment and EU identifier (EU ID) are transmitted to the scheduling level of the requesting processor core; and Instructions issued in the instruction queue of the scheduling level within the active processor core are transmitted to the assigned EU via the EE Network on-Chip (NOC) (EE NOC).

11. The method of claim 10, wherein upon receiving the request, the method further comprises: Identify the idle EUs of the inactive processor core from the idle EU directory in the processor core cluster; as well as The idle EUs of the inactive processor cores are allocated as the assigned EUs.

12. The method of claim 10, wherein the method further transmits an unallocated acknowledgment (ACK) to the scheduling level of the requesting processor core if an idle EU from the processor core cluster is unavailable.

13. The method according to claim 10, further comprising: The issued instructions are executed by the assigned EU to generate the results; as well as The result is then transmitted to the destination buffer address.

14. The method of claim 13, further comprising deactivating the allocated EU after transmitting the result to the destination buffer address.

15. The method of claim 10, wherein transmitting the control signal includes sending multiplexer (MUX) coordinates to connect the input / output (I / O) ports of the allocated EU of the inactive processor core.

16. The method of claim 10, further comprising storing the results from the allocated EU in a result buffer after executing the published instruction, until the published instruction is submitted.

17. The method of claim 10, wherein the control signal includes power and clock signals for the allocated EU.

18. The method of claim 10, further comprising connecting the input / output (I / O) ports of the allocated EU to the execution engine (EE) network on-chip (NOC) (EE NOC).

19. The method of claim 10, wherein transmitting the issued instructions further comprises: The transaction receives the source operand containing the issued instruction and the word address of the result buffer as the destination operand; as well as The transaction is sent to the allocated EU of the inactive processor core.

20. The method of claim 10, wherein monitoring comprises: Detect the idle EUs in the processor core cluster; as well as Add the idle EU to the idle EU directory in the processor core cluster.