Devices, methods, and systems for determining credit-based access to shared circuit resources
By assigning credit to each processor circuit and allocating resource access rights based on credit, the problem of unfair allocation of coprocessor resources in multiprocessor systems is solved, and the system efficiency is improved.
Patent Information
- Application Number
- CN202411750737.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-27
AI Technical Summary
In multiprocessor systems, the resource allocation of coprocessors is unfair, resulting in an increase in execution waiting time and affecting system efficiency.
By assigning credit to each processor circuit, access rights to shared circuit resources are determined based on the size of credit, and a credit-based solution is used to allocate resource access.
The fair allocation of shared circuit resources is achieved, the execution waiting time is reduced, and the system efficiency is improved.
Smart Images

Figure CN120216035A_ABST
Abstract
Description
Background 1. Technical Field
[0001] The present disclosure generally relates to processors, and more specifically but not exclusively to sharing computing resources among multiple processor circuits. 2. Background Art
[0002] In many computer systems, for any of a variety of reasons, a coprocessor is shared among multiple processor units, such as to enable execution of additional instruction sets on demand. Such sharing of coprocessors generally aids in the efficient use of silicon area. However, system efficiency typically depends to a large extent on the fairness of allocating access to the coprocessor by the various processor units. In many cases, an unfair allocation scheme increases the likelihood that one processing unit will prevent the coprocessor from performing tasks on behalf of any other processing unit among the processing units. Thus, in some multiprocessor systems, an unfair allocation scheme results in overall execution latency. As successive generations of multiprocessor architectures continue to increase in number, variety, and capabilities, improvements in how to make shared circuit resources available to different processors in different ways are expected to receive increasing attention. Brief Description of the Drawings
[0003] Embodiments of the present invention are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which:
[0004] Figure 1 is a block diagram illustrating features of a system 100 for determining access to resources of a computing circuit according to an embodiment.
[0005] Figure 2 is a flowchart illustrating features of a method for determining access to computing circuit resources according to an embodiment.
[0006] Figure 3 is a flowchart illustrating features of a method for managing credit information to assist in determining access to computing circuit resources according to an embodiment.
[0007] Figure 4 is a flowchart illustrating features of a method for managing a pool of processor circuits that share access to computing circuit resources according to an embodiment.
[0008] Figure 5 is a flowchart illustrating features of a method for a processor circuit to assign tasks to computing circuit resources based on a current credit accumulation according to an embodiment.
[0009] Figure 6 is a block diagram illustrating features of a system 600 for providing access to a computing circuit according to an embodiment.
[0010] Figure 7 is a flowchart illustrating features of a method for selecting a scenario according to an embodiment, where computing circuit resources are to be accessed according to the scenario.
[0011] Figure 8 is a flowchart illustrating features of a method for assigning tasks to computing circuit resources based on priorities corresponding to processor circuits according to an embodiment.
[0012] Figure 9 Illustrates an exemplary system.
[0013] Figure 10 Illustrates a block diagram of an example processor that may have more than one core and may have an integrated memory controller.
[0014] Figure 11A is a block diagram illustrating an exemplary in-order pipeline and both an exemplary register renaming, out-of-order issue / execution pipeline according to an example.
[0015] Figure 11B is a block diagram illustrating an exemplary example of an in-order architecture core and both an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor according to an example. Detailed Description
[0016] The embodiments discussed herein provide techniques and mechanisms in various ways for any one of a plurality of processor circuits to access shared circuit resources. Under at least some conditions, such access is determined based on each of one or more processor circuits being associated with a respective amount of credit.
[0017] In an embodiment, a first processor circuit that has requested access to a shared circuit resource accumulates credit while the first processor circuit waits for the access. For example, the rate of such credit accumulation is based on the current total number of one or more processor circuits having a current request to access the shared resource. In contrast, when the shared resource performs a task on behalf of the processor circuit, the same (or another) processor circuit consumes (i.e., loses) credit at a relatively high rate. In one such embodiment, a credit-based scheme is used to select an access request for servicing on behalf of a requesting processor circuit, e.g., where the selection is based on determining that the requesting processor circuit is currently the most highly accredited processor circuit. Some embodiments make various transitions between performing access allocation according to a credit-based scheme and performing access allocation according to an alternative such as a priority-based scheme.
[0018] The techniques described herein can be implemented in one or more electronic devices. Non-limiting examples of electronic devices that can utilize the techniques described herein include any kind of mobile device and / or stationary device, such as cameras, cellular phones, computer terminals, desktop computers, e-readers, fax machines, self-service machines, laptop computers, netbook computers, notebook computers, Internet devices, payment terminals, personal digital assistants, media players and / or recorders, servers (e.g., blade servers, rack-mounted servers, combinations thereof, etc.), set-top boxes, smart phones, tablet personal computers, ultra-mobile personal computers, landline telephones, combinations of the foregoing, and so on. More generally, the techniques described herein can be employed in any of a variety of electronic devices, including logic for allocating access to shared circuit resources.
[0019] Figure 1 System 100 for determining access to resources of a computing circuit in accordance with an embodiment is shown. System 100 illustrates features of one example embodiment in which shared circuit resources can be used by any of a plurality of processor circuits, where access to the circuit resources by one such processor circuit is allocated according to a credit-based scheme. When a given processor circuit is waiting for a requested access to the circuit resources, the amount of credit attributed to that processor circuit increases at a first rate, which is based on the total number of one or more processor circuits currently waiting for resource access. In one such embodiment, when access to the circuit resources is provided on behalf of the processor circuit, the amount of credit decreases at a second rate that is, for example, faster (i.e., has a greater magnitude) than the first rate.
[0020] As Figure 1As shown, system 100 includes multiple processor circuits (such as the illustrative processor circuits 110a, …, 110x shown), where each processor circuit includes a corresponding one or more processor cores. By way of illustration and not limitation, one or more of the processor circuits 110a, …, 110x are each a corresponding one of a single-core or multi-core processing unit. For example, a given such processing unit is a central processing unit (CPU). In another embodiment, some or all of the processor circuits 110a, …, 110x are different corresponding cores of the same processing unit. In various embodiments, some or all of the processor circuits 110a, …, 110x are on the same integrated circuit (IC) chip. For example, a network-on-chip (NoC) or system-on-chip (SoC) includes two or more of the processor circuits 110a, …, 110x. Alternatively or additionally, two or more of the processor circuits 110a, …, 110x are on different corresponding IC chips - for example, IC chips included in the same packaging device and / or IC chips of different corresponding packaging devices.
[0021] System 100 further includes any one of various types of circuits (referred to herein as “computing circuits”), or alternatively, couplings adapted to any one of various types of circuits, which are capable of providing computing functions to support a given one of the processor circuits 110a, …, 110x. By way of illustration and not limitation, computing circuit 120 includes any one of various suitable coprocessors to perform floating-point (or other) arithmetic operations, graphics operations, signal processing operations, string processing operations, cryptographic operations, input / output (I / O) interface operations, formatting operations, machine learning operations, etc. Alternatively or additionally, computing circuit 120 includes any one of various suitable accelerator circuits, such as a digital streaming accelerator (DSA). Alternatively or additionally, for example, computing circuit 120 includes any one of various suitable graphics processing units. However, some embodiments are not limited to specific computing circuits that can be shared for access by multiple processor circuits.
[0022] During operation of system 100, some or all of processor circuits 110a, …, 110x each execute a respective one or more software processes—for example, including an operating system (OS), one or more software applications, a virtual machine manager, one or more virtual machines, and the like. To facilitate such software execution, a given one of processor circuits 110a, …, 110x (generally referred to herein as processor circuit 110) explicitly or implicitly requests access to resources 122 implemented by some or all of computing circuit 120. In an illustrative scenario according to one embodiment, processor circuits 110a, 110x submit respective requests 111a, 111x for access to computing circuit 120 differently.
[0023] In an embodiment, system 100 includes logic (e.g., including circuit hardware, firmware, and / or executing software) to allocate access to resources 122 differently on behalf of processor circuits 110a, 110x (and similarly, on behalf of corresponding access requests 111a, 111x). This access is allocated according to a credit-based scheme, where credits are attributed differently based on the current total number of one or more processor circuits, each processor circuit having a current request for such access. In the illustrated example embodiment, such logic includes some or all of task assignment unit 130, pool manager 140, and credit manager 150.
[0024] In one such embodiment, a given processor circuit 110 submits a request (referred to herein as an “access request”) to pool manager 140 to be allocated access to shared resources 122 (also referred to herein as “computing circuit resources” or, for brevity, simply “resources”). In some embodiments, task assignment unit 130, pool manager 140, and credit manager 150 each differently include any one of one or more microcontrollers, state machines, application-specific integrated circuits (ASICs), programmable gate arrays (PGAs), and / or various other types of circuitry systems that are adapted to service such access requests at least in part by assigning one or more tasks to be performed with shared resources 122 of computing circuit 120.
[0025] An access request from the processor circuit 110 is a "current request" until the requested access to the resource 122 has been completed - for example, until each task assigned by the task assignment unit 130 on behalf of the requesting processor circuit 110 has been completed using the shared resource 122. At a given time, a current request can be in either of two states, which are referred to herein as the "pending" state and the "active" state. An active request is the current access request for which the resource 122 is currently executing the corresponding task. When a current access request is active, any other current access request is a "pending request", which is, for example, waiting to be assigned a task to be executed on behalf of the corresponding processor circuit 110 using the resource 122. When the resource 122 has completed each task for a given access request, that access request is a "completed request" and is no longer current. To facilitate the quick retrieval of task results, some embodiments notify the access requester of completed tasks even before all tasks for the access request in question have been executed.
[0026] To facilitate fair allocation of access to the shared computing circuit resources, the pool manager 140 supports the management of a pool of those one or more processor circuits (if any), each of the one or more processor circuits corresponding to a current respective access request (e.g., provided by each processor circuit). The term "pool member" (or, for brevity, simply "member") is used herein to refer to a processor circuit corresponding to a current access request and thus currently belonging to the pool of processor circuits (or simply "pool"). A given pool member corresponds to the respective access request at least in that the pool member generates (or otherwise causes the generation of) the access request, e.g., where the request is for accessing a shared resource on behalf of the corresponding pool member.
[0027] The term "pending member" is used herein to refer to a pool member for which the corresponding access request (made by the member in question) is currently pending. The term "active member" is used herein to refer to a pool member for which the corresponding access request is currently active. The term "pool size" (or, for brevity, simply "size") is used herein to refer to the current total number (N) of one or more members of the pool of processor circuits. Over time, the pool size changes - for example, where processor circuits enter and leave the pool differently as the access requests corresponding to the respective processor circuits are generated and completed differently.
[0028] The pool manager 140 includes pool status information 142 that specifies or otherwise indicates which one or more of the processor circuits 110 (if any) are currently pool members. The pool manager 140 further includes circuitry and / or other suitable logic for periodically updating the pool status information 142 as access requests are generated and completed over time in various ways. The pool status information 142 is provided, for example, by any of a variety of one or more data structures (such as, including tables) adapted to identify the current pool members. In some embodiments, the pool status information 142 (and / or other status information available to the task assignment unit 130 and / or the credit manager 150) further includes other status information that facilitates the allocation of access to the resources 122.
[0029] In the illustrated example embodiment, the pool status information 142 includes or otherwise represents entries, each corresponding to a different respective one of the processor circuits 110a, …, 110x. A given one of such entries includes, for example, a field (Pid) for providing an identifier of the corresponding processor circuit, and another field (Mem) for identifying whether the corresponding processor circuit is currently a member of the pool. In one such embodiment, a given entry further includes a field (Cdt) for providing a variable (referred to herein as a “credit variable”) for indicating the amount of credit currently (or, for example, most recently) attributed to the corresponding processor circuit. Alternatively or additionally, a given entry further includes another field (Ract) for identifying whether the corresponding processor circuit is currently an active member or a pending member.
[0030] In an illustrative scenario according to one embodiment, the first entry of the pool status information 142 includes: an identifier Pu1 of the processor circuit 110a, a flag set to specify that the processor circuit 110a is currently a pool member, an identifier n1 of the amount of credit currently attributed to the processor circuit 110a (the “amount of credit” herein), and another flag set to specify that the processor circuit 110a is currently a pending member. Additionally, the second entry of the pool status information 142 includes: an identifier Pux of the processor circuit 110x, a flag set to specify that the processor circuit 110x is currently a pool member, an identifier nx of the amount of credit currently attributed to the processor circuit 110x, and another flag set to specify that the processor circuit 110x is currently an active member. Although some embodiments are not limited in this regard, the third entry of the pool status information 142 includes: an identifier Pu2 of another processor circuit 110, and a flag specifying that the other processor circuit 110 is currently not a pool member. In an alternative embodiment, the entries of the pool status information 142 are only for current pool members (i.e., rather than for any processor circuit that is not currently a pool member).
[0031] The credit manager 150 includes any one of various types of circuitry adapted to maintain up-to-date status information indicating the amount of credit currently attributed to a given pool member (and thus also to the access request corresponding to that pool member). For example, the credit manager 150 periodically (e.g., continuously, incrementally, or otherwise) updates one or more credit variables in respective different Cdt fields of the pool status information 142.
[0032] In the illustrated example embodiment, the credit manager 150 includes accumulation logic 152 that determines the amount by which a given credit variable is to be increased or otherwise changed to indicate a greater amount of credit. For example, when the corresponding pool member is a pending pool member, the given amount of credit is to be increased with the accumulation logic 152. In various embodiments, a given credit variable is updated once per cycle in such a periodic sequence (also referred to herein as an "authentication cycle" or "evaluation cycle") for successively determining each amount of credit. For example, one such given cycle includes one or more cycles of a clock signal for the circuitry of the operating system 100.
[0033] As used herein, "credit accumulation rate" (or, for simplicity, simply "accumulation rate") refers to the additional amount by which the credit attributed to a pool member will increase in each cycle. In some embodiments, the accumulation rate of a pool member is based on the current size of the pool. By way of illustration and not limitation, the "pool rate" is the total amount of all additional credit that is further attributed to one or more pool members (or at least to one or more pending pool members) in a given cycle. For example, the pool rate is represented by K credit units, and N is the current size of the pool (where K is some positive number and N is a non-negative integer). In one such embodiment, in a given cycle, one or more pool members (or at least one or more pending pool members) will each be further attributed a corresponding portion of K credit units. For example, in some embodiments, one or more pool members (or at least one or more pending pool members) will each receive K / N credit units in a given cycle.
[0034] In another embodiment, one or more pool members each accumulate additional credits at different respective cumulative rates, e.g., where the different rates are each based on the pool size and further based, e.g., on different respective weight factors (also referred to herein as "scaling factors"). In one such embodiment, some or all of the one or more pool members (and / or corresponding access requests of such one or more pool members) each correspond to a respective factor that is used to weight how K credit units are to be differently distributed among the pool members in a given period. For example, a given such weight factor is based on the operational characteristics of the corresponding pool member - e.g., based on the quality of service to be provided with the pool member, the specific hardware thread capabilities of the pool member, etc.
[0035] By way of illustration and not limitation, at least some pool members correspond to different respective weight values (e.g., values w1, w2, …, wx), where the total weight value Wtot is equal to the sum of the weight values.
[0036] In an illustrative scenario according to one such embodiment, the cumulative rate of a such pool member is proportional to (e.g., equal to) the product of the ratio K / N and the ratio of the corresponding weight value to the total weight value Wtot. For example, in one such embodiment, a first pool member corresponding to the weight value w1 accumulates credits at a first per-period rate k1, where: k1∝(K·w1) / (N·Wtot).
[0037] Alternatively or additionally, a second pool member corresponding to the weight value w2 accumulates credits at a second per-period rate k2, where: k2∝(K·w2) / (N·Wtot).
[0038] However, in other embodiments, the accumulation of credits at different respective rates by the various pending pool members is according to any of various additional or alternative suitable weighting schemes.
[0039] The credit manager 150 further includes consumption logic 154 that determines, for a given credit variable, the amount by which the credit variable is to be decreased or otherwise changed to indicate a lesser amount of credit. In one such embodiment, a given amount of credit will be decreased when the corresponding pool member is an active pool member. As used herein, "credit consumption rate" (or, for simplicity, simply "consumption rate") refers to the amount by which the credit attributed to a pool member will be decreased in each period.
[0040] In some embodiments, the amount of credit of an active pool member during a given period decreases at a consumption rate that is greater than the accumulation rate that the pool member would have when in a pending state. In one such embodiment, the consumption rate is independent of the pool size, e.g., where the amount of credit of an active pool member changes by a negative K (i.e., “-K”) credit units per period, and where the total amount of credit of only the one or more pending pool members changes by a positive K (i.e., “+K”) credit units per period. In an alternative embodiment, the consumption rate of an active pool member depends on the pool size, e.g., where the consumption rate is equal to [K·(1–(1 / N))] credit units, and where the accumulation rate of each of the one or more pending pool members is equal to (K / N) credit units. In some embodiments, the consumption rate is greater than the rate of [K·(1–(1 / N))] credit units (e.g., is an integer multiple thereof), e.g., where the same active pool member is represented as performing accesses to multiple resources simultaneously. However, in other embodiments, the consumption of credit by an active pool member is according to any of a variety of additional or alternative schemes.
[0041] The task assignment unit 130 represents any of various types of logic, e.g., including hardware (such as circuitry), firmware, and / or executing software, that provides functionality to effectuate access to the resource 122 on behalf of a given pool member. For example, such access is effectuated by tasks assigned by the task assignment unit 130 to be performed using the resource 122, where the tasks are based on current access requests.
[0042] In various embodiments, the task assignment unit 130 uses a credit-based scheme to identify a pool member to transition from a pending state to an active state. By way of illustration and not limitation, the selection logic 132 of the task assignment unit 130 provides the functionality to select one pool member from all (and e.g., only) one or more current pool members as the next active pool member. In an embodiment, selecting a given pool member includes selecting an access request from (or otherwise representative of) that pool member. For example, such selection is performed based on the selection logic 132 searching for or otherwise accessing credit information (such as the credit information provided in the Cdt field of the pool state information 142) that identifies the “most accredited pool member,” i.e., the pool member currently attributed with a greater amount of credit than any other pool member.
[0043] In one such embodiment, selection logic 132 identifies the most authenticated pool member to assignment logic 134 of task assignment unit 130. Assignment logic 134 tracks one or more current access requests 136, and in response to selection logic 132, assignment logic 134 searches the current request(s) 136 to identify the "most authenticated request", i.e., the access request corresponding to the most authenticated pool member. In an embodiment, the requested access includes one or more execution cycles or execution time periods (e.g., a time period adapted to one or more execution cycles). Alternatively or additionally, the requested access includes some or all of the data stream bandwidth and / or the amount of time the data stream bandwidth is used. However, some embodiments are not limited to providing a particular type of access to resource 122.
[0044] In an illustrative scenario according to one embodiment, where it is determined that the most authenticated request is currently a pending access request, task assignment unit 130 transforms resource 122 from performing a task for another access request to performing a different task for the most authenticated request (which will become the next active access request). For example, assignment logic 134 identifies the task to be performed on behalf of the most authenticated access request. Based on such identification, task assignment unit 130 issues one or more communications (e.g., including the illustrative task assignment 121 shown) to suspend or otherwise stop the current task being performed using resource 122, and further, to resume, start, or otherwise perform another task on behalf of the most authenticated access request. In one such embodiment, assignment logic 134 (or another suitable logic of task assignment unit 130) further signals to pool manager 140 to update one or more Ract fields of pool status information 142, e.g., to indicate that the previously active access request (and corresponding pool member) is now pending, and that the previously pending access request (and corresponding pool member) is now active.
[0045] In an alternative scenario where it is determined that the active access request is the currently most authenticated request, task assignment unit 130 forgoes causing resource 122 to be transformed from performing the currently assigned task.
[0046] Instead, task assignment unit 130 simply waits until the next evaluation cycle has completed, such that the updated credit of the pool members can be reviewed again to possibly identify a different pool member as the most authenticated member.
[0047] Figure 2Method 200 for determining access to computing circuit resources according to an embodiment is shown. Method 200 illustrates an example of an embodiment in which tasks to be performed by a shared resource are assigned based on respective amounts of credit differently attributed to members of a pool of processor circuits. For a given one of the pool members, when a corresponding access request is pending, the amount of credit attributed to that member increases at a rate based on the size of the pool. Operations such as those of method 200 are performed using any combination of suitable hardware (e.g., circuitry), firmware, and / or executing software that provides some or all of the functionality of, for example, task assignment unit 130, pool manager 140, and / or credit manager 150.
[0048] As Figure 2 shown, method 200 includes (at 210) determining the current size of a pool of one or more processor circuits (if any) that are currently each requesting access to a shared resource of a computing circuit. By way of illustration and not limitation, the computing circuit includes one of a coprocessor, an accelerator (such as a digital stream accelerator), or a graphics processing unit. In one such embodiment, the shared resource includes execution cycles, execution time, stream bandwidth, and the like.
[0049] Method 200 further includes (at 212) detecting a first condition, where a first request to access a resource of the computing circuit (a first request on behalf of a first pool member) is currently pending, e.g., not currently being serviced by the shared resource. Based on the first condition, method 200 (at 214) increases a first amount of credit corresponding to the first pool member at a first rate based on the current pool size. In one such embodiment, increasing the first amount of credit at 214 includes updating one or more credit variables, each corresponding to a different respective pool member. In one such embodiment, the one or more credit variables each increase at the same credit accumulation rate (such as the accumulation rate (K / N) described herein with reference to system 100) based on whether the corresponding one or more pool members are pending pool members. In some embodiments, the credit variables of pending pool members (and accordingly, their corresponding access requests) are increased by method 200 until the pool member becomes an active pool member, or until, for example, the credit variable reaches some predetermined value representing a maximum possible amount of credit.
[0050] Method 200 further includes (at 216) detecting a second condition, where a second request to access computing circuit resources (the second request representing a second pool member) is currently being serviced. Based on the second condition, method 200 (at 218) reduces a second credit amount at a second rate, where the second credit amount corresponds to the second pool member. In one such embodiment, the second credit amount is reduced at a credit consumption rate such as the credit consumption rate described herein with reference to system 100, e.g., where the credit consumption rate is independent of the pool size. In some embodiments, the credit variable of an active pool member (and correspondingly, its corresponding access request) is reduced by method 200 until the pool member becomes a pending pool member, or e.g., until the credit variable reaches some predetermined value representing the minimum possible credit amount.
[0051] Method 200 further includes (at 220) performing a selection of a first pool member (e.g., with respect to the second pool member and any other current pool members) based on the first credit amount and the second credit amount. For example, the selection is performed at 220 based on determining that the first pool member is currently the highest authenticated pool member. Based on the selection at 220, method 200 (at 222) assigns a first task to the computing circuit resources on behalf of the first pool member.
[0052] In various embodiments, the assignment at 222 includes transforming the resources from performing a second task on behalf of the second pool member to performing the first task. In one embodiment, such a transformation is based on determining that the second credit amount has fallen below the first credit amount. In various embodiments, additionally or alternatively, such a transformation is performed based on (e.g.) determining that the execution of the second task utilizing the shared resources has reached or exceeded some threshold maximum time slice and / or other such limit. In another embodiment, additionally or alternatively, such a transformation is performed based on (e.g.) the completion of the second task (e.g., when the second pool member is the highest authenticated pool member).
[0053] In various embodiments, method 200 further includes operations (not shown) for transforming between allocating resource access according to a credit-based scheme and allocating resource access according to a priority-based scheme. In this particular context, "priority", "priority scheduling", "priority-based scheme" and related terms generally refer differently to the priority scheduling of a given pool member (and correspondingly, the priority scheduling of the corresponding access request), where such priority scheduling is distinguished from any priority scheduling that may be due to the credit amount attributed to the pool member. For example, in one such embodiment, a given pool member (or some other priority-scheduled resource) can specify the priority to be given to a corresponding memory access request. A priority is e.g., one of a range of possible priority values, e.g., where relatively high-priority access requests will be selected for access to the computing circuit resources relative to relatively low-priority access requests.
[0054] In an illustrative scenario according to one embodiment, a given pool member (corresponding to an access request) is not assigned a priority, while, for example, one or more other pool members each have a corresponding priority ranking. In one such embodiment, the selection of an access request with an unranked priority for accessing a shared resource will be based on the amount of credit of the corresponding pool member, for example, where the selection is made after each currently prioritized access request (if any) has been completed. However, in other embodiments, even if a prioritized access request is to be selected for servicing before an un-prioritized access request (and / or accumulating at a higher rate than an un-prioritized access request), when the prioritized access request runs out of credit, the servicing of the prioritized access request is still suspended.
[0055] Figure 3 Method 300 for managing credit information to assist in determining access to computing circuit resources according to an embodiment is shown. Method 300 illustrates an example of an embodiment in which the amount of credit for each different corresponding member of a processor circuit pool is kept up-to-date at least in part based on the current size of the processor circuit pool. Operations such as those of method 300 are performed using any of a variety of combinations of suitable hardware (e.g., circuitry), firmware, and / or executing software, such as providing some or all of the functions of credit manager 150 and / or other logic of system 100, for example, where method 200 includes or is combined with the operations of method 300 to be performed.
[0056] As Figure 3 shown, method 300 includes (at 310) determining the current size N of the processor circuit pool, where N is an integer representing the total number of one or more processor circuits that are currently members of the pool. For example, the determination at 310 includes counting one or more processor circuits that have provided corresponding current requests for accessing the shared resources of the computing circuit. Method 300 further performs an evaluation (at 312) to determine whether N is a positive number. In the case where it is determined at 312 that N is not a positive number (e.g., N equals zero), method 300 performs the next instance of the determination at 310.
[0057] Conversely, in the case where N is determined to be a positive number at 312, method 300 (at 314) begins a set of operations (referred to herein as an "evaluation period", or for brevity, just a "period") to determine the change in the amount of credit attributed to each processor that is currently a member of the pool. In an embodiment, beginning the period includes or otherwise causes an update of reference information (e.g., including one or more flag bits, status variables, etc.) to indicate that each current pool member has not been evaluated during the period. Subsequently, method 300 identifies (at 316) the next pool member (e.g., the first pool member) to be evaluated during the current period.
[0058] Then, method 300 performs another evaluation (at 318) to determine whether there is currently an access to a shared resource of the computing circuit on behalf of the pool member that is currently being evaluated. In the case where it is determined at 318 that there is currently such an access to the shared resource, method 300 (at 320) reduces the amount of credit corresponding to (e.g., attributed to) the pool member that is currently being evaluated. For example, the credit variable corresponding to the member being evaluated (which is determined to be an active pool member at 318) is decreased or otherwise updated to indicate a lower amount of credit. By way of illustration and not limitation, the corresponding credit variable is decreased by some amount K of credit units, or for example, by [K(N–1) / N] credit units. In one such embodiment, K other credit units (or for example, [K(N–1) / N] other credit units) will be distributed among the other (pending) pool members (if any) during the same period.
[0059] Conversely, in the case where it is determined at 318 that the access to the shared resource on behalf of the pool member being evaluated is only pending (and not actually occurring), method 300 (at 322) increases the amount of credit corresponding to the currently evaluated pool member. For example, the credit variable corresponding to the member being evaluated is increased by an amount based on the current size N of the pool. In one such embodiment, the corresponding amount of credit is increased by K / N credit units (or for example, by [K / (N–1)] credit units). In some embodiments, since this increase at 320 will be performed one or more times during a given evaluation period, each time for a different corresponding pending pool member, the total amount of distributed credit during that given evaluation period will be K credit units (or for example, [K(N–1) / N] credit units).
[0060] Method 300 further performs another evaluation (at 324) to determine whether, for example, each pool member has been evaluated in the current cycle after the decrease at 320 or the increase at 322. In the case where it is determined at 324 that at least one current pool member has not been evaluated in the current cycle, method 300 performs the next instance of the identification at 316. Conversely, in the case where it is determined at 324 that each member of the pool has been evaluated in the current cycle, method 300 performs one or more operations (at 326) to end the cycle. In an embodiment, the one or more operations include (re)setting one or more flags and / or other suitable reference information to indicate that each current pool member is to be (re)evaluated in the next cycle of method 300.
[0061] In the case where it is determined during the current cycle that the cycle is to end, method 300 further performs another evaluation (at 328) to determine whether the pool size has changed since the most recent instance of the determination at 310. In the case where it is determined at 328 that the pool size has changed, method 300 performs the next instance of the determination at 310. Conversely, in the case where it is determined at 328 that the pool size has remained unchanged since the most recent determination at 310, method 300 starts the next evaluation cycle at 314.
[0062] Figure 4 Method 400 for managing a pool of processor circuits that share access to computing circuit resources according to an embodiment is shown. Method 400 illustrates an example of an embodiment that tracks the current size of a pool of processor circuits and allocates an initial credit amount to processor circuits to be added to the pool. Operations such as those of method 400 are performed using any one of a suitable combination of hardware (e.g., circuitry), firmware, and / or executing software that provides some or all of the functionality of, for example, pool manager 140 and / or other logic of system 100, e.g., where method 200 or method 300 includes or is combined with the operations of method 400 to be performed.
[0063] As Figure 4As shown in, method 400 includes (at 410) determining the current size N of a pool of processor circuits. Method 400 further includes performing an evaluation (at 412) to determine whether N is a positive number. The size N will vary over time, so after identifying some first one or more access requests at 410, the evaluation at 412 will eventually give a positive result. In the case where it is determined at 412 that N is not a positive number (e.g., N equals zero), method 400 performs the next instance of the determination at 410. Conversely, in the case where it is determined at 412 that N is a positive number, method 400 performs another evaluation (at 414) to determine whether a request to access computing circuit resources has been received since the pool size was last determined at 412, and the request is from a requester that is not currently in the pool.
[0064] In the case where it is determined at 414 that such a new resource access request has been received, method 400 (at 416) adds the requester processor circuit (i.e., the processor circuit that sent the resource access request in question) as a member of the pool of processor circuits. For example, the addition at 416 includes changing a register, a flag, and / or any of various other types of reference information suitable for classifying a given processor circuit as currently being (or not being) a member of the pool. Additionally, method 400 (at 418) allocates an amount of credit to be attributed to the requester processor. Further still, method 400 (at 420) increments or otherwise changes a variable to indicate an increase in the pool size value N. After changing the pool size value at 420, method 400 performs another instance of the evaluation at 414.
[0065] Conversely, in the case where it is determined at 414 that no new resource access request has been received, method 400 performs another evaluation (at 422) to determine whether the servicing of the request to access computing circuit resources has been completed. In the case where it is determined at 422 that the servicing of the resource access request has not been completed (i.e., has not been completed since at least the most recent previous evaluation at 422), method 400 (at 414) performs another instance of the evaluation at 414.
[0066] Conversely, in the case where it is determined at 422 that the servicing of the resource access request has been completed, method 400 (at 424) deallocates the amount of credit currently attributed to the requester processor from the requester processor. Additionally, method 400 (at 426) removes the requester processor circuit from the pool, and (at 428) decrements or otherwise changes a variable to indicate a decrease in the pool size value N. After changing the pool size value at 428, method 400 performs another instance of the evaluation at 412.
[0067] Figure 5A method 500 for assigning tasks to computing circuit resources based on the current credit accumulation by a processor circuit according to an embodiment is shown. Method 500 illustrates an example of an embodiment, in which tasks are assigned differently to shared computing circuit resources based on one or more amounts of credit each belonging to different respective pool members. Operations such as those of method 500 are performed using any combination of suitable hardware (e.g., circuitry), firmware, and / or executing software that utilizes some or all of the functionality of, for example, selection logic 132, allocation logic 134, and / or other logic of system 100. In some embodiments, one of methods 200, 300, 400 includes the operations of method 500 or is performed in combination with the operations of method 500.
[0068] As Figure 5 shown, method 500 includes performing an evaluation (at 510) to determine whether any resource access requests are currently pending. In the case where it is determined at 510 that no resource access requests are currently pending, method 500 performs the next instance of the evaluation (at 510) at 510. Conversely, in the case where it is determined at 510 that there is some pending resource access request, method 500 (at 512) identifies the access request holding the highest credentials, i.e., the access request from the processing unit that currently has the largest amount of credit attributed to it among all current pool members.
[0069] Method 500 then performs an evaluation (at 514) to determine whether any resource access requests are currently being served, for example, by determining whether there is some task currently being executed on the computing circuit resources on behalf of a pool member. In the case where it is determined at 514 that no such request is currently being served, method 500 (at 520) assigns a task to the computing circuit resources on behalf of the highest-credential request recently identified at 512.
[0070] Conversely, in the case where it is determined at 514 that such a request is currently being served, method 500 performs another evaluation (at 516) to determine whether the request currently being served is the highest-credential request recently identified at 512. In the case where it is determined at 516 that a request with lower credentials is currently being served, method 500 (at 518) pauses the task currently assigned to the resource, i.e., the task serving the request with lower credentials. Subsequently (at 520), method 500 assigns another task to the computing circuit resources on behalf of the highest-credential request recently identified at 512.
[0071] After the assignment at 520, or conversely, in the case where it is determined at 516 that the request holding the highest credential is currently being served, method 500 (at 522) performs another evaluation to determine whether the task most recently assigned to the resource has been completed. In the case where it is determined at 522 that the most recently assigned task has been completed, method 500 performs the next instance of the evaluation at 510. Conversely, in the case where it is determined at 522 that the most recently assigned task has not been completed, method 500 performs the next instance of the identification at 512.
[0072] Figure 6 System 600 for providing access to a computing circuit according to an embodiment is shown. System 600 illustrates the features of an exemplary embodiment for accessing shared circuit resources according to either a credit-based scheme or a priority-based scheme. In some embodiments, system 600 provides functions such as those of system 100 - for example, where the operations of one or more of methods 200, 300, 400, 500 utilize some or all of system 600 to be performed.
[0073] As Figure 6 shown, system 600 includes processor circuitry (such as the illustrative processor circuitry 610a, …, 610x shown) and computing circuitry 620, where the resources 622 of computing circuitry 620 are available for access by processor circuitry 610a, …, 610x at different times and differently. In one such embodiment, processor circuitry 610a, …, 610x (collectively also referred to herein as "processor circuitry 610" and each generally referred to as "processor circuitry 610") functionally corresponds to processor circuitry 110, for example, where resources 622 include the features of resources 122. In an illustrative scenario according to one embodiment, processor circuitry 610a and 610x (respectively) generate requests 611a and 611x for accessing resources 622, or otherwise contribute to the conveyance of requests 611a and 611x.
[0074] To promote fair distribution of access to resource 622 by processor circuitry 610, system 600 further includes a task assignment unit 630, a pool manager 640, and a credit manager 650, which provide functions such as those of computing circuitry 120, task assignment unit 130, pool manager 140, and credit manager 150, respectively, for example, in different ways. By way of illustration and not limitation, pool manager 640 includes pool status information 642, is coupled to access pool status information 642, or otherwise operates based on pool status information 642, which has some or all of the characteristics of pool status information 142, for example. Alternatively or additionally, credit manager 650 includes accumulation logic 652 and consumption logic 654, which provide functions such as those of accumulation logic 152 and consumption logic 154, respectively, for example, in different ways. Alternatively or additionally, credit scheme unit 631 of task assignment unit 630 includes selection logic 632 and allocation logic 634, which provide functions such as those of selection logic 132 and allocation logic 134, respectively, for example, in different ways. Task assignment unit 630 provides the function of differently assigning tasks to be executed using resource 622, where the tasks each serve to represent a respective one of one or more current requests 639 for accessing 622 on behalf of a corresponding processor circuitry 610.
[0075] In an embodiment, a given processor circuitry 610 (or other suitable logic of system 600) provides the function of identifying a corresponding access request as being associated with any one of a variety of possible priority levels. An access request given a relatively high priority ranking will be selected relative to another access request (if any) that is given a lower priority ranking or no priority ranking, for example. Such priority ranking of access requests is to be distinguished from, for example, the amount of credit attributed to the access request. For example, in some embodiments, when an access request is current, such amount of credit varies over time, while the priority ranking level (if any) given to the access request remains the same over time.
[0076] The priority level is determined based on, for example, the type of execution thread that generated the access request under discussion or is otherwise associated with the access request under discussion. Alternatively or additionally, the priority level is provided based on the quality of service to be provided by a given processor circuitry. However, some embodiments are not limited as to the particular basis on which a priority ranking level (if any) is associated with a given access request. The term "prioritized access request" (or, for brevity, "prioritized request") is used herein to refer to an access request associated with some level of priority ranking. The term "prioritized pool member" is used herein to refer to a pool member that generates a prioritized access request or otherwise contributes to the delivery of a prioritized access request.
[0077] In various embodiments, the task assignment unit 630 provides the function of switching between allocating access to the resource 622 according to a credit-based scheme and allocating access to the resource 622 according to a priority-based scheme. By way of illustration and not limitation, the task assignment unit 630 further includes a priority scheme unit 635, and the priority scheme unit 635 includes a selection logic 636 and an allocation logic 638.
[0078] The selection logic 636 provides the function of selecting one pool member as the next active pool member. For example, the selection is made from all (and, for example, only) the one or more currently prioritized pool members (if any). In an embodiment, selecting a given prioritized pool member includes selecting a prioritized access request from (or otherwise representing) that pool member. For example, this selection is performed based on the selection logic 636 searching for or otherwise accessing prioritization information (such as the prioritization information provided in the Pty field of the pool status information 642), which identifies the "highest-priority pool member", that is, the pool member currently assigned a higher priority level than any other pool member.
[0079] In one such embodiment, the selection logic 636 identifies the highest-priority pool member to the allocation logic 638 of the priority scheme unit 635. The allocation logic 638 (e.g., in combination with the allocation logic 634) tracks one or more current access requests 639. In response to the selection logic 636, the allocation logic 638 searches the (one or more) current requests 639 to identify the "highest-priority request", that is, the access request corresponding to the highest-priority pool member. In some cases where at least one of the current requests 639 is prioritized, the allocation logic 638 signals the task assignment unit 630 to participate in one or more communications (e.g., including the illustrated exemplary task assignment 621) to assign a task to be performed using the resource 622 to service the highest-priority access request. Conversely, in some alternative scenarios where none of the current requests 639 are prioritized, the allocation logic 634 instead signals the task assignment unit 630 to assign a task to be performed using the resource 622 to service the most authenticated access request.
[0080] In an illustrative scenario according to one embodiment, the first entry of the pool status information 642 includes: an identifier Pu1 of the processor circuit 610a, a member (Mem) flag set to specify that the processor circuit 610a is currently a pool member, an identifier n1 of the amount of credit currently attributed to the processor circuit 610a, a priority (Pty) value that does not identify any possible priority level, and an active (Ract) flag set to specify that the processor circuit 610a is currently a pending member. Further, the second entry of the pool status information 642 includes: an identifier Pux of the processor circuit 110x, a Mem flag set to specify that the processor circuit 610x is currently a pool member, an identifier nx of the amount of credit currently attributed to the processor circuit 610x, a priority (Pty) value indicating a third-level priority scheduling, and an Ract flag set to specify that the processor circuit 610x is currently a pending member. Still further, the third entry of the pool status information 642 includes: an identifier Pu2 of a third processor circuit (not shown), a Mem flag set to specify that the third processor circuit is currently a pool member, an identifier n2 of the amount of credit currently attributed to the third processor circuit, a priority (Pty) value indicating a fifth-level priority scheduling, and an Ract flag set to specify that the third processor circuit is currently an active member.
[0081] At the time in this scenario, task assignment using the priority scheme unit 635 (e.g., instead of using the credit scheme unit 631) is performed based on the determination by the task assignment unit 630 that at least one current access request is prioritized, e.g., where the determination is based on the Pty field of the entries in the pool status information 642. Since the third processor circuit is the pool member with the highest priority, and since the processor circuit 110a is not a prioritized pool member, the priority scheme unit 635 determines that the task assignment unit 630 will assign a first task to be executed on behalf of the third processor circuit to the resource 622.
[0082] Subsequently, at some other time after the first task is completed, the priority scheme unit 635 determines that the task assignment unit 630 will assign a second task to be executed on behalf of the processor circuit 110x to the resource 622 (i.e., when the processor circuit 110x is the pool member with the highest priority). Subsequently, at some other time after the first task is completed, the credit scheme unit 631 determines that the task assignment unit 630 will assign a second task to be executed on behalf of the processor circuit 110a to the resource 622 (i.e., when no other access requests are prioritized and when the processor circuit 110a is the highest-certified pool member).
[0083] Figure 7Illustrates a method 700 for selecting a scheme according to which computing circuit resources are to be accessed. Method 700 illustrates an example of an embodiment where the basis for task assignment is different from and prior to the basis for authentication for task assignment. Operations such as those of method 700 are performed using any combination of suitable hardware (e.g., circuitry), firmware, and / or software-executing, such as to provide some or all of the functionality of system 600. In some embodiments, method 200 includes some or all of the operations of method 700, or is otherwise performed in combination with some or all of the operations of method 700.
[0084] As Figure 7 shown, method 700 includes performing an evaluation (at 710) to determine whether any resource access requests are currently pending. If it is determined at 710 that no resource access requests are currently pending, method 700 performs the next instance of the determination at 710. Conversely, if it is determined at 710 that at least one resource access request is currently pending, method 700 performs another evaluation (at 712) to determine whether any of the currently pending requests are prioritized requests.
[0085] If it is determined at 712 that no currently pending requests are prioritized requests, method 700 performs task assignment (at 716) according to a credit-based scheme. By way of illustration and not limitation, the task assignment performed at 716 includes some or all of the features of method 300, e.g., where at least one task is assigned based on the amount of credit currently attributed to the corresponding pool member. In one such embodiment, the assignment at 716 includes performing method 300 until the receipt of a new prioritized access request causes such execution to terminate or otherwise be at least temporarily suspended.
[0086] Conversely, if it is determined at 712 that at least one currently pending request is a prioritized request, method 700 performs task assignment (at 714) according to a priority-based scheme. By way of illustration and not limitation, the task assignment performed at 714 includes some or all of the features of method 800, e.g., where at least one task is assigned based on the priority level assigned to the corresponding access request by the pool member (or another suitable agent). In one such embodiment, the assignment at 714 includes performing method 800 until no other currently access requests are prioritized. After the priority-based task assignment at 714, or after the credit-based task assignment at 716, method 700 performs the next instance of the evaluation at 710.
[0087] In various embodiments, credit accumulation and / or credit consumption is changed due to a transformation between resource access allocation according to a credit-based scheme and resource access allocation according to a priority-based scheme. For example, during priority-based access allocation, credit accumulation is relatively better at mitigating the likelihood that a given prioritized access request will deplete its corresponding credit amount. By way of illustration and not limitation, during priority-based access allocation (e.g., at 714), relatively high-priority access requests accumulate credit at a greater rate than relatively low-priority access requests. In one such embodiment, access requests with the same prioritization accumulate additional credit at the same rate. Alternatively or additionally, un-prioritized (and pending) access requests stop accumulating additional credit, or at least accumulate additional credit at a rate lower than any prioritized access request. Additionally or alternatively, prioritized access requests will be selected for service before un-prioritized access requests (and / or accumulate at a rate higher than un-prioritized access requests), but when a prioritized access request depletes its credit, service of the prioritized access request remains suspended.
[0088] Figure 8 Illustrated is a method 800 for assigning tasks to computing circuit resources based on priorities corresponding to processor circuits. Method 800 illustrates one example of an embodiment in which resource access requests (i.e., requests to access computing circuit resources shared by each member of a pool of processor circuits) are prioritized, which prioritization is the basis for assigning tasks to the shared resources. Operations such as those of method 800 are performed using any combination of suitable hardware (e.g., circuitry), firmware, and / or software executing, for example, some or all of the functionality provided by priority scheme unit 635 and / or other circuitry of system 600.
[0089] As Figure 8 shown, method 800 includes (at 810) performing an evaluation to determine whether at least one current resource access request is prioritized. In the case where it is determined at 810 that no current resource access request is prioritized, method 800 ends (e.g., where task assignment returns to executing on a credit-based scheme). Conversely, in the case where it is determined at 810 that at least one current resource access request is prioritized, method 800 (at 812) identifies the highest-priority request among one or more current prioritized requests.
[0090] In one such embodiment where multiple access requests each have the same priority scheduling, the identification at 812 includes selecting from such multiple access requests according to some other criteria, such as the amount of credit attributed to the corresponding pool member, the order in which the multiple access requests are received, and the like. After the identification at 812, method 800 performs an evaluation (at 814) to determine whether any requests representing pool members accessing the shared computing circuit resources are currently being served. In the case where it is determined at 814 that no such access requests are currently being served, method 800 (at 820) assigns a task to the shared computing circuit resources on behalf of the highest-priority request.
[0091] Conversely, in the case where it is determined at 814 that an access request is currently being served, method 800 performs another evaluation (at 816) to determine whether the currently served request is the highest-priority request. In the case where it is determined at 816 that the currently served request is a request other than the highest-priority request, method 800 (at 818) terminates or otherwise suspends the currently assigned task to the shared resources. After such a suspension, method 800 (at 820) assigns another task to the shared resources on behalf of the highest-priority request.
[0092] After the assignment at 820 (or conversely, in the case where it is determined at 816 that the currently served request is the highest-priority request), method 800 (at 822) performs another evaluation to determine whether the most recently assigned task has been completed. In the case where it is determined at 822 that the most recently assigned task has been completed, method 800 performs the next instance of the evaluation at 810. Conversely, in the case where it is determined at 822 that the most recently assigned task has not been completed, method 800 (at 812) performs the next instance of the identification at 812.
[0093] Figure 9 Illustrative exemplary system. Multiprocessor system 900 is a point-to-point interconnect system and includes multiple processors, which include a first processor 970 and a second processor 980 coupled via a point-to-point interconnect 950. In some examples, the first processor 970 and the second processor 980 are homogeneous. In some examples, the first processor 970 and the second processor 980 are heterogeneous. Although the exemplary system 900 is shown as having two processors, the system can have three or more processors, or can be a single-processor system.
[0094] Processors 970 and 980 are shown as including integrated memory controller (IMC) circuitry 972 and 982, respectively. Processor 970 also includes point-to-point (P-P) interfaces 976 and 978 as part of its interconnect controller; similarly, second processor 980 includes P-P interfaces 986 and 988. Processors 970, 980 may exchange information via point-to-point (P-P) interconnect 950 using point-to-point (P-P) interface circuits 978, 988. IMCs 972 and 982 couple processors 970, 980 to respective memories, namely memory 932 and memory 934, which may be portions of main memories locally attached to the respective processors.
[0095] Processors 970, 980 may each exchange information with chipset 990 via respective P-P interconnects 952, 954 using point-to-point interface circuits 976, 994, 986, and 998. Chipset 990 may optionally exchange information with coprocessor 938 via interface 992. In some examples, coprocessor 938 is a specialized processor, such as, for example, a high throughput processor, a network or communication processor, a compression engine, a graphics processor, a general purpose graphics processing unit (GPGPU), a neural-network processing unit (NPU), an embedded processor, and so on.
[0096] A shared cache (not shown) may be included in either processor 970, 980, or external to both processors but connected to these processors via P-P interconnects such that if a processor is placed in a low power mode, local cache information of either or both processors may be stored in the shared cache.
[0097] The chipset 990 can be coupled to a first interconnect 916 via an interface 996. In some examples, the first interconnect 916 can be a Peripheral Component Interconnect (PCI) interconnect or an interconnect such as a PCI Express interconnect or another I / O interconnect. In some examples, one of the interconnects is coupled to a power control unit (PCU) 917, which can include circuitry, software, and / or firmware for performing power management operations related to processors 970, 980, and / or coprocessor 938. The PCU 917 provides control information to a voltage regulator (not shown) to cause the voltage regulator to generate an appropriate regulated voltage. The PCU 917 also provides control information to control the generated operating voltage. In various examples, the PCU 917 can include various power management logic units (circuitry) for performing hardware-based power management. Such power management can be fully controlled by the processor (e.g., controlled by various processor hardware and can be triggered by workload and / or power, thermal constraints, or other processor constraints), and / or the power management can be performed in response to an external source such as a platform or a power management source or system software.
[0098] The PCU 917 is illustrated as existing as a logic separate from processors 970 and / or processor 980. In other cases, the PCU 917 can execute on a given one or more cores in the core (not shown) of processor 970 or 980. In some cases, the PCU 917 can be implemented as a (dedicated or general-purpose) microcontroller or other control logic configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other examples, the power management operations to be performed by the PCU 917 can be implemented external to the processor, such as by a separate power management integrated circuit (PMIC) or another component external to the processor. In still other examples, the power management operations to be performed by the PCU 917 can be implemented within the BIOS or other system software.
[0099] A variety of I / O devices 914 may be coupled to a first interconnect 916 along with a bus bridge 918 that couples the first interconnect 916 to a second interconnect 920. In some examples, one or more additional processors 915 (such as, a coprocessor, a high throughput many integrated core (MIC) processor, a GPGPU, an accelerator (such as a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor) are coupled to the first interconnect 916. In some examples, the second interconnect 920 may be a low pin count (LPC) interconnect. A variety of devices may be coupled to the second interconnect 920, including, for example, a keyboard and / or mouse 922, a communication device 927, and a storage circuitry 928. The storage circuitry 928 may be one or more non-transitory machine-readable storage media as described below, such as a disk drive or other mass storage device that may include instructions / code and data 930 in some examples. Further, audio I / O 924 may be coupled to the second interconnect 920. Note that other architectures are possible in addition to the point-to-point architecture described above. For example, instead of a point-to-point architecture, a system such as the multiprocessor system 900 may implement a multi-branch interconnect or other such architectures. Exemplary core architectures, processors, and computer architectures.
[0100] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores can include: 1) general-purpose in-order cores intended for general computing; 2) high-performance general-purpose out-of-order cores intended for general computing; 3) specialized cores intended primarily for graphics and / or scientific (throughput) computing. Implementations of different processors can include: 1) a CPU that includes one or more general-purpose in-order cores intended for general computing and / or one or more general-purpose out-of-order cores intended for general computing; and 2) a coprocessor that includes one or more specialized cores intended primarily for graphics and / or scientific (throughput) computing. Such different processors give rise to different computer system architectures, which can include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor in the same package as the CPU but on a separate die; 3) a coprocessor on the same die as the CPU (in which case such a coprocessor is sometimes referred to as specialized logic or as a specialized core, such as integrated graphics and / or scientific (throughput) logic); and 4) a system on a chip (SoC) that can include the described CPU (sometimes referred to as the (one or more) application core or (one or more) application processor), the above-described coprocessor, and additional functionality on the same die. Exemplary core architectures are then described, followed by exemplary processors and computer architectures.
[0101] Figure 10 FIG. shows a block diagram of an example processor 1000 that can have more than one core and can have an integrated memory controller. The solid block diagram shows a processor 1000 having a single core 1002A, a system agent unit circuitry 1010, and a set of one or more interconnect controller unit circuitries 1016, while the optional addition of the dashed block diagram shows an alternative processor 1000 having a plurality of cores 1002A - 1002N, a set of one or more integrated memory controller units 1014 in the system agent unit circuitry 1010, and specialized logic 1008 and a set of one or more interconnect controller unit circuitries 1016. Note that the processor 1000 can be Figure 9 one of the processors 970 or 980, or coprocessors 938 or 915.
[0102] Accordingly, different implementations of the processor 1000 may include: 1) a CPU, where the dedicated logic 1008 is an integrated graphics device and / or scientific (throughput) logic (which may include one or more cores, not shown), and the cores 1,002A - 1302N are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where the cores 1,002A - 1302N are a large number of dedicated cores designed primarily for graphics and / or scientific (throughput); and 3) a coprocessor, where the cores 1,002A - 1302N are a large number of general-purpose in-order cores. Accordingly, the processor 1000 may be a general-purpose processor, a coprocessor, or a special-purpose processor, such as, for example, a network or communication processor, a compression engine, a graphics processor, a GPGPU (general purpose graphics processing unit circuitry), a high-throughput many integrated core (MIC) coprocessor (including 30 or more cores), an embedded processor, etc. The processor may be implemented on one or more chips. The processor 1000 may be part of one or more substrates and / or implemented on one or more substrates using any of a variety of process technologies, such as, for example, complementary metal-oxide-semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).
[0103] The memory hierarchy includes one or more levels of (one or more) cache unit circuitry 1004A - 1004N within cores 1002A - 1002N, a collection of one or more shared cache unit circuitry 1006, and external memory (not shown) coupled to a collection of (one or more) integrated memory controller unit circuitry 1014. The collection of one or more shared cache unit circuitry 1006 may include one or more intermediate levels of cache (such as second level (L2), third level (L3), fourth level (L4)) or other levels of cache (such as, last level cache (LLC)) and / or combinations of the foregoing. Although in some examples, a ring - based interconnect network circuitry 1012 interconnects dedicated logic 1008 (e.g., integrated graphics logic), the collection of (one or more) shared cache unit circuitry 1006, and system agent unit circuitry 1010, alternative examples use any number of well - known techniques for interconnecting such units. In some examples, coherence is maintained between the collection of (one or more) shared cache unit circuitry 1006 and one or more of cores 1002A - 1002N.
[0104] In some examples, one or more of cores 1002A - 1002N are capable of implementing multithreaded operation. The system agent unit circuitry 1010 includes those components that coordinate and operate cores 1002A - 1002N. The system agent unit circuitry 1010 may include, for example, power control unit (PCU) circuitry and / or display unit circuitry (not shown). The PCU may be or may include the logic and components required to regulate the power states of cores 1002A - 1002N and / or dedicated logic 1008 (e.g., integrated graphics logic). The display unit circuitry is used to drive one or more externally connected displays.
[0105] Cores 1002A - 1002N may be homogeneous in terms of instruction set architecture (ISA). Alternatively, cores 1002A - 1002N may be heterogeneous in terms of ISA; that is, a subset of cores 1002A - 1002N may be capable of executing an ISA, while other cores may be capable of executing only a subset of that ISA or another ISA. Exemplary core architecture - in-order and out-of-order core block diagrams.
[0106] Figure 11A is a block diagram illustrating an exemplary in - order pipeline and an exemplary register renaming, out - of - order issue / execution pipeline according to an example. Figure 11BFIG. 0 is a block diagram illustrating both an exemplary in-order architecture core to be included in a processor, and an exemplary register renaming, out-of-order issue / execution architecture core according to an example. Figures 11A - 11B The solid boxes in FIG. 2 illustrate an in-order pipeline and an in-order core, while the optional additions of the dashed boxes illustrate register renaming, out-of-order issue / execution pipelines and cores. Given that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.
[0107] In Figure 11A FIG. 7, processor pipeline 1100 includes a fetch stage 1102, an optional length decoding stage 1104, a decode stage 1106, an optional allocation (Alloc) stage 1108, an optional rename stage 1110, a schedule (also known as dispatch or issue) stage 1112, an optional register read / memory read stage 1114, an execution stage 1116, a write-back / memory write stage 1118, an optional exception handling stage 1122, and an optional commit stage 1124. One or more operations may be performed in each of these processor pipeline stages. For example, during the fetch stage 1102, one or more instructions are fetched from an instruction memory, and during the decode stage 1106, the one or more fetched instructions may be decoded, an address (e.g., a load store unit (LSU) address) using the forwarded register ports may be generated, and branch forwarding (e.g., immediate offset or link register (LR)) may be performed. In one example, the decode stage 1106 and the register read / memory read stage 1114 may be combined into one pipeline stage. In one example, during the execution stage 1116, the decoded instructions may be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AMB) interface may be performed, multiplication and addition operations may be performed, arithmetic operations with branch results may be performed, and so on.
[0108] As an example, Figure 11BAn exemplary register renaming, out-of-order issue / execution architecture core may implement pipeline 1100 as follows: 1) Instruction fetch circuitry 1138 performs the fetch stage 1102 and the length decoding stage 1104; 2) Decoding circuitry 1140 performs the decode stage 1106; 3) The rename / allocator unit circuitry 1152 performs the allocation stage 1108 and the rename stage 1110; 4) The (one or more) scheduler circuitry 1156 performs the schedule stage 1112; 5) The (one or more) physical register file circuitry 1158 and the memory unit circuitry 1170 perform the register read / memory read stage 1114; The (one or more) execution clusters 1160 perform the execution stage 1116; 6) The memory unit circuitry 1170 and the (one or more) physical register file circuitry 1158 perform the writeback / memory write stage 1118; 7) Various circuitries may be involved in the exception handling stage 1122; and 8) The retirement unit circuitry 1154 and the (one or more) physical register file circuitry 1158 perform the commit stage 1124.
[0109] Figure 11B Illustrated is a processor core 1190 that includes a front end unit circuitry 1130 coupled to an execution engine unit circuitry 1150, and both are coupled to a memory unit circuitry 1170. The core 1190 may be a reduced instruction set architecture computing (RISC) core, a complex instruction set architecture computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 1190 may be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general purpose graphics processing unit (GPGPU) core, a graphics core, and so on.
[0110] The front-end unit circuit system 1130 may include a branch prediction circuit system 1132 coupled to an instruction cache circuit system 1134, the instruction cache circuit system 1134 being coupled to an instruction translation lookaside buffer (TLB) 1136, the instruction translation lookaside buffer 1136 being coupled to an instruction fetch circuit system 1138, and the instruction fetch circuit system 1138 being coupled to a decoding circuit system 1140. In one example, the instruction cache circuit system 1134 is included in the memory unit circuit system 1170 rather than in the front-end circuit system 1130. The decoding circuit system 1140 (or decoder) may decode the instruction and generate, as output, one or more micro-operations, micro-code entry points, micro-instructions, other instructions, or other control signals that are decoded from, otherwise reflect, or are derived from the original instruction. The decoding circuit system 1140 may further include address generation unit (AGU, not shown) circuitry. In one example, the AGU generates an LSU address using the forwarded register ports and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). The decoding circuit system 1140 may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, look-up tables, hardware implementations, programmable logic arrays (PLAs), micro-code read only memories (ROMs), etc. In one example, the core 1190 includes a micro-code ROM (not shown) or other medium (e.g., in the decoding circuit system 1140 or otherwise within the front-end circuit system 1130) that stores micro-code for certain macro-instructions. In one example, the decoding circuit system 1140 includes a micro-operation (micro-op) or operation cache (not shown) to save / cache the decoded operations, micro-tags, or micro-operations generated during the decoding stage or other stages of the processor pipeline 1100. The decoding circuit system 1140 may be coupled to a rename / allocator unit circuit system 1152 in the execution engine circuit system 1150.
[0111] The execution engine circuitry 1150 includes a rename / allocator unit circuitry 1152 that is coupled to a retirement unit circuitry 1154 and a set 1156 of one or more scheduler circuitries. The one or more scheduler circuitries 1156 represent any number of different schedulers, including reservation stations, a central instruction window, and the like. In some examples, the one or more scheduler circuitries 1156 may include an arithmetic logic unit (ALU) scheduler / scheduling circuitry, an ALU queue, an arithmetic generation unit (AGU) scheduler / scheduling circuitry, an AGU queue, and so on. The one or more scheduler circuitries 1156 are coupled to the one or more physical register file circuitries 1158. Each of the one or more physical register file circuitries 1158 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point, status (e.g., an instruction pointer that is the address of the next instruction to be executed), and so on. In one example, the one or more physical register file circuitries 1158 include a vector register unit circuitry, a write mask register unit circuitry, and a scalar register unit circuitry. These register units may provide architectural vector registers, vector mask registers, general-purpose registers, and so on. The one or more physical register file circuitries 1158 are coupled to the retirement unit circuitry 1154 (also referred to as a retirement queue (“retire queue” or “retirement queue”)) to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., using one or more reorder buffers (ROBs) and one or more retirement register files; using one or more future files, one or more history buffers, and one or more retirement register files; using register maps and register pools, and so on). The retirement unit circuitry 1154 and the one or more physical register file circuitries 1158 are coupled to the one or more execution clusters 1160. The one or more execution clusters 1160 include a set of one or more execution unit circuitries 1162 and a set of one or more memory access circuitries 1164. The one or more execution unit circuitries 1162 may perform various arithmetic, logical, floating-point, or other types of operations (e.g., shift, add, subtract, multiply) and may operate on various data types (e.g., scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point).While some examples may include multiple execution units or execution unit circuitry dedicated to a particular function or set of functions, other examples may include only one execution unit circuitry or multiple execution units / execution unit circuitry that all perform all functions. The (one or more) scheduler circuitry 1156, the (one or more) physical register file circuitry 1158, and the (one or more) execution clusters 1160 are shown as potentially being multiple because some examples create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / tight integer / tight floating point / vector integer / vector floating point pipelines, and / or memory access pipelines each having their own scheduler circuitry, (one or more) physical register file circuitry, and / or execution clusters—and in the case of separate memory access pipelines, some examples where only the execution cluster of that pipeline has the (one or more) memory access circuitry 1164). It should also be understood that in the case of using separate pipelines, one or more of these pipelines may be out-of-order issue / execution, and the remaining pipelines may be in-order issue / execution.
[0112] In some examples, the execution engine unit circuitry 1150 may perform load store unit (LSU) address / data pipelining to an Advanced Microcontroller Bus (AMB) interface (not shown), as well as address staging and write-back, data staging load, store, and branch.
[0113] A set of memory access circuitry 1164 is coupled to a memory unit circuitry 1170 that includes data TLB circuitry 1172, which is coupled to a data cache circuitry 1174, which is coupled to a second-level (L2) cache circuitry 1176. In one exemplary example, the memory access circuitry 1164 may include a load unit circuitry, a store address unit circuit, and a store data unit circuitry, each of which is coupled to the data TLB circuitry 1172 in the memory unit circuitry 1170. The instruction cache circuitry 1134 is further coupled to the second-level (L2) cache circuitry 1176 in the memory unit circuitry 1170. In one example, the instruction cache 1134 and the data cache 1174 are combined into a single instruction and data cache (not shown) in the L2 cache circuitry 1176, a third-level (L3) cache circuitry (not shown), and / or main memory. The L2 cache circuitry 1176 is coupled to one or more other levels of cache and ultimately to main memory.
[0114] The core 1190 can support one or more instruction sets (e.g., the x86 instruction set architecture (optionally with some extensions added with newer versions); the MIPS instruction set architecture; the ARM instruction set architecture (optionally with optional additional extensions such as NEON)), which includes the instruction(s) described herein. In one example, the core 1190 includes logic for supporting a SIMD instruction set architecture extension (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using SIMD data.
[0115] The description herein includes numerous details to provide a more thorough explanation of embodiments of the present disclosure. However, it will be apparent to those skilled in the art that embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the present disclosure.
[0116] Note that in the corresponding drawings of the embodiments, lines are used to represent signals. Some lines may be thicker to indicate a greater number of component signal paths, and / or have arrows at one or more ends to indicate the direction of information flow. Such indications are not intended to be restrictive. Instead, the lines are used in conjunction with one or more exemplary embodiments to facilitate a more easily understood of the circuit or logic unit. As dictated by design requirements or preferences, any represented signal may actually include one or more signals that may travel in either direction and may be implemented using any suitable type of signal scheme.
[0117] Throughout the specification and in the claims, the term "connected" means a direct connection between the connected objects such as an electrical, mechanical, or magnetic connection without any intermediate devices. The term "coupled" means a direct or indirect connection, such as a direct electrical, mechanical, or magnetic connection between the connected objects or an indirect connection through one or more passive or active intermediate devices. The term "circuit" or "module" may refer to one or more passive and / or active components arranged to cooperate with each other to provide a desired function. The term "signal" may refer to at least one current signal, voltage signal, magnetic signal, or data / clock signal. The meaning of "a / an" and "the" includes plural references. The meaning of "in" includes "in" and "on".
[0118] The term "device" generally can refer to an apparatus according to the context in which that term is used. For example, a device can refer to a stack of layers or structures, a single structure or layer, a connection of various structures having active and / or passive elements, and so on. Generally speaking, a device is a three-dimensional structure having a plane in the x-y direction along the x-y Cartesian coordinate system and a height in the z direction. The plane of the device can also be the plane of the apparatus that includes the device.
[0119] The term "scaling" generally refers to the conversion of a design (schematic and layout) from one process technology to another and subsequent reduction in the layout area. The term "scaling" generally also refers to the reduction in size of a layout and a device within the same technology node. The term "scaling" can also refer to the adjustment (e.g., deceleration or acceleration - i.e., reduction or magnification, respectively) of a signal frequency relative to another parameter (e.g., power supply level).
[0120] The terms "substantially", "close to", "approximate", "near", and "about" generally refer to being within + / - 10% of a target value. For example, unless otherwise specified in the explicit context in which it is used, the terms "substantially equal", "about equal", and "approximately equal" mean that there are only incidental variations between the objects so described. In the art, such variations typically are not greater than + / - 10% of a predetermined target value.
[0121] It should be understood that the terms so used are interchangeable where appropriate, such that embodiments of the invention described herein can operate in other orientations different from those illustrated or otherwise described herein.
[0122] Unless otherwise specified, the use of the ordinal adjectives "first", "second", "third", etc. to describe a common object merely indicates that different instances of like objects are being referred to and is not intended to imply that the objects so described must be in a given order in time, space, ranking, or in any other way.
[0123] The terms "left", "right", "front", "back", "top", "bottom", "above", "below", etc. (if any) in the specification and claims are used for descriptive purposes and are not necessarily used to describe permanent relative positions. For example, as used herein, the terms "above", "below", "front side", "back side", "top", "bottom", "above", "below", and "on" refer to the relative position of one component, structure, or material with respect to other referenced components, structures, or materials within the device, where such physical relationships are significant. These terms are adopted herein solely for descriptive purposes and are primarily within the context of the z-axis of the device, and thus these terms can be relative to the orientation of the device. Thus, the first material "above" the second material in the context of the figures provided herein can also be "below" the second material when the device is oriented upside down relative to the context of the figures provided. In the context of materials, a material disposed above or below another material can be in direct contact or can have one or more intervening materials. Additionally, a material disposed between two materials can be in direct contact with both of those layers or can have one or more intervening layers. In contrast, the first material "on" the second material is in direct contact with the second material. Similar distinctions are made in the context of component assemblies.
[0124] The term "between" can be employed in the context of the z-axis, x-axis, or y-axis of the device. A material between two other materials can be in contact with one or both of those two materials, or the material can be separated from both of the other two materials by one or more intervening materials. Thus, a material "between" two other materials can be in contact with either of the other two materials, or the material can be coupled to the other two materials through intervening materials. A device between two other devices can be directly connected to one or both of those two devices, or the device can be separated from both of the other two devices by one or more intervening devices.
[0125] As used throughout the specification and in the claims, a list of items joined by the terms "at least one of..." or "one or more of..." can mean any combination of the listed items. For example, the phrase "at least one of A, B, or C" can mean A; B; C; A and B; A and C; B and C; or A, B, and C. It should be noted that those elements of the figures having the same reference numerals (or names) as elements of any other figure can operate or function in any manner similar to the manner described, but are not limited thereto.
[0126] In addition, the various elements of combinational and sequential logic discussed in this disclosure may relate to physical structures (such as AND gates, OR gates, or XOR gates), or to synthesized or otherwise optimized collections of devices implementing a logical structure that is a Boolean equivalent of the discussed logic.
[0127] Techniques and architectures for providing access to shared resources of a computing circuit are described herein. In the foregoing description, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of certain embodiments. However, it will be apparent to those skilled in the art that some embodiments may be practiced without these specific details. In other instances, structures and devices are shown in block diagram form to avoid obscuring the description.
[0128] References to "an embodiment" or "embodiments" in the specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. The appearances of the phrase "in an embodiment" in various places in the specification are not necessarily all referring to the same embodiment.
[0129] Some portions of the detailed descriptions herein are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the computer art to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. These steps are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. For the sake of common usage, it has proven convenient at times to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.
[0130] However, it should be borne in mind that all such and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless stated otherwise explicitly, as will be apparent from the discussion herein, it is to be appreciated that throughout the specification, discussions using terms such as "processing" or "computing" or "operating" or "determining" or "displaying" etc. refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities within the registers and memories of the computer system and transforms it into other data similarly represented as physical quantities within the computer system memory or registers or other such information storage, transmission, or display devices.
[0131] Certain embodiments also relate to apparatus for performing the operations herein. The apparatus can be specially constructed for the required purposes, or it can comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such computer programs can be stored in computer-readable storage media, such as but not limited to any type of disk, including floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM) (such as dynamic RAM (DRAM)), EPROM, EEPROM, magnetic or optical cards, or any type of media suitable for storing electronic instructions and coupled to a computer system bus.
[0132] In one or more first embodiments, an apparatus includes: a first circuitry for determining a size of a pool of one or more processor circuits, each of the one or more processor circuits corresponding to a respective current request to access resources of an access computing circuit; a second circuitry coupled to the first circuitry, the second circuitry for: detecting a first condition where a first request to access resources is pending, where the first request represents a first pool member; based on the first condition, increasing a first credit amount by a first ratio based on the size, where the first credit amount corresponds to the first pool member; detecting a second condition where a second request to access resources of the access computing circuit is being serviced, where the second request represents a second pool member; based on the second condition, decreasing a second credit amount at a second rate, where the second credit amount corresponds to the second pool member; a third circuitry for performing a selection of the first pool member based on the first credit amount and the second credit amount; and a fourth circuitry for assigning a first task to the access computing circuit resources based on the selection representing the first pool member.
[0133] In one or more second embodiments, further to the first embodiment, the second rate is independent of the size of the pool.
[0134] In one or more third embodiments, further to the first embodiment or the second embodiment, the second circuitry is further for increasing the first credit amount and the second credit amount at the first rate respectively when both the first request and the second request are pending.
[0135] In one or more fourth embodiments, further to any one of the first embodiment to the third embodiment, the first circuitry is further for detecting an increase in the size of the pool, and based on the first condition, the second circuitry is for increasing the first credit amount at a third rate less than the first rate, where the third rate is based on the size of the pool after the increase.
[0136] In one or more fifth embodiments, further to any one of the first to fourth embodiments, the first circuit system is further configured to detect a reduction in the size of the pool, and based on a first condition, the second circuit system is configured to increase a first credit amount at a third rate that is greater than the first rate, where the third rate is based on the size of the pool after the reduction.
[0137] In one or more sixth embodiments, further to any one of the first to fifth embodiments, the computing circuit includes one of a coprocessor, an accelerator, or a graphics processing unit.
[0138] In one or more seventh embodiments, further to any one of the first to sixth embodiments, allocating access to resources according to a credit-based allocation scheme includes selection of a first pool member; a third circuit system is further configured to rank a priority level as corresponding to a third request from a third pool member, and based on the ranked priority level, the third circuit system is configured to perform a transformation from the credit-based allocation scheme to a priority-based allocation scheme.
[0139] In one or more eighth embodiments, further to the seventh embodiment, the third circuit system is further configured to detect completion of the prioritized request to access the resource, and based on the completion, the third circuit system is further configured to perform another transformation from the priority-based allocation scheme to the credit-based allocation scheme.
[0140] In one or more ninth embodiments, further to the seventh embodiment, after the transformation, the second circuit system is configured to increase a third credit amount at a rate based on the ranked priority level, where the third credit amount corresponds to the third pool member.
[0141] In one or more tenth embodiments, a method includes: determining the size of a pool of one or more processor circuits, each of the one or more processor circuits corresponding to a respective current request to access resources of a computing circuit; detecting a first condition where a first request to access the resources is pending, where the first request represents a first pool member; based on the first condition, increasing a first credit amount at a first ratio based on the size, where the first credit amount corresponds to the first pool member; detecting a second condition where a second request to access the computing circuit resources is being serviced, where the second request represents a second pool member; based on the second condition, reducing a second credit amount at a second rate, where the second credit amount corresponds to the second pool member; performing selection of the first pool member based on the first credit amount and the second credit amount; and assigning a first task to the computing circuit resources based on the selection representing the first pool member.
[0142] In one or more eleventh embodiments, further to the tenth embodiment, the second rate is independent of the size of the pool.
[0143] In one or more twelfth embodiments, further to the tenth or eleventh embodiment, the method further includes increasing a first credit amount and a second credit amount at a first rate respectively while both the first request and the second request are pending.
[0144] In one or more thirteenth embodiments, further to any one of the tenth to twelfth embodiments, the method further includes detecting an increase in the size of the pool, and increasing the first credit amount at a third rate less than the first rate based on a first condition, wherein the third rate is based on the size of the pool after the increase.
[0145] In one or more fourteenth embodiments, further to any one of the tenth to thirteenth embodiments, the method further includes detecting a decrease in the size of the pool, and increasing the first credit amount at a third rate greater than the first rate based on a first condition, wherein the third rate is based on the size of the pool after the decrease.
[0146] In one or more fifteenth embodiments, further to any one of the tenth to fourteenth embodiments, the computing circuit includes one of a coprocessor, an accelerator, or a graphics processing unit.
[0147] In one or more sixteenth embodiments, further to any one of the tenth to fifteenth embodiments, allocating access to resources according to a credit-based allocation scheme includes selection of a first pool member; and the method further includes designating a priority ranking identifier as corresponding to a third request from a third pool member, and performing a transformation from the credit-based allocation scheme to a priority-based allocation scheme based on the priority ranking.
[0148] In one or more seventeenth embodiments, further to the sixteenth embodiment, the method further includes detecting completion of the prioritized request to access the resources, and performing another transformation from the priority-based allocation scheme to the credit-based allocation scheme based on the completion.
[0149] In one or more eighteenth embodiments, further to the sixteenth embodiment, after the transformation, the accumulation of the credit of the third pool member is at a rate based on the priority ranking.
[0150] In one or more nineteenth embodiments, a system includes: a plurality of processor circuits; a computing circuit; a first circuitry coupled to the plurality of processor circuits and the computing circuit, the first circuitry for determining the size of a pool of one or more processor circuits, each of the one or more processor circuits corresponding to a respective current request to access a resource of the computing circuit; and a second circuitry coupled to the first circuitry, the second circuitry for: detecting a first condition, wherein a first request to access a resource is pending, and wherein the first request represents a first pool member; based on the first condition, increasing a first credit amount at a first ratio based on size, wherein the first credit amount corresponds to the first pool member; detecting a second condition, wherein a second request to access a computing circuit resource is being serviced, and wherein the second request represents a second pool member; based on the second condition, decreasing a second credit amount at a second rate, wherein the second credit amount corresponds to the second pool member.
[0151] In one or more twentieth embodiments, further to the nineteenth embodiment, each of the plurality of processor circuits includes a respective one or more processor cores.
[0152] In one or more twenty-first embodiments, further to the nineteenth or twentieth embodiment, the second rate is independent of the size of the pool.
[0153] In one or more twenty-second embodiments, further to any one of the nineteenth to twenty-first embodiments, the second circuitry is further for increasing the first credit amount and the second credit amount at the first rate when both the first request and the second request are pending.
[0154] In one or more twenty-third embodiments, further to any one of the nineteenth to twenty-second embodiments, the first circuitry is further for detecting an increase in the size of the pool, and based on the first condition, the second circuitry is for increasing the first credit amount at a third rate less than the first rate, wherein the third rate is based on the size of the pool after the increase.
[0155] In one or more twenty-fourth embodiments, further to any one of the nineteenth to twenty-third embodiments, the first circuitry is further for detecting a decrease in the size of the pool, and based on the first condition, the second circuitry is for increasing the first credit amount at a third rate greater than the first rate, wherein the third rate is based on the size of the pool after the decrease.
[0156] In one or more twenty-fifth embodiments, further to any one of the nineteenth to twenty-fourth embodiments, the computing circuit includes one of a coprocessor, an accelerator, or a graphics processing unit.
[0157] In one or more twenty-sixth embodiments, further to any one of the nineteenth to twenty-fifth embodiments, the system further includes a third circuitry for performing a selection of a first pool member based on a first credit amount and a second credit amount; and a fourth circuitry for assigning a first task to the computing circuit resources on behalf of the first pool member based on the selection.
[0158] In one or more twenty-seventh embodiments, further to the twenty-sixth embodiment, allocating access to resources according to a credit-based allocation scheme includes a selection of a first pool member; the third circuitry is further configured to rank a priority ranking identifier as corresponding to a third request from a third pool member, and based on the priority ranking, the third circuitry is configured to perform a transformation from the credit-based allocation scheme to a priority-based allocation scheme.
[0159] In one or more twenty-eighth embodiments, further to the twenty-seventh embodiment, the third circuitry is further configured to detect the completion of a prioritized request for accessing a resource, and based on the completion, the third circuitry is further configured to perform another transformation from the priority-based allocation scheme to the credit-based allocation scheme.
[0160] In one or more twenty-ninth embodiments, further to the twenty-seventh embodiment, after the transformation, the second circuitry is configured to increase a third credit amount at a rate based on the priority ranking, wherein the third credit amount corresponds to the third pool member.
[0161] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for various such systems will be presented from the description herein. Additionally, certain embodiments are not described with reference to any particular programming language. It will be appreciated that various programming languages may be used to implement the teachings of such embodiments described herein.
[0162] Except for what is described herein, various modifications may be made to the disclosed embodiments and their implementations without departing from their scope. Accordingly, the descriptions and examples herein should be construed as illustrative and not restrictive. The scope of the invention should be defined only by reference to the appended claims.
Claims
1. An apparatus for determining access to a shared circuit resource, the apparatus comprising: first circuitry for determining a size of a pool of one or more processor circuits that each correspond to a respective current request for access to a resource of a computing circuit; A second circuit system is coupled to the first circuit system, and the second circuit system is configured to: detecting a first condition wherein a first request to access the resource is pending, wherein the first request is on behalf of a first pool member; based on the first condition, increasing a first credit amount at a first ratio based on the size, wherein the first credit amount corresponds to the first pool member; detecting a second condition wherein a second request to access the computing circuit resource is being serviced, wherein the second request is on behalf of a second pool member; based on the second condition, reducing a second credit at a second rate, wherein the second credit corresponds to the second pool member; third circuitry for performing selection of the first pool member based on the first credit amount and the second credit amount; as well as A fourth circuit system is configured to assign a first task to the computing circuit resource on behalf of the first pool member based on the selection.
2. The device according to claim 1, wherein The second rate is independent of the size of the pool.
3. The device according to any one of claims 1 or 2, wherein: The second circuit system is further configured to increase the first credit and the second credit at the first rate, respectively, while both the first request and the second request are pending.
4. The apparatus according to any one of claims 1 to 3, wherein: The first circuitry is further configured to detect an increase in the size of the pool; and Based on the first condition, the second circuit system is used to increase the first amount of credit at a third rate that is less than the first rate, wherein the third rate is based on a size of the pool after the increase.
5. The apparatus according to any one of claims 1 to 3, wherein: The first circuit system is further configured to detect a decrease in the size of the pool; and Based on the first condition, the second circuit system is used to increase the first amount of credit at a third rate greater than the first rate, wherein the third rate is based on a size of the pool after the reduction.
6. The device according to any one of claims 1 to 3, wherein: The computing circuit includes one of a coprocessor, an accelerator, or a graphics processor unit.
7. The apparatus according to any one of claims 1 to 3, wherein: allocating access to said resource according to a credit-based allocation scheme comprises said selection of said first pool member; The third circuitry is further operable to identify a prioritization level as corresponding to a third request from a third pool member; and Based on the prioritization level, the third circuitry is operable to perform a transition from the credit-based allocation scheme to a priority-based allocation scheme.
8. The apparatus of claim 7, wherein: The third circuitry is further for detecting completion of a prioritized request to access the resource; and Based on the completion, the third circuit system is further for performing another transition from the priority-based allocation scheme to the credit-based allocation scheme.
9. The device according to claim 7, wherein: After the transitioning, the second circuit system is to increase a third amount of credit at a rate based on the prioritization level, wherein the third amount of credit corresponds to the third pool member.
10. A method for determining access to a shared circuit resource, the method comprising: determining a size of a pool of one or more processor circuits that each correspond to a respective current request for access to a resource of the computing circuit; detecting a first condition wherein a first request to access the resource is pending, wherein the first request is on behalf of a first pool member; based on the first condition, increasing a first credit amount at a first ratio based on the size, wherein the first credit amount corresponds to the first pool member; detecting a second condition wherein a second request to access the computing circuit resource is being serviced, wherein the second request is on behalf of a second pool member; based on a second condition, reducing a second credit at a second rate, wherein the second credit corresponds to the second pool member; performing selection of the first pool member based on the first credit amount and the second credit amount; as well as A first task is assigned to the computing circuit resource on behalf of the first pool member based on the selection.
11. The method according to claim 10, wherein: The second rate is independent of the size of the pool.
12. The method according to any one of claims 10 or 11, further comprising: While both the first request and the second request are pending, the first credit and the second credit are each increased at the first rate.
13. The method according to any one of claims 10 to 12, further comprising: detecting an increase in the size of the pool; as well as Based on the first condition, the first amount of credit is increased at a third rate that is less than the first rate, wherein the third rate is based on a size of the pool after the increase.
14. The method according to any one of claims 10 to 12, further comprising: detecting a decrease in the size of the pool; as well as Based on the first condition, the first amount of credit is increased at a third rate that is greater than the first rate, wherein the third rate is based on a size of the pool after the reduction.
15. The method according to any one of claims 10 to 12, wherein: The computing circuit includes one of a coprocessor, an accelerator, or a graphics processor unit.
16. The method according to any one of claims 10 to 12, wherein: allocating access to said resource according to a credit-based allocation scheme comprises said selecting of said first pool member; and The method further comprises: identifying a prioritization level as corresponding to a third request from a third pool member; and Based on the prioritization level, a transition from the credit-based allocation scheme to a priority-based allocation scheme is performed.
17. The method according to claim 16, further comprising: detecting completion of a prioritized request to access the resource; as well as Based on the completion, another transition from the priority-based allocation scheme to the credit-based allocation scheme is performed.
18. The method according to claim 16, wherein: After the transition, accumulation of credits for the third pool member is at a rate based on the prioritization level.
19. A system for determining access to a shared circuit resource, the system comprising: a plurality of processor circuits; Computational circuits; a first circuit system coupled to the plurality of processor circuits and the computation circuit, the first circuit system to determine a size of a pool of one or more processor circuits, the one or more processor circuits each corresponding to a respective current request to access a resource of the computation circuit; as well as A second circuit system is coupled to the first circuit system, and the second circuit system is configured to: detecting a first condition wherein a first request to access the resource is pending, wherein the first request is on behalf of a first pool member; based on the first condition, increasing a first credit amount at a first ratio based on the size, wherein the first credit amount corresponds to the first pool member; detecting a second condition wherein a second request to access the computing circuit resource is being serviced, wherein the second request is on behalf of a second pool member; Based on the second condition, a second credit amount is reduced at a second rate, wherein the second credit amount corresponds to the second pool member.
20. The system of claim 19, wherein: The plurality of processor circuits each include a respective one or more processor cores.
21. A system according to any one of claims 19 or 20, wherein: The second rate is independent of the size of the pool.
22. A system according to any one of claims 19 to 21, wherein: The second circuit system is further configured to increase the first credit and the second credit at the first rate, respectively, while both the first request and the second request are pending.
23. A system according to any one of claims 19 to 22, wherein: The first circuitry is further configured to detect an increase in the size of the pool; and Based on the first condition, the second circuit system is used to increase the first amount of credit at a third rate that is less than the first rate, wherein the third rate is based on a size of the pool after the increase.
24. A system according to any one of claims 19 to 22, wherein: The first circuit system is further configured to detect a decrease in the size of the pool; and Based on the first condition, the second circuit system is used to increase the first amount of credit at a third rate greater than the first rate, wherein the third rate is based on a size of the pool after the reduction.
25. The system of any one of claims 19 to 22, wherein: The computing circuit includes one of a coprocessor, an accelerator, or a graphics processor unit.