System, method and apparatus for controlling power state of computing resources

Optimizing the power state conversion of computing resources through artificial intelligence and machine learning technology, solving the energy consumption problem of computing resources when switching between power states, achieving more efficient energy management.

CN120106155APending Publication Date: 2025-06-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411771811.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-22
Filing Date
2024-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Computing resources may consume a lot of time and energy when switching between power states, and the break-even energy may be greater than the energy saved by operating in a reduced power state, resulting in a disadvantageous power management.

Method used

Using artificial intelligence and machine learning technology, by collecting activity information of computing resources and training models, control information is generated to optimize power state transitions of computing resources, avoiding unnecessary conversions to save energy.

Benefits of technology

By intelligently controlling the power state of computing resources, unnecessary conversions are reduced, energy consumption is reduced, and the system's energy efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106155A_ABST
    Figure CN120106155A_ABST
Patent Text Reader

Abstract

An apparatus may include at least one control circuit configured to receive activity information for one or more computing resources, and generate control information based on the activity information using a model to control a power state of at least one of the one or more computing resources. The at least one control circuit may include a multiply-accumulate circuit. The at least one control circuit may include a neural processing unit. The model may include a neural network. The activity information may include first activity information, and the at least one control circuit may be further configured to collect second activity information for the one or more computing resources and transmit the second activity information. The at least one control circuit may also be configured to receive one or more parameters of the model based on transmitting the second activity information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63 / 606,593 filed on December 5, 2023 and U.S. Patent Application Serial No. 18 / 957,651 filed on November 22, 2024, which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates generally to computing resources and, more particularly, to systems, methods, and apparatus for utilizing artificial intelligence to control the power state of computing resources. Background Art

[0004] A computing system may include one or more computing resources, such as a central processing unit (CPU), a graphics processing unit (GPU), etc. The one or more computing resources may be configured to run one or more computing workloads for artificial intelligence (AI), machine learning (ML), etc., such as training, inference, etc. Depending on the type of workload, some computing resources may consume relatively large amounts of power and / or energy.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the principles of the invention and therefore it may contain information that does not constitute the prior art. Summary of the invention

[0006] An apparatus may include: at least one control circuit configured to receive activity information of one or more computing resources and generate control information for controlling a power state of at least one of the one or more computing resources using a model based on the activity information. The at least one control circuit may include a multiply-accumulate circuit. The at least one control circuit may include a neural processing unit (NPU). The model may include a neural network. The activity information may include first activity information, and the at least one control circuit may also be configured to collect second activity information of the one or more computing resources and send the second activity information. The at least one control circuit may also be configured to receive one or more parameters of the model based on sending the second activity information. The at least one control circuit may include a buffer for storing the activity information. The at least one control circuit may also be configured to generate a timestamp for the activity information. The at least one control circuit may also be configured to generate the control information based on a characteristic of at least one of the one or more computing resources. The characteristic may include a breakeven energy.

[0007] An apparatus may include: one or more computing resources configured to operate in a first power state and to operate in a second power state based on control information; and at least one control circuit configured to receive activity information of at least one of the one or more computing resources and to generate control information using a model based on the activity information. The apparatus may also include a power circuit configured to control the second power state based on the control information. The at least one control circuit may include a multiply-accumulate circuit. The at least one control circuit may include a neural processing unit.

[0008] A method may include: using at least one control circuit connected to one or more computing resources to collect first activity information of the one or more computing resources; using the first activity information and a characteristic of at least one of the one or more computing resources to train a model; using at least one control circuit to collect second activity information of the one or more computing resources; using the model and the second activity information to generate control information; and using the control information to control a power state of at least one of the one or more computing resources. The training may include determining a first value corresponding to a first portion of the first activity information based on the characteristic and a first portion of the first activity information, determining a second value corresponding to a second portion of the first activity information based on the characteristic and a second portion of the first activity information; and generating one or more parameters of the model using the first portion of the first activity information, the second portion of the first activity information, the first value, and the second value. The first value may include a label. The label may include information for converting a power state of at least one of the one or more computing resources. The first value may include a quantity. The characteristic may include an amount of energy. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings are not necessarily drawn to scale, and in some drawings, for the purpose of illustration, elements or parts thereof of similar structure or function may generally be represented by reference indicators ending with and / or containing the same number, letter, etc. throughout the drawings. The drawings are intended only to facilitate the description of the various embodiments described herein. The drawings do not describe every aspect of the teachings disclosed herein and do not limit the scope of the claims. In order to prevent the drawings from becoming obscure, not all components, connections, etc. may be shown, and not all components may have reference numerals. However, the pattern of the component configuration can be easily apparent from the drawings. The drawings, together with the specification, illustrate example embodiments of the present disclosure and are used together with the specification to explain the principles of the present disclosure.

[0010] Figure 1 A computing system according to an example embodiment of the present disclosure is shown.

[0011] Figure 2 Graph illustrating an example embodiment of power states of computing resources and transitions between power states according to an example embodiment of the present disclosure.

[0012] Figure 3 A graph illustrating an example embodiment of activity states of multiple computing resources and transitions between activity states according to an example embodiment of the present disclosure.

[0013] Figure 4 An embodiment of a computing system having a model according to an example embodiment of the present disclosure is shown.

[0014] Figure 5 An embodiment of a scheme for training a model of a computing system according to an example embodiment of the present disclosure is shown.

[0015] Figure 6 An embodiment of a computing system having a model and a management controller according to an example embodiment of the present disclosure is shown.

[0016] Figure 7 An example embodiment of a computing system with power gating according to an example embodiment of the present disclosure is shown.

[0017] Figure 8 An example embodiment of machine learning training data according to an example embodiment of the present disclosure is shown.

[0018] Fig. 9 An example embodiment of NPU monitoring according to an example embodiment of the present disclosure is shown.

[0019] Fig.10 An example embodiment of a method for training a machine learning model for power gating according to an example embodiment of the present disclosure is shown.

[0020] Fig.11 An example embodiment of a method for power gating using a trained ML model according to an example embodiment of the present disclosure is shown.

[0021] Fig.12 An example embodiment of a system that can implement power gating system instructions and / or NPU monitoring instructions according to an example embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0022] A computing system may include one or more computing resources configured to run a computing workload. Some computing resources may consume relatively large amounts of power and / or energy, particularly when running certain computing workloads, such as training and / or inference of artificial intelligence (AI), machine learning (ML), etc. In order to reduce power and / or energy consumption, the computing resources may be configured to transition to a reduced power state based on the activity level of the computing resources. For example, for some or all of a portion of a workload during which the computing resources may be relatively inactive, power to the computing resources may be turned off (which may be referred to as power gating).

[0023] However, transitioning computing resources between power states may consume time, energy, etc. In some cases, the amount of energy associated with transitioning computing resources to and / or from a reduced power state (which may be referred to as a breakeven energy) may be greater than the amount of energy that may be saved by operating the computing resources in a reduced power state. In such cases, it may be beneficial to avoid transitioning computing resources to a reduced power state.

[0024] Some computing systems may decide whether to transition one or more computing resources between power states by performing an online calculation (e.g., during real-time operation) to determine whether the amount of energy saved by transitioning the one or more computing resources between power states may exceed the break-even energy. However, depending on implementation details, performing an online break-even calculation may be difficult and / or expensive because, for example, it may be difficult to predict the activity and / or idle duration of the computing resources.

[0025] Some computing systems according to example embodiments of the present disclosure may implement one or more models using artificial intelligence, machine learning, etc. to control the power state of one or more computing resources. For example, the control circuit may collect activity information of one or more computing resources and apply it as input to a machine learning model, which may generate control information (e.g., recommendations, decisions, etc.) that may be used to transition one or more computing resources between different power states.

[0026] In some embodiments, data collected by the control circuitry may be used to train a model according to an example embodiment of the present disclosure. For example, the control circuitry may collect a data set (e.g., historical data) of activity information of one or more computing resources while the one or more computing resources are running one or more computing workloads (e.g., target workloads).

[0027] In some embodiments, the model may be trained using one or more offline operations that may perform calculations that may be too time consuming, resource intensive, etc. to be performed during online operation of one or more computing resources, such as energy break-even calculations. For example, a data set of activity information collected by the control circuit may be loaded into a data processing system (e.g., a database system), which may use computing resources of one or more CPUs, servers, data centers, etc. to process the activity information and / or other information to generate a training data set, which may include values ​​(e.g., labels, quantities, etc.) of corresponding portions of the activity information. Examples of other information that may be used to generate a training data set may include one or more characteristics of one or more computing resources, such as break-even energy, the amount of power consumed while active, the amount of power consumed while idle, etc.

[0028] Examples of labels that may be generated for the training data set may include binary labels, such as a enter or do not enter (dne) recommendation and / or a decision to enter a certain power state. Examples of quantities that may be generated for the training data set may include one or more numbers indicating a probability that energy savings exceed energy breakeven if one or more computing resources are transitioned to a different power state.

[0029] The training data set can be used to train the model, for example, using an offline process, in which parameters of the model (which may include hyperparameters), such as weights, biases, etc., can be generated, adjusted, optimized, etc. The trained model can be loaded (e.g., by loading one or more model parameters) into the control circuit, which can use the model to control one or more power states of one or more computing resources during operation. The control circuit can collect activity information and apply it as input to the trained model, which can generate one or more control outputs to control one or more power states of one or more computing resources.

[0030] The control circuitry according to an example embodiment of the present disclosure may include one or more processors, and may implement a model, for example, by performing operations such as applying weights to input data, combining intermediate results, applying an activation function to combined results, and the like. In some embodiments, the control circuitry may include one or more NPUs, and the one or more NPUs may include circuits such as multiply-accumulate (MAC) units that may be particularly suitable for implementing one or more models. Depending on the implementation details, the use of a processor such as an NPU may enable the control circuitry to implement relatively complex (and therefore potentially more accurate) prediction models. Additionally, or alternatively, depending on the implementation details, the use of a processor such as an NPU may enable the control circuitry to operate using a relatively wide range of training techniques, inference techniques, models, usage (e.g., activity) patterns of computing resources, and the like.

[0031] Some computing systems according to example embodiments of the present disclosure may include management controller circuitry that receives a recommendation from a model (e.g., to transition one or more computing resources between power states) and decides whether to at least partially implement the recommendation. For example, the management controller circuitry may receive a recommendation from a model to transition a cluster of computing resources to a reduced power state. Depending on one or more additional considerations, the management controller circuitry may send control information to the power circuitry to transition a subset of the computing resource cluster to a reduced power state.

[0032] The present disclosure covers many aspects related to the power state of converting computing resources. The various aspects disclosed herein can have independent utility and can be embodied individually, and not every embodiment can utilize every aspect. In addition, these aspects can also be embodied in various combinations, some of which can amplify some benefits of various aspects in a collaborative manner.

[0033] Figure 1 A computing system according to an example embodiment of the present disclosure is shown. Figure 1 The computing system 100 shown in FIG. 1 may include one or more computing resources 102, a control circuit 104 that may generate control information 108 based on activity information 106 about the one or more computing resources 102, and / or a power circuit 110 that may control one or more power states of the one or more computing resources 102 based on the control information 108. In some embodiments, the control circuit 104 may also use one or more characteristics 112 and / or other information of some or all of the one or more computing resources 102 to generate the control information 108.

[0034] One or more computing resources 102 may be configured to run one or more computing workloads. Examples of computing resources 102 may include processing units such as central processing units (CPUs), graphics processing units (GPUs), neural processing units (NPUs), tensor processing units (TPUs), digital signal processors (DSPs), and the like. Additional examples of computing resources 102 may include circuits such as combinatorial logic, sequential logic, gate arrays, timers, counters, registers, state machines, accelerators, and the like. Examples of computing workloads may include training, inference, and the like for artificial intelligence (AI), machine learning (ML), neural networks, deep learning, and the like, which may be collectively and / or individually referred to as AIML. Some embodiments may be implemented using one or more of a die (which may also be referred to as a chip), a dielet (which may also be referred to as a chiplet), a SoC, a system-in-package (SIP), a multi-chip module, a chip on a wafer on a substrate (CoWoS) (e.g., with or without a semiconductor interposer), and the like, or a combination thereof.

[0035] Some computing resources 102 may consume relatively large amounts of power and / or energy, particularly when running certain computing workloads (e.g., training, inference, etc. for artificial intelligence, machine learning, etc.). In order to reduce power and / or energy consumption, control circuitry 104 and / or power circuitry 110 may be configured to cause some or all of computing resources 102 to operate in one or more power states based on activity information of computing resources 102, workload, etc. For example, for a portion of the workload, computing resources 102 may be relatively inactive (e.g., idle), and thus may perform little or no useful work even though computing resources 102 may continue to consume power (e.g., standby power, leakage power, etc.). Therefore, for some or all of the portion of the workload for which computing resources 102 may be relatively inactive, computing resources 102 may be placed in a reduced power state (e.g., power to computing resources 102 may be reduced or shut off, which may be referred to as power gating).

[0036] The power circuit 110 may be implemented using any device that can control the power state of and / or the flow of power to one or more computing resources 102. Examples may include a power supply for all or a portion of a circuit, a power regulator (e.g., a voltage regulator) for all or a portion of a circuit, a switch, a bias current signal, etc. The power control circuit may be implemented using one or more components that are separate and / or integrated with the one or more computing resources 102. The power circuit 110 may be configured to control the power state of one or more computing resources 102 at the level of individual computing resources 102 (e.g., processing units, logic circuits, state machines, etc.), clusters of computing resources 102, dies having one or more computing resources 102, packages having one or more dies, etc., or combinations thereof.

[0037] Transitioning computing resource 102 between power states may consume time, energy, etc. For example, turning off power to computing resource 102 may involve storing information such as register values, cache contents, program counters, etc. to memory before turning off power to computing resource 102 (which may be referred to as saving the state of the computing resource). Storing this information may consume time, energy, etc. Additionally, or alternatively, loading such information from memory after power is turned on to computing resource 102 (which may be referred to as restoring the state of computing resource 102, e.g., so that computing resource 102 can resume a workload) may consume time, energy, etc. Additionally, or alternatively, one or more electrical processes for applying and / or removing power from computing resource 102 may consume time, energy, etc.

[0038] In some cases, the amount of energy associated with transitioning computing resource 102 to and / or out of a reduced power state (which may be referred to as transition energy, transition energy penalty, and / or breakeven energy) may be greater than the amount of energy that may be saved by operating computing resource 102 in a reduced power state. Additionally, or alternatively, the amount of time associated with transitioning computing resource 102 to and out of a reduced power state may be greater than the amount of time during which computing resource 102 may operate in a reduced power state. In such cases, and depending on implementation details, it may be beneficial to avoid transitioning computing resource 102 to a reduced power state.

[0039] Figure 2 A graph illustrating an example embodiment of power states of a computing resource and transitions between power states according to an example embodiment of the present disclosure. In graph 202, the vertical axis may indicate power P, e.g., in Watts. The horizontal axis may indicate time. In some embodiments, when a computing resource is inactive (e.g., idle), the computing resource may transition to a reduced power state (e.g., a power saving mode), e.g., to prevent loss of data that may be stored in volatile memory, loss of program counter position, etc.

[0040] In some embodiments, the energy cost may be related to at most Figure 2 204. The power gated time period (power_gated) 204 shown in FIG. 204 is associated with a transition from an active power state to a reduced power state (e.g., a power gated mode). Active time period 208 shows a time period during which the computing resource is active. Upon entering the power gated mode, power may be increased during the time period tr_in, as indicated by the transition in 204. The energy associated with entering the power gated mode (which may be referred to and / or characterized as a transition energy, energy cost, energy penalty, etc.) may be determined, for example, by multiplying the power associated with the transition by the time tr_in (e.g., the area of ​​the transition 204).

[0041] When exiting the power gating mode, power may similarly be consumed in the transition, as indicated by transition exit 206. The transition energy associated with exiting the power gating mode may be determined, for example, by the product of the power associated with the transition and the time tr_out (eg, the area of ​​transition 206).

[0042] The power savings 210 indicates how much power can be saved by converting the computing resource to the power gated mode. The power gated period (power_gated) 214 shows an example amount of time that power gating can save power (in some embodiments, this can be implemented as and / or referred to as a sleep mode 212 of the computing resource). In this example, the power consumed by the computing resource during normal mode (e.g., steady state) is shown as 1 watt, and the power consumed during the power gated mode is shown as 0.2 watts. Therefore, the power savings during the power gated mode can be 0.8 watts. These numbers are just examples, and any amount of power can be used in each period based on the size, type, utility, etc. of the circuit. The energy savings during the power gated mode can be determined, for example, by the product of the power savings 210 and the power gated period power_gated 214. In some embodiments, the break-even amount of energy used for energy savings can be equal to the sum of the energy costs of the conversion 204 and the conversion 206.

[0043] Figure 3 A graph illustrating an example embodiment of activity states of multiple computing resources and transitions between activity states according to an example embodiment of the present disclosure. Figure 3 The embodiment 302 shown in FIG. 302 may be shown in the context of computing resources implemented using four NPUs (NPU0-NPU3), but the principles may be applied to any number and / or type of computing resources. In the graph 302, the activity states (e.g., active and / or inactive states, power consumption levels, etc.) of NPU0-NPU3 are plotted according to Figure 3 The legend at the bottom is indicated by dashed and / or dotted lines.

[0044] The no-entry time period (“dne”) 304 may indicate a time period during which NPU0-NPU3 may be inactive (e.g., idle), but the total transition energy used to enter and exit the power gating mode of NPU0-NPU3 may exceed the total energy savings of NPU0-NPU3. Additionally, or alternatively, the no-entry (dne) time period 304 may indicate a time period during which an entry transition period (tr_in) and / or an exit transition period (tr_out) of one or more of NPU0-NPU3 may be greater than or equal to the time period dne 304.

[0045] The time period "Break Even" 307 indicates a minimum amount of time that one or more (e.g., all) of NPU0-NPU3 may be in a power gated mode, because the total transition energy for entering and exiting the power gated mode of NPU0-NPU3 may be equal to the total energy savings of NPU0-NPU3. The time period "Enter" 306 may indicate a time range within which all NPU0-NPU3 may be inactive (e.g., idle), and the total transition energy for entering and exiting the power gated mode of NPU0-NPU3 may be less than the total energy savings for placing NPU0-NPU3 in the power gated mode. Therefore, in some embodiments, it may be beneficial to transition some or all of NPU0-NPU3 to a power gated mode during the time period "Enter" 306.

[0046] Reference again Figure 1 , some computing systems 100 according to example embodiments of the present disclosure may decide whether to transition one or more computing resources 102 between power states based on one or more considerations such as activity information 106 of the one or more computing resources 102, one or more characteristics 112 of the one or more computing resources 102, and / or other considerations. The activity information 106 may include current activity, recent activity, historical activity, etc. The characteristics 112 may include breakeven energy, time to save the state of the computing resource, time to store the state of the computing resource, etc.

[0047] In some embodiments, computing system 100 may perform online calculations (e.g., calculations during real-time operation) to determine the energy break-even duration of one or more computing resources 102, and thus determine whether to transition some or all of one or more computing resources 102 to a different power state (e.g., into a power-gated state). However, depending on implementation details, performing online break-even calculations may be difficult and / or expensive (e.g., in terms of time, energy, etc.) because, for example, it may be difficult to predict the duration of activity modes (e.g., active, idle, etc.) of some or all of one or more computing resources 102. As another example, in some embodiments, computing system 100 may use a heuristic algorithm (e.g., a relatively simple heuristic) to make power transition decisions. However, depending on implementation details, such an algorithm may not produce acceptable results.

[0048] Some computing systems 100 according to example embodiments of the present disclosure may implement one or more models that use artificial intelligence, machine learning, or the like to control the power state of some or all of one or more computing resources 102 .

[0049] Figure 4 An embodiment of a computing system having a model according to an example embodiment of the present disclosure is shown. Figure 4The computing system 400 shown in FIG. 4 may include a computer system similar to Figure 1 One or more elements of those shown in , wherein elements having similar construction, operation, etc. may be indicated by reference numerals ending with and / or containing the same numbers, letters, etc.

[0050] exist Figure 4 In the illustrated computing system 400 (which in some embodiments may be referred to and / or characterized as an accelerator architecture), the control circuitry 404 may include a model 414 that may generate control information 408 based on activity information 406 about one or more computing resources 402. The power circuitry 410 may control one or more power states of the one or more computing resources 402 based on the control information 408. In some embodiments, the model 414 may also use one or more characteristics 412 and / or other information of some or all of the one or more computing resources 402 to generate the control information 408. In some embodiments, the control circuitry 404 may also include monitoring circuitry 416 that may collect and / or store the activity information 406 of the one or more computing resources 402. The model 414 may use the activity information 406 of some or all of the one or more computing resources 402 as input to generate the control information 408 (e.g., one or more recommendations, decisions, etc.), which may be used by the power circuitry 410 to transition some or all of the one or more computing resources 402 between different power states.

[0051] Models according to example embodiments of the present disclosure may be trained, for example, using data collected by monitoring circuitry 416. For example, monitoring circuitry 416 may collect a data set (e.g., historical data) of activity information 406 of some or all of one or more computing resources 402 while one or more computing resources 402 are running one or more example (e.g., target) computing workloads. Activity information 406 may include, for example, timestamp data indicating when various computing resources 402 were active and / or inactive (e.g., idle), activity levels (e.g., percentage of processing capacity) when computing resources 402 were active, activity types (e.g., computing, transmitting data, etc.) when computing resources 402 were active, how much power various computing resources 402 consumed for different operations, and the like.

[0052] Figure 5 An embodiment of a scheme for training a model of a computing system according to an example embodiment of the present disclosure is shown. Figure 5 The scheme 500 shown in FIG. 5 may include a method similar to Figure 1 and / or Figure 4One or more elements of those shown in , wherein elements having similar construction, operation, etc. may be indicated by reference numerals ending with and / or containing the same numbers, letters, etc.

[0053] Figure 5 The illustrated scheme 500 may include one or more data processing systems 518 that may perform one or more operations associated with the training mode 514. The model 514 may be trained, for example, using one or more offline operations that may perform calculations, such as energy break-even calculations, that may be too time consuming, resource intensive, etc. to perform during online operation of the one or more computing resources 502 and / or monitoring circuitry 516. For example, a data set of activity information 506 collected by the monitoring circuitry 516 may be loaded into one or more data processing systems 518 (e.g., a database system), which may perform training data operations 520 using computing resources of one or more CPUs, servers, data centers, etc., to process the activity information 506 and / or other information to generate a training data set 522 that may include values ​​(e.g., labels, quantities, etc.) for corresponding portions of the activity information 506. Examples of other information that may be used to generate training data set 522 may include one or more characteristics 512 of some or all of one or more computing resources 502, such as breakeven energy, amount of power consumed while active, amount of power consumed while inactive (e.g., idle), and the like.

[0054] Examples of labels that may be generated for the training data set 522 may include a (e.g., binary) label for a number such as a enter or no-entry (DNE) decision, recommendation, etc. for entering a certain power state, a digital label (e.g., using more than one binary bit) such as a conditional version of an enter and / or DNE decision, recommendation, etc., a digital label that may indicate a partitioning decision, recommendation, etc., where some computing resources 502 may transition from a first power state to a second power state and some other computing resources 502 may remain in the first power state and / or transition to a third power state, etc. Examples of quantities that may be generated for the training data set 522 may include one or more numbers indicating a probability that energy savings exceed energy breakeven if some or all of one or more computing resources 502 transition to a different power state, a number of computing resources 502 transitioning to a different power state, etc.

[0055] The training data set 522 may be used to train the model 514, for example, using a training process 524 (e.g., an offline process), in which one or more parameters 526 (which may include hyperparameters), such as weights, biases, etc., of the model 514 may be generated, adjusted, optimized, etc. For example, the trained model 514 as described herein may be loaded into the control circuitry 504, which may use the model 514 to control one or more power states of one or more computing resources 502 during operation. The trained model 514 may be loaded, for example, by loading one or more parameters 526, such as weights, biases, etc., into the control circuitry 504, which may include monitoring circuitry 516 to collect activity information 506 (e.g., real-time or online activity information) of some or all of the one or more computing resources 502. The monitoring circuitry 516 may apply the collected activity information as input to the model 514 during online operation, and the model 514 may generate control information 508 to control one or more power states of some or all of the one or more computing resources 502.

[0056] The control circuit 504 according to an example embodiment of the present disclosure may include one or more processors, which may implement a model 514 (e.g., a neural network), for example, by performing operations such as applying weights to input data (e.g., multiplication), combining intermediate results (e.g., addition), applying an activation function to the combined results, and the like.

[0057] Although the control circuitry 504 is not limited to any particular type or amount of circuitry for implementing the model 514, in some embodiments, the control circuitry 504 may include one or more NPUs, which may include circuitry that may be particularly suitable for implementing one or more models. For example, in some embodiments, the NPU may include one or more multiply-accumulate (MAC) units that may efficiently perform multiplication and / or addition at relatively high speeds, relatively low power, etc. Depending on the implementation details, the use of a processor such as an NPU may enable the control circuitry 504 to implement relatively complex (and therefore potentially more accurate) prediction models 514. Additionally, or alternatively, depending on the implementation details, the use of a processor such as an NPU may enable the control circuitry 504 to operate with a relatively wide range of training techniques, inference techniques, models 514 (e.g., types, sizes, etc.), usage patterns (e.g., activity patterns) of the computing resources 502, etc., as compared to, for example, a general-purpose CPU.

[0058] In some embodiments, one or more NPUs used to implement any of the models disclosed herein may have any number of the following characteristics and / or implement any number of the following features, components, operations, etc.

[0059] In some embodiments, the NPU may include one or more of the following components: a MAC unit (e.g., a MAC engine), an activation unit, a weight decoding circuit, local and / or shared memory, an element-by-element engine, a memory controller (e.g., a direct memory access (DMA) controller), etc.

[0060] In some embodiments, the MAC unit may perform calculations for multiplication (e.g., matrix multiplication) and / or addition, convolution, dot product, and / or other functions. For example, in some embodiments, the MAC unit may multiply the activity information 608 by the corresponding weight and sum the results of the multiplication operations to generate an intermediate result.

[0061] In some embodiments, the activation unit may scale intermediate results from the MAC unit, apply an activation function to the intermediate results, scale element-wise operations, perform resizing operations, and the like.

[0062] In some embodiments, the element-by-element engine may perform element-by-element arithmetic operations.

[0063] In some embodiments, one or more of the components described herein may perform operations on integers, floating point numbers, combinations thereof, and the like.

[0064] In some embodiments, one or more of the components described herein may perform one or more operations on an element-by-element basis, on a layer-by-layer basis (e.g., across layers of a neural network), on a depth-wise basis, and the like.

[0065] In some embodiments, the NPU may perform operations with relatively low-precision arithmetic (eg, eight bits or less), eg, to reduce computational complexity, improve energy efficiency, etc.

[0066] In some embodiments, the weight decoding circuit can preload (e.g., prefetch) and / or decompress weights that can be compressed, for example to reduce the amount of memory that can store the weights. Depending on the implementation details, this can enable larger models than can otherwise be handled by specific control circuits, memory, etc.

[0067] In some embodiments, the NPU may be adapted to perform AIML tasks and workloads such as computing neural network layers using scalar, vector, and / or tensor math followed by one or more activation functions (e.g., non-linear activation functions).

[0068] In some embodiments, the NPU can perform relatively low-latency parallel computations (e.g., executing multiple concurrent neural network operations).

[0069] In some embodiments, the NPU can utilize relatively high bandwidth memory (e.g., on-die memory) and / or acceleration hardware (e.g., systolic array architecture and / or tensor processing unit). In some embodiments, the NPU can pre-fetch weights, activations, etc.

[0070] In some embodiments, the NPU may implement one or more features, such as a long short-term memory (LSTM) network that may implement a recurrent neural network (e.g., for problems involving learning sequential dependencies in sequence prediction), a gated recurrent unit (GRU) for the vanishing gradient problem, and the like.

[0071] Figure 6 An embodiment of a computing system having a model and a management controller according to an example embodiment of the present disclosure is shown. Figure 6 The computing system 600 shown in FIG. 6 may include a computer system similar to Figure 1 , Figure 4 and / or Figure 5 One or more elements of those shown in , wherein elements having similar construction, operation, etc. may be indicated by reference numerals ending with and / or containing the same numbers, letters, etc.

[0072] exist Figure 6 In the computing system 600 shown in FIG. 6 , the control circuitry 604 may include management circuitry 628 that may control one or more aspects of the operation of one or more computing resources 602. For example, in some embodiments, the management circuitry 628 may receive control information in the form of recommendations 608A from the model 614. The management circuitry 628 may use the recommendations 608A and / or other information available to the management circuitry 628 to decide whether to fully or partially implement the recommendations 608A and send control information 608B to the power circuitry 610 to implement the decision.

[0073] For example, if the model 614 sends a recommendation 608A (e.g., a binary recommendation) indicating that it may be beneficial to transition all of the one or more computing resources 602 to a different power state (e.g., a reduced power state), the management circuit 628 may check whether there is sufficient available space in the memory 632 to save the one or more states of the one or more computing resources 602. If there is sufficient space, the management circuit 628 may implement the recommendation by sending control information 608B to the power circuit 610 to cause the power circuit to, for example, transition the one or more computing resources 602 to a different power state for a specified period of time. Additionally, or alternatively, the management circuit 628 may send one or more save indications 630 (e.g., digital signals) to cause the one or more computing resources 602 to save the one or more states to the memory 632. However, if there is insufficient space in the memory 632, the management circuit 628 may avoid sending the control information 608B to the power circuit, thereby maintaining the one or more computing resources 602 in their current power state.

[0074] As another example, if model 614 sends recommendation 608A in the form of a number indicating a probability that energy savings exceed energy breakeven if all of one or more computing resources 602 are transitioned to a different power state, management circuitry 628 may compare the probability to a threshold to decide whether to transition one or more computing resources 602 to a different state. For example, if the probability is relatively low, management circuitry 628 may refrain from transitioning one or more computing resources 602 to a reduced power state, e.g., because the relatively low probability of saving energy may be outweighed by one or more other considerations, such as a quality of service (QoS) arrangement that may provide an incentive to keep one or more computing resources 602 running at full operating speed.

[0075] In some embodiments, model 614 may make decisions, recommendations, etc. and / or management circuitry 628 may make decisions, implement recommendations, etc. at the level of an individual computing resource, a cluster of computing resources, a die, multiple dies within a package, etc.

[0076] For purposes of illustration, some specific implementation details may be described. Figure 7-11 Some example embodiments shown in , such as NPUs arranged in a cluster, power circuits implemented with regulators, specific signals, etc. However, the principles of the present disclosure are not limited to these or any other implementation details.

[0077] Figure 7 An example embodiment of a computing system with power gating according to an example embodiment of the present disclosure is shown. Figure 7The computing system 702 shown may include a power supply, Vdd 704, a low dropout regulator (LDO) 706, an NPU cluster 708, a system management controller 710, and / or an NPU monitor 712. One or more of the NPU cluster 708, the system management controller 710, and the NPU monitor 712 may be on a separate power domain. For example, the NPU cluster 708 may be on a cluster power domain, the system management controller 710 may be on a system management power domain, and the NPU monitor 712 may be on a power domain (e.g., a normally open power domain) that enables the NPU monitor 712 to monitor and / or control the NPU cluster 708. One or more (e.g., each) power domain may be connected to a separate or identical power supply, such as Vdd 704. In the NPU cluster 708, each box of the array may indicate one NPU of the NPU cluster. The LDO 706 may operate like a power switch, allowing power to be disconnected from the entire NPU cluster 708. In other embodiments, the LDO 706 may be configured to selectively disconnect power from different NPUs. The LDO 706 may be implemented using any type of voltage regulation block, power gating block, power gating technique, etc. Examples of the LDO 706 may include a linear voltage regulator (such as a fixed voltage regulator or an adjustable voltage regulator), a switching voltage regulator (such as a buck converter, a boost converter, etc.), a charge pump power management IC (PMIC), etc.

[0078] The array of NPUs in 708 may include any number of NPUs, depending on, for example, system requirements or architecture. One or more (e.g., all) NPUs in a given cluster may be idle, for example, to enter a cluster-level power gating mode. In some embodiments, little or no work or processing may occur on each NPU in the cluster to shut down the power of the cluster, for example, to avoid interruption of work. In addition, to prevent interruption of work, the NPU monitor 712 may monitor the activity of the NPU cluster 708.

[0079] The NPU monitor 712 may monitor the activity of the NPU cluster 708 by receiving busy / idle information (e.g., data) 720 and training one or more ML models to operate in the NPU monitor 712. After the ML model is trained, for example, according to one or more methods described herein, the ML model may monitor the activity of the NPU cluster 708 based on the busy / idle information 720. If the ML model determines that one or more (e.g., all) NPUs in the cluster 708 may become idle for a period of time equal to or exceeding the break-even time, the NPU monitor 712 may notify the system management controller 710 to stop the NPU cluster 708. An interrupt 724 may be sent from the NPU monitor 712 to send a power-off control signal 714 and / or a regulation control signal 716 to shut down power to the NPU cluster 708.

[0080] In addition to shutting down the LDO 706 or power-gating switches, the system management controller 710 may also handle save / restore 718 data movement. The system management controller 710 may, for example, save the state of the NPU cluster 708 for when power is returned. To implement power gating, the NPU monitor 712 may execute an efficient method of ML and decide when it is appropriate to turn the NPU cluster on and off.

[0081] Some embodiments may improve efficiency by building an offline model using the power breakeven duration and / or building an ML model to determine when it is appropriate to enter or prohibit power gating. The ML model may be built offline using information collected by the NPU monitor 712. For example, in some embodiments, the ML algorithm may be trained while offline, and the input to the ML model may be the aggregated NPU busy / idle data 720. In some embodiments, the power breakeven formula may be used to create labels, and the labels are used to train the ML algorithm.

[0082] Figure 8 802 shows an example embodiment of machine learning training data according to an example embodiment of the present disclosure. Each row in the ML training data 802 may represent a time increment. In some embodiments, the NPU history buffer (such as Fig. 9 The data in the active buffer 912 shown in FIG. 8 can be transformed and labeled so that it can adapt to a variety of ML and neural network models. The example ML training data 802 can transpose the NPU data so that each row can contain a number of NPUs (e.g., Figure 7) and provide labels (e.g., decisions) for training. This is just one example of data transformation and labeling to accommodate one or more ML models. In some embodiments, 0 may indicate that the NPU is active and 1 may indicate that the NPU is inactive or idle. In some other embodiments, 1 may indicate that the NPU is active and 0 may indicate that the NPU is inactive or idle. Each row may include a label indicating a decision or recommendation to enter or prohibit entering a power gating mode based on, for example, a break-even calculation. In some embodiments, the ML training data 802 may be referred to as and / or characterized as vectorized data (e.g., activity information from the NPU cluster 708 may be vectorized to create the ML training data 802).

[0083] During the training phase (e.g., offline mode), activity and / or idle information may be streamed into an activity buffer such as Fig. 9 The ML model may reside in an NPU (e.g., such as an NPU monitor 712) such as Fig. 9 In some embodiments, Fig. 9 The NPU 906 shown in FIG. 1 can be used with Fig. 9 , wherein the I / O circuitry 904 of the system management controller shown in FIG. 7 is connected to an interface, where the I / O circuitry can send signals (e.g., interrupts) to the system management controller 710. In some embodiments, the signal to the system management controller can include a power gate entry or power gate prohibit entry prediction from the ML model.

[0084] In an example embodiment, the NPU cluster 708 can run for a period of time while busy and idle data is collected and collected in the NPU monitor 712. During offline operation, the NPU monitor 712 can use its own NPU 906 to train the ML model. Additionally, or alternatively, the multivariate training data can be recorded using, for example, an activity buffer 912. Using the break-even formula, a prediction of one or more idle durations can be calculated. Labels can be created during the offline process to indicate when to enter and when to prohibit entering a power gating mode. Examples of ML models can include random forests, deep neural networks (DNNs), convolutional neural networks (CNNs), logistic regression, and / or any classification algorithm. These are intended to be examples and are not intended to be limiting in any way.

[0085] Fig. 9An example embodiment of an NPU monitor according to an example embodiment of the present disclosure is shown. The NPU monitor 902 may exist on an independent power domain that may be sufficiently turned on (e.g., always on) to ensure the monitoring capability of the system. A separate voltage regulator may be provided for the NPU monitor 902. An activity buffer 912 may be used to store activity and / or idle information, and signals received from the NPU cluster 708. The activity buffer 912 may be one or more storage units (e.g., static random access memory (SRAM) and / or dynamic random access memory (DRAM)), but any type of memory may be used for the activity buffer 912.

[0086] The NPU monitor 902 may include a time stamp counter (TSC) 910, which may be implemented, for example, with a clock or crystal oscillator to record event times. The NPU monitor 902 may execute an ML model (e.g., an ML algorithm) to determine and / or predict beneficial times for power gating. An I / O 904 to a system management controller may enable the NPU monitor 902 to communicate with the system management controller 710. The NPU 906 may include, for example, one or more multiply-accumulate (MAC) units.

[0087] Activity buffer 912 may have one or more sampling rates. The sampling rate may be variable. Activity buffer 912 may store activity information shown, for example, in ML training data 802. In some embodiments, each row in ML training data 802 may indicate a different time slice. Timestamp counter 910 may assign timestamps to, for example, each row of ML training data 802, one or more time steps within a row (e.g., Figure 8 In some embodiments, the size of the NPU 906 can be appropriately adjusted for the amount of idle data generated. In some embodiments, the NPU 906 can be customized to perform vector math.

[0088] During the training phase of the ML algorithm, the active and / or idle information may be streamed into the active buffer 912. In some embodiments, the ML model may reside or be programmed into the NPU 906. The NPU 906 may interface with the I / O 904 to the system management controller, which in turn may send signals to the system management controller 710. In some embodiments, the data pre / post processor 908 may process data from and / or to the NPU 906. For example, the data may be quantized, compressed, decompressed, etc. to match the format used by the NPU 906, the system management controller 710, etc. In some embodiments, the data may be refined. Some example ML algorithms that may be run in the NPU 906 include random forests, deep neural networks, convolutional neural networks, logistic regression, etc. Any ML or AI algorithm, including classification algorithms, may be executed and considered.

[0089] Fig.10 An example embodiment of a method for training a machine learning model for power gating according to an example embodiment of the present disclosure is shown. Although the example method 1000 depicts a specific sequence of operations, the sequence may be changed without departing from the scope of the present disclosure. For example, some of the depicted operations may be performed in parallel or in a different order that does not substantially affect the functionality of the method 1000. In other examples, different components of an example device or system implementing the method 1000 may perform functions substantially simultaneously or in a specific order.

[0090] In some embodiments, the example method 1000 may be performed offline. In operation 1002, the NPU monitor 902 may collect idle and active histories across a target workload set from an NPU monitoring history buffer. Collecting idle and active histories may include, for example, receiving busy / idle data 720 from the NPU cluster 708.

[0091] In operation 1004, the NPU monitor 902 may post-process the data. The NPU monitor 902 may utilize a power break-even formula to determine "entry" and "no-entry" data labels. In some embodiments, the data pre / post processor 908 may perform post-processing.

[0092] In some embodiments, the power breakeven formula may be specified as:

[0093] Energy saved ≥ conversion entry power + conversion exit power

[0094] Herein, the conversion entry power + the conversion exit power may be referred to and / or characterized as a conversion energy penalty.

[0095] In operation 1006, the NPU monitor 902 may use the labeled data to train one or more ML models. For example, any ML or AI training model may be used. Some examples may include DNN, CNN, classification methods, linear regression, and / or random forest techniques, but any model may be used.

[0096] In operation 1008, the NPU monitor 902 may evaluate the ML model(s). The NPU monitor 902 may tune or improve the data until a satisfactory model and prediction accuracy is achieved.

[0097] Fig.11 An example embodiment of a method for power gating with a trained ML model according to an example embodiment of the present disclosure is shown. Although the example method 1100 depicts a specific sequence of operations, the sequence may be changed without departing from the scope of the present disclosure. For example, some of the depicted operations may be performed in parallel or in a different order that does not substantially affect the functionality of the method 1100. In other embodiments, different components of an example device or system implementing the method 1100 may perform functions substantially simultaneously or in a specific order. In some embodiments, the example method 1100 may be performed online while the system is active.

[0098] In some embodiments, at operation 1102, the NPU monitor 902 may collect idle and / or active data in the activity buffer 912 using the timestamps provided from the timestamp counter 910. In some embodiments, at operation 1104, the NPU monitor 902 may process the idle and / or active data to adapt one or more ML models loaded into the NPU 906 hardware.

[0099] In some embodiments, at operation 1106, the NPU monitor 902 may predict one or more enter / no-entry power gating decisions using the ML model executed on the NPU 906. In some embodiments, at operation 1108, the NPU monitor 902 may send the prediction to the system management controller 710 via the I / O 904 hardware to the system management controller.

[0100] In some embodiments, at operation 1110, the system management controller 710 may make a decision to enter / disable cluster-level power gating using the prediction from the NPU monitor 902. For example, when the prediction indicates that the energy saved may be higher than the energy consumed by the conversion entry power + conversion exit power, the system management controller 710 may send a signal to shut down the NPU cluster 708.

[0101] In some embodiments, at operation 1112, the system management controller 710 may coordinate saving the state of the NPU cluster 708 to a temporary storage device (such as DRAM) when the decision to power gate is made, and shutting down the power gate or LDO 706 that supplies power to the cluster of NPU clusters 708. In some embodiments, at operation 1114, the system management controller 710 may maintain power to the NPU cluster 708 based on the decision not to enter a power gated mode.

[0102] Fig.12 An example embodiment of a system that may implement power-gating system instructions and / or NPU monitor instructions according to an example embodiment of the present disclosure is shown. Fig.12 The system shown in may include a processing circuit 1202 coupled to a memory 1204, an integrated circuit (IC) and / or a system on a chip (SoC) 1210, a user interface 1206 (or GUI), and / or a network interface 1212. In some embodiments, Fig.12 The components of the system shown in the figure can be communicatively connected via a system bus 1208. The system bus 1208 can be any type of data or system interconnect or structure. The system bus 1208 can include interfaces and / or protocols such as Peripheral Component Interconnect Express (PCIe), Non-Volatile Memory Express (NVMe), NVMe-over-Fabric (NVMe-oF), Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Remote Direct Memory Access (RDMA), RDMA over Converged Ethernet (ROCE), Fibre Channel, InfiniBand, Serial ATA (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), iWARP, Hypertext Transfer Protocol (HTTP), Compute Express Link (CXL), etc., or any combination thereof.

[0103] The processing circuit 1202 may be implemented using one or more hardware logic components and / or circuits. For example, but not limited to, illustrative types of hardware logic components that can be used include NPUs, NPU clusters, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), general purpose microprocessors, microcontrollers, digital signal processors (DSPs), graphics processing units (GPUs), etc., or any other hardware logic components capable of performing calculations or other operations on information.

[0104] The memory 1204 may be volatile (eg, RAM, etc.), non-volatile (eg, ROM, flash memory, etc.), or a combination thereof. In some configurations, computer readable instructions for implementing one or more embodiments disclosed herein may be stored in the IC / SoC 1210 .

[0105] In another embodiment, the memory 1204 may be configured to store software. In some embodiments, software may refer to any type of instructions, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. The instructions may include code (e.g., in source code format, binary code format, executable code format, or any other suitable code format). The instructions, when executed by the processing circuit 1202, may cause the processing circuit 1202 to perform the various processes described herein. Specifically, the instructions, when executed, cause the processing circuit 1202 to execute an ML algorithm, track and collect data about processor / NPU idle and active states, train an ML model, and / or perform power gating based on a learned model.

[0106] IC / SoC 1210 may be one or more integrated circuits (ICs) or SoCs that include the components described herein as well as storage memory (e.g., flash memory or other memory technology), or any other medium that can be used to store desired information. In some embodiments, IC / SoC 1210 may include one or more power gated devices with one or more NPU / GPU clusters, power supplies, voltage regulators, one or more system management controllers, and NPU monitoring. In some embodiments, IC / SoC 1210 may include, for example, all or part of power gated device 702.

[0107] IC / SoC 1210 may store and maintain power gating system instructions 1214 that may be executed according to method 1000 and / or method 1100, and monitoring NPU service instructions 1216 that may be executed according to appropriate ML algorithms discussed herein. Network interface 1212 may enable Fig.12 The system shown in is capable of communicating with the Internet, an intranet, a cloud server network, etc. for the purpose of receiving data, sending data, etc.

[0108] Some embodiments include a system of one or more computers that can be configured to perform specific operations or actions by installing software, firmware, hardware, or a combination thereof on the system, wherein the software, firmware, hardware, or a combination thereof causes the system to perform actions in operation. One or more computer programs can be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform actions. A general aspect includes a device having a first NPU cluster. The device can also be an NPU controller circuit configured to turn the power of the first NPU cluster on and off. The device can also include an NPU monitoring circuit configured to send instructions to the NPU controller circuit to turn the power of the first NPU cluster on and off.

[0109] In some embodiments, the NPU monitoring circuit may include an NPU configured to process activity and idle signals of a first NPU cluster received from an NPU controller circuit. The NPU monitoring circuit may also include a timestamp counter that timestamps entries input to the NPU. The NPU monitoring circuit may also include an activity buffer configured to store activity and idle signals of the first NPU cluster. The NPU may also be configured to train a machine learning model using the activity signal and the idle signal. The machine learning model may be used to power gate the first NPU cluster. The NPU monitoring circuit may also include an input / output interface to the NPU controller circuit configured to receive activity and idle signals from the NPU monitoring circuit. The NPU monitoring circuit may also include an input / output interface to the NPU controller circuit configured to transmit a power gating decision to the NPU monitoring circuit. The NPU monitoring circuit may include a data preprocessor and / or a postprocessor configured to process data entering and leaving the NPU. The NPU may also include a second NPU cluster.

[0110] In some embodiments, the device may include an NPU cluster. The device may include an NPU controller circuit configured to turn on and off the power of the NPU cluster. In some embodiments, the device may include an NPU monitoring circuit configured to collect idle and active data in an activity buffer with a timestamp in the NPU monitoring circuit. In some embodiments, the device may send instructions to the NPU controller circuit to turn on and off the power of the NPU cluster based on the idle and active data.

[0111] In some embodiments, the NPU monitoring circuit may also be configured to process idle and active data to adapt to a machine learning model in the NPU loaded into the NPU monitoring circuit, and predict entry or prohibition of power gating performed by the ML model executed on the NPU. The NPU monitoring circuit may also be configured to transmit the prediction to the NPU controller circuit via the NPU monitoring circuit, and utilize the prediction from the NPU monitoring circuit by the NPU controller circuit to enter power gating of the NPU cluster. The NPU monitoring circuit may also be configured to save the state of the NPU cluster to a memory. The NPU monitoring circuit may also be configured to turn off the power gating or LDO that supplies power to the cluster.

[0112] According to some embodiments, the method may include receiving data from a buffer on an NPU monitoring circuit that monitors the NPU cluster. The method may include training the data based on the reception by an NPU in the NPU monitoring circuit. The method may also include sending a signal based on the training by the NPU monitoring circuit to shut down the power of the NPU cluster.

[0113] The method may also include processing the data by the NPU monitoring circuit, wherein the training is based on the processing. According to some embodiments, receiving may include collecting idle and active histories across the target workload set from a buffer on the NPU monitoring circuit. The processing may also include assigning tags to the data. Training may be performed based on the assigned tags.

[0114] Some embodiments disclosed above have been described in the context of various implementation details, but the principles of the present disclosure are not limited to these or any other specific details. For example, some functions have been described as being implemented by certain components, but in other embodiments, functions may be distributed between different systems and components in different locations and have various user interfaces. Certain embodiments have been described as having specific processes, operations, etc., but these terms also cover embodiments in which specific processes, operations, etc. can be implemented using multiple processes, operations, etc., or embodiments in which multiple processes, operations, etc. can be integrated into a single process, step, etc. References to components or elements may indicate only a portion of a component or element. For example, references to blocks may indicate an entire block or one or more sub-blocks. Terms such as "first" and "second" used in the present disclosure and claims may be used only for the purpose of distinguishing the elements they modify, and may not indicate any spatial or temporal order unless otherwise apparent from the context. In some embodiments, references to elements may indicate at least a portion of an element, for example, "based on" may indicate "based at least in part on", etc. References to a first element may not imply the presence of a second element. The principles disclosed herein have independent utility and may be embodied separately, and not every embodiment may utilize every principle. However, these principles can also be embodied in various combinations, some of which can amplify the benefits of the individual principles in a synergistic manner. The various details and embodiments described above can be combined to produce additional embodiments according to the inventive principles of this patent disclosure.

[0115] In some embodiments, a portion of an element may indicate less than or all of the element. A first portion of an element and a second portion of an element may indicate the same portion of an element. A first portion of an element and a second portion of an element may overlap (e.g., a portion of the first portion may be the same as a portion of the second portion).

[0116] Since the inventive principles of this patent disclosure may be modified in arrangement and detail without departing from the inventive concept, such changes and modifications are considered to be within the scope of the following claims.

Claims

1. A device comprising: At least one control circuit configured to: receiving activity information of one or more computing resources; and Control information is generated using the model based on the activity information to control a power state of at least one of the one or more computing resources.

2. The device according to claim 1, wherein: At least one control circuit includes a multiply-accumulate circuit.

3. The device according to claim 1, wherein: At least one control circuit includes a neural processing unit.

4. The device according to claim 1, wherein: The model includes a neural network.

5. The device according to claim 1, wherein: The activity information includes first activity information, and the at least one control circuit is further configured to: collecting second activity information of the one or more computing resources; and Send the second activity information.

6. The device according to claim 5, wherein: The at least one control circuit is further configured to receive one or more parameters of the model based on sending the second activity information.

7. The device according to claim 1, wherein: At least one control circuit includes a buffer for storing activity information.

8. The device according to claim 1, wherein: The at least one control circuit is further configured to generate a time stamp for the activity information.

9. The device according to claim 1, wherein: The at least one control circuit is further configured to generate control information based on a characteristic of at least one of the one or more computing resources.

10. The device according to claim 9, wherein: Characteristics include breakeven energy.

11. An apparatus comprising: One or more computing resources, configured as: operating in a first power state; and operating in a second power state based on the control information; as well as At least one control circuit configured to: receiving activity information of at least one of the one or more computing resources; and The control information is generated using the model based on the activity information.

12. The apparatus according to claim 11, further comprising: The power circuit is configured to control the second power state based on the control information.

13. The device according to claim 11, wherein: At least one control circuit includes a multiply-accumulate circuit.

14. The device according to claim 11, wherein: At least one control circuit includes a neural processing unit.

15. A method comprising: collecting first activity information of the one or more computing resources using at least one control circuit coupled to the one or more computing resources; training a model using the first activity information and a characteristic of at least one of the one or more computing resources; collecting second activity information of the one or more computing resources using at least one control circuit; generating control information using the model and the second activity information; as well as The control information is used to control a power state of at least one of the one or more computing resources.

16. The method according to claim 15, wherein: Training includes: determining a first value corresponding to the first portion of the first activity information based on the characteristic and the first portion of the first activity information; determining a second value corresponding to the second portion of the first activity information based on the characteristic and the second portion of the first activity information; and One or more parameters of the model are generated using the first portion of the first activity information, the second portion of the first activity information, the first value, and the second value.

17. The method according to claim 16, wherein: The first value includes a label.

18. The method according to claim 17, wherein: The tag includes information to transition a power state of at least one of the one or more computing resources.

19. The method according to claim 16, wherein: The first value includes an amount.

20. The method according to claim 15, wherein: Properties include the amount of energy.