Dynamic frequency scaling (DFS) thermal management with stochastic gradient optimization (SGD)
Stochastic gradient optimization (SGD) is used to determine optimal clock frequencies, addressing the challenge of maintaining performance within thermal constraints by iteratively adjusting frequencies, thus optimizing electronic system efficiency and temperature management.
Patent Information
- Application Number
- PCT/CN2024/097690
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2025-12-11
AI Technical Summary
Electronic systems face challenges in maintaining optimal performance while adhering to thermal constraints, as dynamic frequency scaling techniques often degrade performance to reduce temperature without considering efficiency.
Employing stochastic gradient optimization (SGD) to determine optimal clock frequencies that balance performance and thermal constraints by iteratively adjusting clock frequencies using gradient increment parameters and learning rates.
Achieves optimal performance within thermal limits by dynamically adjusting clock frequencies, enhancing efficiency and reducing thermal impact on electronic components.
Smart Images

Figure CN2024097690_11122025_PF_FP_ABST
Abstract
Description
DYNAMIC FREQUENCY SCALING (DFS) THERMAL MANAGEMENT WITH STOCHASTIC GRADIENT OPTIMIZATION (SGD)TECHNICAL FIELD
[0001] This disclosure relates generally to the field of power management, and, in particular, to determination of clock frequencies under thermal constraint.BACKGROUND
[0002] An electronic system may require active control to maintain a proper thermal environment. for nominal operation. Some of the electrical energy supplied to the electronic system is unavoidably converted into thermal energy. The thermal energy modifies temperature distribution in the electronic system and its ambient environment. Therefore, an effective thermal management active control technique is needed for improving performance.SUMMARY
[0003] The following presents a simplified summary of one or more aspects of the present disclosure, in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated features of the disclosure, and is intended neither to identify key or critical elements of all aspects of the disclosure nor to delineate the scope of any or all aspects of the disclosure. Its sole purpose is to present some concepts of one or more aspects of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[0004] In one aspect, the disclosure provides determination of clock frequencies under thermal constraint. Accordingly, an apparatus including: a controller configured to select a first clock frequency and a second clock frequency using a stochastic gradient optimization (SGD) ; a processor coupled to the controller, the processor configured to receive the first clock frequency; and a memory coupled to the controller, the memory configured to receive the second clock frequency.
[0005] In one example, the controller is further configured to initialize the first clock frequency and the second clock frequency with a randomly selected initial first clock frequency and a randomly selected initial second clock frequency. In one example, the controller is further configured to compute a gradient increment parameter based on a first constituent power function evaluated at an updated first clock frequency and on a second constituent power function evaluated at an updated second clock frequency. In one example, the controller is further configured to determine an updated performance metric using the gradient increment parameter at the updated first clock frequency and at the updated second clock frequency and with a learning rate parameter.
[0006] Another aspect of the disclosure provides an apparatus for determining one or more clock frequencies under a thermal constraint, the apparatus including: means for initializing a first clock frequency and a second clock frequency with a randomly selected initial first clock frequency and a randomly selected initial second clock frequency; means for computing a gradient increment parameter based on a first constituent power function evaluated at an updated first clock frequency and on a second constituent power function evaluated at an updated second clock frequency; and means for determining an updated performance metric using the gradient increment parameter at the updated first clock frequency and at the updated second clock frequency and with a learning rate parameter.
[0007] In one example, the apparatus further includes: means for comparing an updated objective function based on the updated performance metric and an allowable dc power characteristic to a performance threshold; means for determining the first constituent power function of the first clock frequency as a first nth order polynomial function; means for determining the second constituent power function of the second clock frequency as a second nth order polynomial function; and means for determining the allowable dc power characteristic as a function of a temperature.
[0008] Another aspect of the disclosure provides a method including: initializing a first clock frequency and a second clock frequency with a randomly selected initial first clock frequency and a randomly selected initial second clock frequency; computing a gradient increment parameter based on a first constituent power function evaluated at an updated first clock frequency and on a second constituent power function evaluated at an updated second clock frequency; and determining an updated performance metric using the gradient increment parameter at the updated first clock frequency and at the updated second clock frequency and with a learning rate parameter.
[0009] In one example, the gradient increment parameter numerically represents a gradient of the first constituent power function and the second constituent power function. In one example, the gradient increment parameter is represented as a two-dimensional vector variable. In one example, the gradient increment parameter is computed as a first-order derivative approximation.
[0010] In one example, the method further includes selecting the learning rate parameter to balance an optimization convergence and an optimization speed. In one example, the method further includes comparing an updated objective function based on the updated performance metric and an allowable dc power characteristic to a performance threshold.
[0011] In one example, the updated objective function is an Euclidean norm squared of the updated first clock frequency and the updated second clock frequency. In one example, the method further includes determining the first constituent power function of the first clock frequency as a first nth order polynomial function. In one example, the first nth order polynomial function is a first cubic function of the first clock frequency. In one example, the first nth order polynomial function is a first quadratic function of the first clock frequency.
[0012] In one example, the method further includes determining the second constituent power function of the second clock frequency as a second nth order polynomial function. In one example, the second nth order polynomial function is a second cubic function of the second clock frequency. In one example, the second nth order polynomial function is a second quadratic function of the second clock frequency. In one example, the method further includes determining the allowable dc power characteristic as a function of a temperature.
[0013] These and other aspects of the present disclosure will become more fully understood upon a review of the detailed description, which follows. Other aspects, features, and implementations of the present disclosure will become apparent to those of ordinary skill in the art, upon reviewing the following description of specific, exemplary implementations of the present invention in conjunction with the accompanying figures. While features of the present invention may be discussed relative to certain implementations and figures below, all implementations of the present invention can include one or more of the advantageous features discussed herein. In other words, while one or more implementations may be discussed as having certain advantageous features, one or more of such features may also be used in accordance with the various implementations of the invention discussed herein. In similar fashion, while exemplary implementations may be discussed below as device, system, or method implementations it should be understood that such exemplary implementations can be implemented in various devices, systems, and methods.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] FIG. 1 illustrates an example information processing system.
[0015] FIG. 2 illustrates an example of a two-dimensional performance function for a plurality of clock frequencies.
[0016] FIG. 3 illustrates an example polynomial function of clock frequency.
[0017] FIG. 4 illustrates an example iteration procedure for determination of optimal clock frequencies under a thermal constraint.
[0018] FIG. 5 illustrates an example flow diagram for determination of clock frequencies under a thermal constraint.DETAILED DESCRIPTION
[0019] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well known structures and components are shown in block diagram form in order to avoid obscuring such concepts.
[0020] While for purposes of simplicity of explanation, the methodologies are shown and described as a series of acts, it is to be understood and appreciated that the methodologies are not limited by the order of acts, as some acts may, in accordance with one or more aspects, occur in different orders and / or concurrently with other acts from that shown and described herein. For example, those skilled in the art will understand and appreciate that a methodology could alternatively be represented as a series of interrelated states or events, such as in a state diagram. Moreover, not all illustrated acts may be required to implement a methodology in accordance with one or more aspects.
[0021] Elevated temperatures may limit electronic system performance. Every electronic component in an electronic system has a maximum allowable temperature or temperature constraint for its proper operation. One common active thermal control technique to comply with a temperature constraint is the usage of dynamic frequency scaling (DFS) . In DFS the electronic component clock frequency is adjusted according to the workload and thermal state of the electronic component. For example, the dc power consumption of the electronic component (e.g., CPU, DDR memory, etc. ) may be directly proportional to clock frequency. Since thermal dissipation is proportional to dc power consumption, temperature is directly dependent on the clock frequency. One example technique to maintain a proper thermal environment is to apply DFS without any consideration of performance degradation. When DFS lowers the clock frequency, the temperature is reduced, which is beneficial, but the performance of the electronic system may be degraded, due to the reduced operating speed.
[0022] An alternative active thermal control technique applies DFS with a performance optimization criterion. That is, the clock frequency is adjusted according to the workload, thermal state and performance. In one example, performance is directly related to dc power consumption.
[0023] In one example, stochastic gradient optimization (SGD) (a. k. a. stochastic gradient descent) is an optimization technique which may be employed in certain scenarios, for example, machine learning. In one example, SGD may be employed to determine an optimal operating condition in an electronic system within a plurality of temperature constraints. In one example, the optimal operating condition is a plurality of model parameters derived from a plurality of empirical parameters.
[0024] In one example, SGD may be formulated according to a multi-variable optimization procedure. For example, if a plurality of independent variables is denoted by a vector variable q with coordinates (q1, q2, …, qN) , an objective function (e.g., loss function) of q may be denoted by J (q) . In one example, the objective function J (q) may be a Euclidean norm squared or L2 norm squared of q, denoted as || q ||2 = q1 2 + q2 2 + qN 2. In one example, minimization the objective function J (q) is used to determine the plurality of model parameters In one example, the objective function J (q) measures a difference between the plurality of model parameters and the plurality of empirical parameters.
[0025] In one example, SGD is initialized with a random set of model parameters, represented by q0, and a sequence of a subsequent set of model parameters, represented by {q1, q2, …} . In one example, if qk represents a kth vector variable in an iteration, then the (k+1) th vector variable where = gradient of objective function J (q) evaluated at qk and η = learning rate parameter. In one example, the learning rate parameter η may be selected to balance optimization convergence and optimization speed. In one example, the gradient may be estimated using a numerical approximation.
[0026] In one example, SGD may be employed to determine an optimal setting of a plurality of clock frequencies within a plurality of temperature constraints. In one example, if a maximum temperature is reached, SGD may search a multi-dimensional energy efficiency model to determine the optimal setting of a plurality of clock frequencies. In one example, DFS may be used to adjust a plurality of clock frequencies according to the determined optimal setting of the plurality of clock frequencies.
[0027] FIG. 1 illustrates an example information processing system 100. In one example, the information processing system 100 includes a plurality of processing engines such as a central processing unit (CPU) 120, a digital signal processor (DSP) 130, a graphics processing unit (GPU) 140, a display processing unit (DPU) 180, etc. In one example, various other functions in the information processing system 100 may be included such as a support system 110, a modem 150, a memory 160, a cache memory 170 and a video display 190.
[0028] In one example, the support system 110 includes a controller (not shown) . In one example, the controller includes a logic circuitry or a microcontroller to manage dc power distribution through the information processing system 100. In one example, the controller manages the dc power distribution to the CPU 120 and the memory 160. In one example, the logic circuitry includes a clock frequency determination subsystem configured to adjust a plurality of clock frequencies according to an algorithm, for example, by using temperature data and / or dc power consumption data. Alternatively, dc current consumption data is used by the algorithm for adjusting the plurality of clock frequencies. In one example, the algorithm inputs a learning rate parameter η and a performance threshold δ. In one example, the performance threshold δ may be determined heuristically or may be determined by an a priori rule, for example, for a particular application. In one example, the learning rate parameter η may be selected to balance optimization convergence and optimization speed.
[0029] In one example, the plurality of processing engines and various other functions may be interconnected by an interconnection databus 105 to transport data and control information. In one example, the memory 160 and / or the cache memory 170 may be shared among the CPU 120, the GPU 140s and the other processing engines. In one example, the CPU 120 may include a first internal memory which is not shared with the other processing engines. In one example, the GPU 140 may include a second internal memory which is not shared with the other processing engines.
[0030] In one example, any processing engine of the plurality of processing engines may have an internal memory (i.e., a dedicated memory) which is not shared with the other processing engines. Although several components of the information processing system 100 are included herein, one skilled in the art would understand that the components listed herein are examples and are not exclusive. Thus, other components may be included as part of the information processing system 100 within the spirit and scope of the present disclosure.
[0031] In one example, one or more processing engines in the information processing system 100 may connect to a plurality of peripheral devices to provide additional functionality. The plurality of peripheral devices may include, for example, cameras, imagers, sensors, displays, speakers, microphones, etc. In one example, processor-peripheral device communications may be implemented by a bidirectional high-speed interface.
[0032] FIG. 2 illustrates an example of a two-dimensional performance function for a plurality of clock frequencies (e.g., two clock frequencies) 200. FIG. 2 shows a first clock frequency (e.g., central processing unit (CPU) frequency) on a first axis (e.g., x) 210, a second clock frequency (e.g., Double Data Rate (DDR) memory frequency) on a second axis (e.g., y) 220 and a performance metric on a third axis 230. In one example, the performance metric represents computing capacity, e.g., instructions per second, inferences per second, etc. In one example, FIG. 2 illustrates a performance function M (x, y) 240 which quantifies performance as a function of the first clock frequency (e.g., x) and the second clock frequency (e.g., y) .
[0033] In one example, an optimal operating condition for highest performance without a plurality of temperature constraints may be obtained by maximizing the plurality of clock frequencies. However, with a plurality of temperature constraints, a constrained optimal operating condition must account for thermal dissipation due to increased dc power consumption. In one example, the constrained optimal operating condition may be determined using SGD.
[0034] In one example, SGD may employ an iterative optimization procedure to achieve maximum performance with a plurality of temperature constraints. For example, SGD may iteratively adjust system parameters to minimize dc power consumption while maximizing performance. In one example, a vector variable q= (x, y) has a first variable x (e.g., first clock frequency) and a second variable y (e.g., second clock frequency) .
[0035] In one example, SGD adjusts system parameters iteratively using a gradient increment parameter Δ to numerically represent a gradient of the performance function M (x, y) at a first vector variable q1 = (x1, y1) to determine at a second vector variable q2 = (x2, y2) . In one example, the gradient increment parameter Δ may be computed as Δ = {Δx, Δy} = {M (x1, y1) / x1, M (x1, y1) / y1} or as a first-order derivative approximation. In one example, a second vector variable q2 = (x2, y2) may be determined via x2 = x1 + Δx and y2 = y1 + Δy.
[0036] In one example, the performance function M (x, y) at the second vector variable q2 may be computed and compared to the performance function M (x, y) at the first vector variable q1. In one example, if the performance function M (x, y) at the second vector variable q2 is greater than the performance function M (x, y) at the first vector variable q1 then the gradient increment parameter Δ improves performance. In one example, if the performance function M (x, y) at the second vector variable q2 is less than or equal to the performance function M (x, y) at the first vector variable q1 then the gradient increment parameter Δ degrades performance and a different gradient increment parameter Δd is computed and the procedure is repeated with the different gradient increment parameter Δd.
[0037] In one example, the gradient increment parameter Δ is an incremental fraction ε of a first range of the first variable x and of a second range of the second variable y. In one example, a first range of the first variable x is defined as a difference between a maximum value of the first variable x and a minimum value of the first variable x. In one example, a second range of the second variable y is defined as a difference between a maximum value of the second variable y and a minimum value of the second variable y. In one example, the gradient increment parameter Δ may have a first component Δx = [max (x) -min (x) ] *ε and a second component Δy = [max (y) -min (y) ] *ε, where the incremental fraction ε < 1. For example, the incremental fraction ε may be 0.01 or 0001.
[0038] In one example, the dc power consumption P of an electronic circuit is proportional to clock frequency f , capacitance C and square of voltage v via the following equation:
[0039] P = v2 f C.
[0040] That is, the dc power consumption P depends linearly on the clock frequency f for a given capacitance C and voltage v. In one example, the linear dependence of dc power consumption P on clock frequency f results in a correlation between voltage and clock frequency f. In one example, the dc power consumption P may be expressed as a polynomial function of clock frequency f with a plurality of coefficients {ak} via an Nth order polynomial of the form (with index k ranging from zero to N) : P (f) = Σk ak fk.
[0041] In one example, the Nth order polynomial is a third order polynomial of the form: P (f) = a3 f3 + a2 f2 + a1 f + a0.
[0042] In one example, the Nth order polynomial is a second order polynomial of the form; P (f) = a2 f2 + a1 f + a0.
[0043] In one example, the plurality of coefficients {ak} may be determined from empirical data and regression analysis (e.g., curve fitting of empirical data) .
[0044] FIG. 3 illustrates an example polynomial function of clock frequency 300. In one example, the example polynomial function has a first axis 310 (e.g., clock frequency) and a second axis 320 (e.g., dc power consumption P) . In one example a regressed polynomial curve 330 is shown as a function of the first axis 310 after curve fitting.
[0045] In one example, dc power consumption P of an electronic system may be expressed as a superposition of constituent power functions for electronic components in the electronic system. For example, for two constituents, the dc power consumption P (x, y) may be expressed as
[0046] P (x, y) = P1 (x) + P2 (y) , where
[0047] P1 (x) = first constituent power function of first clock frequency x,
[0048] P2 (y) = second constituent power function of second clock frequency y.
[0049] In one example, the first constituent power function may be expressed as a first Nth order polynomial and the second constituent power function may be expressed as a second Nth order polynomial. In one example, the first constituent power function may be expressed as a first cubic function of the form: P1 (x) = a3 x3 + a2 x2 + a1 x + a0,
[0050] with x = first clock frequency. In one example, the second constituent power function may be expressed as a second cubic function of the form: P2 (y) = b3 y3 + b2 y2 + b1 y + b0,
[0051] with y = second clock frequency. In one example, performance M (x, y) is maximized by a selection of the first clock frequency x = xopt and the second clock frequency y = yopt. That is, optimization occurs by maximizing performance M (x, y) under the thermal constraint P (x, y) ≤ h (T) by determining optimal clock frequencies x = xopt and y = yopt.
[0052] FIG. 4 illustrates an example iteration procedure for determination of optimal clock frequencies under a thermal constraint 400. In one example, the iteration procedure commences with a start block 410 where an initial plurality of clock frequencies results in an initial temperature Tinit. and continues with a temperature constraint check 420. If the initial temperature Tinit. ≤ a maximum allowable temperature Tmax, then proceed to a termination block 430 with optimal clock frequencies equal to the initial plurality of clock frequencies. If the initial temperature Tinit. > the maximum allowable temperature Tmax, then proceed to block 440. In one example, in block 440 update performance function M (x, y) with an updated plurality of clock frequencies determined by a gradient increment parameter Δ to determine an updated performance function Mupdate..
[0053] In one example, in block 450, calculate an updated loss function L (x, y) using the updated performance function Mupdate. In one example, in block 460, compare the updated loss function L (x, y) to a performance threshold δ. If the updated loss function L (x, y) is less than or equal to the performance threshold δ, then proceed to block 470 with the optimal clock frequencies equal to the updated set of clock frequencies and then proceed to the termination block 430. If the updated loss function L (x, y) is greater than the performance threshold δ, then proceed to block 440 to obtain a subsequent performance function M (x, y) with a subsequent set of clock frequencies determined by the gradient increment parameter Δ. In one example, continue with this iteration until block 460 results in a subsequent loss function L (x, y) being less than or equal to the performance threshold δ.
[0054] FIG. 5 illustrates an example flow diagram 500 for determination of clock frequencies under a thermal constraint. In block 510, determine an allowable dc power characteristic h (T) as a function of temperature T. In one example, an allowable dc power characteristic h (T) is determined as a function of temperature T. In one example, the allowable dc power characteristic h (T) is a monotonic function of temperature T. In one example, the allowable dc power characteristic h (T) is based on a thermal characteristic of CPU 120 or memory 160 in FIG. 1. In one example, the allowable dc power characteristic h (T) may be derived based on an empirical curve fit or by a thermal analysis simulation. In one example, the allowable dc power characteristic h (T) is a linear function of T wherein, h (T) = aT + b, where a and b are model parameters.
[0055] In block 520, determine a first constituent power function P1 (x) of a first clock frequency x as a first nth order polynomial function. In one example, a first constituent power function P1 (x) of a first clock frequency x is determined as a first nth order polynomial function. In one example, the first nth order polynomial function is a first cubic function of the first clock frequency x. In one example, the first nth order polynomial function is a first quadratic function of the first clock frequency x. In one example, the determination may be performed by CPU 120 or another processing engine in FIG. 1.
[0056] In block 530, determine a second constituent power function P2 (y) of a second clock frequency y as a second nth order polynomial function. In one example, a second constituent power function P2 (y) of a second clock frequency y is determined as a second nth order polynomial function. In one example, the second nth order polynomial function is a second cubic function of the second clock frequency y. In one example, the second nth order polynomial function is a second quadratic function of the second clock frequency y. In one example, the determination may be performed by CPU 120 or another processing engine in FIG. 1.
[0057] In block 540, initialize the first clock frequency x and the second clock frequency y with a randomly selected initial first clock frequency xinit and a randomly selected initial second clock frequency yinit. In one example, the first clock frequency x and the second clock frequency y are initialized with a randomly selected initial first clock frequency xinit and a randomly selected initial second clock frequency yinit. In one example, the initialization may be performed by CPU 120 or another processing engine in FIG. 1. In one example, the randomly selected initial first clock frequency and the randomly selected initial second clock frequency may be generated by a random number generator. Alternatively, they may be set to a minimum clock frequency, a middle clock frequency or a maximum clock frequency.
[0058] In block 550, compute a gradient increment parameter Δ for a stochastic gradient optimization (SGD) based on the first constituent power function P1 (x) evaluated at an updated first clock frequency x1 and the second constituent power function P2 (y) evaluated at an updated second clock frequency y1. In one example, a gradient increment parameter Δ for a stochastic gradient optimization (SGD) is computed based on the first constituent power function P1 (x) evaluated at an updated first clock frequency x1 and the second constituent power function P2 (y) evaluated at an updated second clock frequency y1.
[0059] In one example the gradient increment parameter Δ numerically represents a gradient of the first constituent power function P1 (x) and the second constituent power function P2 (y) . In one example, the gradient increment parameter Δ may be represented as a two-dimensional vector variable {Δx, Δy} . In one example, the gradient increment parameter Δ may be computed as Δ = {Δx, Δy} = {P1 (x) / x, P2 (y) / y} . In one example, the gradient increment parameter Δ may be computed as a first-order derivative approximation. In one example, the computation may be performed by CPU 120 or another processing engine in FIG. 1.
[0060] In block 560, determine an updated performance metric M (x, y) using the gradient increment parameter Δ at the updated first clock frequency x1 and the updated second clock frequency y1 and with a learning rate parameter η. In one example, an updated performance metric M (x, y) is determined using the gradient increment parameter Δ at the updated first clock frequency x1 and the updated second clock frequency y1 and with a learning rate parameter η. In one example, the updated first clock frequency x1 = xinit -ηΔx and the updated second clock frequency y1 = yinit -ηΔy . In one example, the learning rate parameter η may be selected to balance optimization convergence and optimization speed. In one example, the determination may be performed by CPU 120 or another processing engine in FIG. 1.
[0061] In block 570, compare an updated objective function J (x, y) based on the updated performance metric M (x, y) and the allowable dc power characteristic h (T) to a performance threshold δ. In one example, an updated objective function J (x, y) is compared based on the updated performance metric M (x, y) and the allowable dc power characteristic h (T) to a performance threshold δ. In one example, the updated objective function J (x, y) is a L2 norm squared (e.g., an Euclidean norm squared) of the updated first clock frequency x1 and the updated second clock frequency y1. If the updated objective function J (x, y) is less than or equal to the performance threshold δ, then terminate and select the first operational clock frequency to the updated first clock frequency and select the second operational clock frequency to the updated second clock frequency. If the updated objective function J (x, y) is greater than the performance threshold δ, proceed to block 560 and compute a second updated performance metric M (x, y) using the gradient increment parameter at a second updated first clock frequency x2 and an updated second clock frequency y2.
[0062] In one example, the comparison may be performed by CPU 120 or another processing engine in FIG. 1. In one example, the updated objective function J (x, y) is used to determine a plurality of model parameters In one example, the objective function J (x, y) measures a difference between the plurality of model parameters and a plurality of empirical parameters.
[0063] In one aspect, one or more of the steps for determining clock frequencies under thermal constraint in FIG. 6 may be executed by one or more processors which may include hardware, software, firmware, etc. The one or more processors, for example, may be used to execute software or firmware needed to perform the steps in the flow diagram of FIG. 6. Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0064] The software may reside on a computer-readable medium. The computer-readable medium may be a non-transitory computer-readable medium. A non-transitory computer-readable medium includes, by way of example, a magnetic storage device (e.g., hard disk, floppy disk, magnetic strip) , an optical disk (e.g., a compact disc (CD) or a digital versatile disc (DVD) ) , a smart card, a flash memory device (e.g., a card, a stick, or a key drive) , a random access memory (RAM) , a read only memory (ROM) , a programmable ROM (PROM) , an erasable PROM (EPROM) , an electrically erasable PROM (EEPROM) , a register, a removable disk, and any other suitable medium for storing software and / or instructions that may be accessed and read by a computer. The computer-readable medium may also include, by way of example, a carrier wave, a transmission line, and any other suitable medium for transmitting software and / or instructions that may be accessed and read by a computer. The computer-readable medium may reside in a processing system, external to the processing system, or distributed across multiple entities including the processing system. The computer-readable medium may be embodied in a computer program product. By way of example, a computer program product may include a computer-readable medium in packaging materials. The computer-readable medium may include software or firmware. Those skilled in the art will recognize how best to implement the described functionality presented throughout this disclosure depending on the particular application and the overall design constraints imposed on the overall system.
[0065] Any circuitry included in the processor (s) is merely provided as an example, and other means for carrying out the described functions may be included within various aspects of the present disclosure, including but not limited to the instructions stored in the computer-readable medium, or any other suitable apparatus or means described herein, and utilizing, for example, the processes and / or algorithms described herein in relation to the example flow diagram.
[0066] Within the present disclosure, the word “exemplary” is used to mean “serving as an example, instance, or illustration. ” Any implementation or aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects of the disclosure. Likewise, the term “aspects” does not require that all aspects of the disclosure include the discussed feature, advantage or mode of operation. The term “coupled” is used herein to refer to the direct or indirect coupling between two objects. For example, if object A physically touches object B, and object B touches object C, then objects A and C may still be considered coupled to one another-even if they do not directly physically touch each other. The terms “circuit” and “circuitry” are used broadly, and intended to include both hardware implementations of electrical devices and conductors that, when connected and configured, enable the performance of the functions described in the present disclosure, without limitation as to the type of electronic circuits, as well as software implementations of information and instructions that, when executed by a processor, enable the performance of the functions described in the present disclosure.
[0067] One or more of the components, steps, features and / or functions illustrated in the figures may be rearranged and / or combined into a single component, step, feature or function or embodied in several components, steps, or functions. Additional elements, components, steps, and / or functions may also be added without departing from novel features disclosed herein. The apparatus, devices, and / or components illustrated in the figures may be configured to perform one or more of the methods, features, or steps described herein. The novel algorithms described herein may also be efficiently implemented in software and / or embedded in hardware.
[0068] It is to be understood that the specific order or hierarchy of steps in the methods disclosed is an illustration of exemplary processes. Based upon design preferences, it is understood that the specific order or hierarchy of steps in the methods may be rearranged. The accompanying method claims present elements of the various steps in a sample order, and are not meant to be limited to the specific order or hierarchy presented unless specifically recited therein.
[0069] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, wherein reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more. ” Unless specifically stated otherwise, the term “some” refers to one or more. A phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover: a; b; c; a and b; a and c; b and c; and a, b and c. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims. No claim element is to be construed under the provisions of 35 U.S.C. §112, sixth paragraph, unless the element is expressly recited using the phrase “means for” or, in the case of a method claim, the element is recited using the phrase “step for. ”
[0070] One skilled in the art would understand that various features of different embodiments may be combined or modified and still be within the spirit and scope of the present disclosure.
Claims
1.An apparatus comprising:a controller configured to select a first clock frequency and a second clock frequency using a stochastic gradient optimization (SGD) ;a processor coupled to the controller, the processor configured to receive the first clock frequency; anda memory coupled to the controller, the memory configured to receive the second clock frequency.2.The apparatus of claim 1 wherein the controller is further configured to initialize the first clock frequency and the second clock frequency with a randomly selected initial first clock frequency and a randomly selected initial second clock frequency.3.The apparatus of claim 2, wherein the controller is further configured to compute a gradient increment parameter based on a first constituent power function evaluated at an updated first clock frequency and on a second constituent power function evaluated at an updated second clock frequency.4.The apparatus of claim 3, wherein the controller is further configured to determine an updated performance metric using the gradient increment parameter at the updated first clock frequency and at the updated second clock frequency and with a learning rate parameter.5.An apparatus for determining one or more clock frequencies under a thermal constraint, the apparatus comprising:means for initializing a first clock frequency and a second clock frequency with a randomly selected initial first clock frequency and a randomly selected initial second clock frequency;means for computing a gradient increment parameter based on a first constituent power function evaluated at an updated first clock frequency and on a second constituent power function evaluated at an updated second clock frequency; andmeans for determining an updated performance metric using the gradient increment parameter at the updated first clock frequency and at the updated second clock frequency and with a learning rate parameter.6.The apparatus of claim 5, further comprising:means for comparing an updated objective function based on the updated performance metric and an allowable dc power characteristic to a performance threshold;means for determining the first constituent power function of the first clock frequency as a first nth order polynomial function;means for determining the second constituent power function of the second clock frequency as a second nth order polynomial function; andmeans for determining the allowable dc power characteristic as a function of a temperature.7.A method comprising:initializing a first clock frequency and a second clock frequency with a randomly selected initial first clock frequency and a randomly selected initial second clock frequency;computing a gradient increment parameter based on a first constituent power function evaluated at an updated first clock frequency and on a second constituent power function evaluated at an updated second clock frequency; anddetermining an updated performance metric using the gradient increment parameter at the updated first clock frequency and at the updated second clock frequency and with a learning rate parameter.8.The method of claim 7, wherein the gradient increment parameter numerically represents a gradient of the first constituent power function and the second constituent power function.9.The method of claim 7, wherein the gradient increment parameter is represented as a two-dimensional vector variable.10.The method of claim 7, wherein the gradient increment parameter is computed as a first-order derivative approximation.11.The method of claim 7, further comprising selecting the learning rate parameter to balance an optimization convergence and an optimization speed.12.The method of claim 7, further comprising comparing an updated objective function based on the updated performance metric and an allowable dc power characteristic to a performance threshold.13.The method of claim 12, wherein the updated objective function is an Euclidean norm squared of the updated first clock frequency and the updated second clock frequency.14.The method of claim 12, further comprising determining the first constituent power function of the first clock frequency as a first nth order polynomial function.15.The method of claim 14, wherein the first nth order polynomial function is a first cubic function of the first clock frequency.16.The method of claim 14, wherein the first nth order polynomial function is a first quadratic function of the first clock frequency.17.The method of claim 14, further comprising determining the second constituent power function of the second clock frequency as a second nth order polynomial function.18.The method of claim 17, wherein the second nth order polynomial function is a second cubic function of the second clock frequency.19.The method of claim 17, wherein the second nth order polynomial function is a second quadratic function of the second clock frequency.20.The method of claim 17, further comprising determining the allowable dc power characteristic as a function of a temperature.
Citation Information
Patent Citations
Power state control of a mobile device
CN111065999A
Software assisted power management
CN112445529A
Adaptive on-chip digital power estimator
CN114424144A
Power balancing and configuration for static data centers
CN114816029A
Power consumption control method and device of GPU, equipment, medium and program product
CN115857655A