Adaptive power management for ai / ML accelerators

By dynamically adjusting voltage values based on monitored pre-processing times and slack time calculations, the system addresses inefficiencies in conventional DVFS, reducing frame drop rates and power consumption in data processing pipelines with nondeterministic operations.

WO2025216741A1PCT designated stage Publication Date: 2025-10-16GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/024107
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Conventional dynamic voltage frequency scaling (DVFS) approaches in integrated circuits fail to account for nondeterministic operations, leading to inefficient power consumption and decreased performance, particularly in data processing pipelines with varying operation times.

Method used

Dynamically adjust voltage values for target processing cores based on monitored pre-processing times and slack time calculations to ensure timely completion of operations while minimizing power consumption.

Benefits of technology

Reduces frame drop rates and overall power consumption, improving user experience and extending battery life in mobile devices by optimizing power management for nondeterministic operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024024107_16102025_PF_FP_ABST
    Figure US2024024107_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing dynamic voltage and / or frequency scaling. One of the methods includes obtaining an input data item; processing the input data item using a data processing pipeline executed on the computing device to generate an output data item, wherein processing the input data item comprises: executing a sequence of one or more pre-processing operations of the data processing pipeline on the input data item to generate a pre-processed data item; during the executing, monitoring a length of time consumed in executing the one or more pre-processing operations; and determining, based on the length of time consumed in executing the one or more pre-processing operations, a voltage value for a target processing core that will execute one or more subsequent processing operations of the data processing pipeline to process the pre-processed data item to generate a processed data item.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ADAPTIVE POWER MANAGEMENT FOR AI / ML ACCELERATORS

[0002] BACKGROUND

[0003] This specification relates to integrated circuit power management, and more particularly to controlling the voltage and / or frequency of signals supplied to integrated circuits on a chip.

[0004] For some systems using integrated circuits, performance of a workload can be improved by increasing clock frequency. One way of increasing clock frequency is to raise a voltage of the accelerator chip. However, this comes at the cost of increasing the temperature and power consumption of the chip, and potentially shortening longevity of the chip.

[0005] In order to strike a balance in the tradeoff between performance and power consumption, dynamic voltage frequency scaling (DVFS) is typically used to dynamically adjust clock frequency through voltage changes, such that clock frequency can be high during computation-heavy periods and low during lighter periods.

[0006] Conventional DVFS approaches typically rely on metrics such as utilization rate or throughput to select the voltage and / or frequency values at which a processing core will operate. Using conventional DVFS approaches to select the voltage values results in inefficient power consumption by the processing core and, in some cases, further results in decreased performance of the processing core, because those approaches fail to account for operations that have nondeterministic timing, i.e., operations that take a varying or nondeterministic time length to complete.

[0007] SUMMARY

[0008] This specification describes techniques for performing dynamic voltage and / or frequency scaling (“DVFS”). In particular, the specification describes techniques for dynamically setting a voltage value for a target processing core of a computing system that executes a data processing pipeline to conserve power consumption without decreasing the performance of the computing system.

[0009] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. Performance of a data processing pipeline that includes multiple operations where at least some of the operations will complete in nondeterministic time length can be improved when executed on a computing device. For example, frame drop rate in an image processing application that repeatedly performs an image processing pipeline can be reduced. Reduced frame drop rate may improve the quality of a stream of output image frames that is being generated by the application, which in turn results in improved user experience with the computing device.

[0010] Conventional DVFS approaches have drawbacks because they typically rely on utilization rate of the independent processing cores or system throughput to select the voltage values at which a target processing core will operate. However, as data processing pipelines become increasingly complicated, e.g., as they include a greater number of machine learning processing operations or other operations that have nondetermimstic timing, using conventional DVFS approaches to select the voltage values results in inefficient power consumption by the system, because those systems fail to account for the variations in the lengths of time taken by the target processing core to complete a corresponding operation in the data processing pipeline across multiple iterations of the data processing pipeline.

[0011] In contrast, using the techniques described in this specification to dynamically adjust the voltage values at which the target processing core will operate across multiple iterations of the data processing pipeline, the described system can reduce the frame drop rate while decreasing overall power consumption of the data processing pipeline. The described system thus extends battery life in mobile computing devices that rely on battery power.

[0012] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

[0013] BRIEF DESCRIPTION OF THE DRAWINGS

[0014] FIG. 1 is a block diagram of an example system.

[0015] FIG. 2 is an illustration of an example image processing pipeline.

[0016] FIG. 3 is a flowchart of an example process for processing an input data item using a data processing pipeline to generate an output data item.

[0017] FIG. 4 shows a quantitative example of the power consumption that can be saved byusing the system described in this specification.

[0018] Like reference numbers and designations in the various drawings indicate like elements.

[0019] DETAILED DESCRIPTION

[0020] FIG. 1 is a diagram of an example system 100. The system 100 includes a processing subsystem 110. The processing subsystem 110 includes a plurality of independent processing cores, e.g., processing core 110a, processing core 110m, processing core 11 On, and so on. The plurality of processing cores are generally capable of providing data processing capabilities, although they may be implemented in different ways and for different purposes.

[0021] Some common example implementations of such cores include: 1) a general-purpose programmable processing core that includes registers, control circuitry, and an arithmetic logic unit (ALU) intended for general-purpose computing; and 2) a special-purpose processing core having specialized hardware intended primarily for graphics computing, signal processing computing, encryption computing, and so forth.

[0022] The system 100 also includes a power management unit 120 for the processing subsy stem 110 and a battery 122. The power management unit 120 includes logic and components needed for regulating the power state and / or the operating frequency of the plurality of processing cores included in the processing subsystem 110. For example, the power management unit 120 controls which processing cores of the processing subsystem 110 receive power and how much power each processing core receives, e.g., by switching off or adjusting the voltage independently to each processing core.

[0023] In some implementations, the plurality of independent processing cores and the power management unit 120 are integrated onto a single system-on-a-chip (SOC). The SOC can be an integrated circuit that includes the aforementioned components, and possibly other components of the system, on a single silicon substrate or on multiple interconnected dies, e.g., using silicon interposers, stacked dies, or interconnect bridges.

[0024] The SOC is an example of a device that can be installed on or integrated into any appropriate computing device. Because the techniques described in this specification are particularly suited to reducing power consumption and increasing performance for the computing device, the SOC can be especially beneficial when installed on computing devices that rely on battery power, e.g., a smart phone, a smart watch or another wearable computing device, a tablet computer, or a laptop computer, to name just a few examples.

[0025] It is noted that the number of components of the SOC may vary from implementation to implementation. For example, there may be more processing cores than the number shown in FIG. 1 that are either included in the same processing subsystem or in separate processing subsystems.

[0026] The system 100 further includes a memory 130. The memory 130 can be any type of transitory or non-transitory computer readable medium capable of storing content accessible by the processing subsystem 110, such as volatile and non-volatile memory. For example, the memory 130 can store data that can be retrieved, manipulated, or stored by the processing subsystem 110, software applications (or “applications’' for short) that can be executed by the processing subsystem 110. or a combination thereof.

[0027] In the example of FIG. 1, the applications include an image processing application 132, which is stored in the memory 130 and executes in the processing subsystem 110. For example, the image processing application 132 can be or be included in a camera application, a video recording or enhancement application, a photo album application, a video conferencing application, a media player application, a social media application, a streaming application, an image filtering application, a video telephone call application, and the like that is capable of processing image data.

[0028] Each of the applications can include instructions that when executed cause the computing device to perform specific tasks or functions. The image processing application 132 can repeatedly execute an image processing pipeline on a stream of input image frames to generate a stream of output image frames, which can for example include a corresponding output image frame for each input image frame, and then present the stream output image frames on a display of the computing device, e.g.. in the order they become available.

[0029] The image processing pipeline executed by the processing application 132 can include multiple operations. Different operations of the image processing pipeline will generally be executed on different processing cores. In some implementations, each operation is executed on a respective one of the plurality of independent processing cores while in other implementations, two or more of the operations are executed on a same processing core. For example, in FIG. 1, the image processing pipeline can include a first operation that is executed on processing core 110a, followed by a second operation that is executed on processing core 110m, followed by a third operation that is executed on processing core 11 On, and so on.

[0030] Different operations of the image processing pipeline will generally take different time lengths to complete. For example, one of the multiple operations of the image processing pipeline can complete faster that another one of the multiple operations of the image processing pipeline. Moreover, at least one of the multiple operations of the image processing pipeline is a non-time-determimstic operation that may complete in a varying or nondeterministic time length, as opposed to others of the multiple operations of the image processing pipeline which can complete in deterministic, known time lengths.

[0031] The multiple operations of the image processing pipeline include one or more preprocessing operations, followed by a target operation, and, optionally, further followed by one or more post-processing operations. At least one of the pre-processing operations is a non-time-deterministic operation. The target operation is executed on a "target" processing core for which the operating voltage values can be selected. Unlike the pre-processing operations, in many situations, the target operation that is executed on the target processing core is a time-deterministic operation, or a relatively more time-deterministic operation than the pre-processing operations, e.g., the lengths of time needed to complete the target operation vary by a smaller amount than the length of time needed to complete any of the pre-processing operations.

[0032] In the example of FIG. 1, the data stored in the memory 130 includes target preprocessing time length data 134, and, optionally, in some implementations, historical performance data 136. The target pre-processing time length data 134 defines or otherwise specifies a target length of time in which to complete the one or more pre-processing operations of the image processing pipeline, i.e., the operations that precede the target operation that is executed on the target processing core.

[0033] In some implementations, the target pre-processing time length data 134 can be defined relative to the beginning of the image processing pipeline. Stated differently, the target pre-processing time length data 134 can define a target latency of the sequence of one or more pre-processing operations within the image processing pipeline. Suppose, for example, the processing core 110m in FIG. 1 is the target processing core for which the operating voltage values can be selected, and the second operation that is executed on the processing core 110m is the target operation of the image processing pipeline, then the target pre-processing time length data 134 can define a target length of time within which to complete first operation that is executed on processing core 110a.

[0034] There are many ways in which system 100 can obtain the target pre-processing time length data 134. For example, the target pre-processing time length data 134 can be input by a user of the system through a (virtual) keyboard or keypad. As another example, the target pre-processing time length data 134 can be received by the system over a data communication network from a remote system, e.g., a server system or another computing device. As yet another example, when the data stored in the memory 130 also includes the historical performance data 136. the target pre-processing time length data 134 can be generated by the system 100 based on the historical performance data 136.

[0035] The historical performance data 136 specifies the lengths of time that have previously been taken by one or more of the plurality of processing cores included in the processing subsystem 110 to complete the sequence of pre-processing operations of the image processing pipeline that were executed on the processing cores over a period of time, such as over the past hour, day, week, month, etc. For example, the target pre-processing time length data 134 can define, for processing core 110a. historical lengths of time it previously took to complete a corresponding pre-processing operation of the image processing pipeline that is being executed on the processing core 110a across multiple executions of an image processing pipeline when generating a stream of output image frames.

[0036] In some implementations, the historical performance data 136 is specific to the computing device on which the system 100 is implemented. That is, the system 100 monitors the lengths of time that have been taken by each of one or more of the plurality of processing cores included in the processing subsystem 110 to complete the corresponding pre-processing operation of the image processing pipeline, and store the monitoring data as the historical performance data 136 in the memory 130.

[0037] In other implementations, the historical performance data 136 is more holistic and offers more comprehensive insight into the performance of the processing cores. For example, a corresponding system implemented on each of a plurality of computing devices coupled to each other over a data communication network can each monitor such lengths of time taken by its corresponding processing cores, and then exchange the monitoring data with each other systems. In this way, the historical performance data 136 includes not only monitoring data specific to the computing device on which the system 100 is implemented, but also monitoring data specific to other computing devices.

[0038] There are many ways in which the system 100 can generate the target pre-processing time length data 134 based on the historical performance data 136. Moreover, the target preprocessing time length data 134 can be updated as often as necessary7based on the most recent historical performance data 136. For example, the system 100 can determine that a predetermined period of time, e.g., one hour, one day. one week, or one month, has elapsed since the target pre-processing time length data 134 was updated and, in response, automatically update the target pre-processing time length data 134 based on the historical performance data 136 that has been collected during that period of time. As another example, the system 100 can determine that the target pre-processing time length data 134 requires update or modification upon receiving an instruction from the image processing application 132 or another application installed on the system 100.

[0039] In some implementations, the system 100 can use a statistical value of the historical lengths of time that the one or more processing cores included in the processing subsystem 110 previously took to complete the sequence of one or more pre-processing operations of the image processing pipeline that are being executed on these processing cores across multiple executions of an image processing pipeline, as a target pre-processing length of time in which to complete the sequence of one or more pre-processing operations of the image processing pipeline.

[0040] For example, the statistical value can be a maximum, a minimum, a mean, a median, or a predetermined percentile value, e.g., a 90thpercentile value, a 95thpercentile value, a 99thpercentile value, etc. As a particular example, the 95thpercentile value means that, for 95% of all runs of the image processing pipeline, the historical lengths of time that have been taken by the one or more processing cores to execute the one or more pre-processing operations are below this value, and, for the remaining 5% of the runs, the historical lengths of time are above this value.

[0041] The data stored in the memory 130 further includes voltage value data 138. The voltage value data 138 defines, for the target processing core included in the processing subsystem 110, different voltage values at which the target processing core can be operating and, in some implementations, different clock frequencies at which the target processing core can be operating.

[0042] The target processing core can generally operate at two or more different voltage values. For example, the voltage value data 138 can define, for processing core 110m, a high voltage value, a middle voltage value, and a low voltage value at which the processing core 110m can be operating. As another example, the voltage value data 138 can define, for processing core 110m. a high voltage value and a low voltage value at which the processing core 110m can be operating.

[0043] Typically, changes in operating voltage for a processing core correspond to changes in operating frequency. For example, increasing the operating voltage value of the processing core will typically increase frequency, which can potentially improve throughput and latency, but at the cost of more power consumed. Conversely, decreasing the operating voltage value of the processing core will typically decrease frequency, which can potentially reduce power consumption but at the cost of lower throughput and latency.

[0044] In some implementations, the different voltage values at which the target processing core can operate can be predetermined, e.g.. by an administrator of the system 100 or the manufacturer of the processing core. In other implementations, the different voltage values at which the target processing core can operate can be set by the system 100 based on the target pre-processing time length data 134. and the voltage value data 138 can be optionally updated based on the most recent performance of the processing subsystem 110. During the execution of each iteration of the image processing pipeline, the power management unit 120 selects, for the target processing core included in the processing subsystem 110, a selected voltage value from the different voltage values defined in the voltage value data 138, and supplies a voltage at the selected voltage value to the target processing core that will execute the target operation of the image processing pipeline.

[0045] There are many ways in which the power management unit 120 can do this. As a general example, for of the target processing core included in the processing subsystem 110. the power management unit 120 can select voltage values at which the target processing core will operate, such that the target operation of the image processing pipeline that is being executed on the target processing core could complete within a certain time length after the target pre-processing length of time as defined by the target pre-processing time length data 134. Moreover, the power management unit 120 can select the voltage values while attempting to conserve power consumption of the processing cores.

[0046] Alternatively or additionally, the power management unit 120 can analogously select the frequencies at which a processing core will operate, since different operating frequencies generally correspond to different voltage values. Therefore, while the discussion below focuses on selecting voltage values, it should be noted that frequencies or some other dynamic voltage and frequency scaling (DVFS) settings that are dependent on the voltage values, can also be selected using the techniques described throughout this specification.

[0047] A concreate example of selecting voltage values for a target processing core while executing an image processing pipeline will now be described. It should be noted that the examples discussed below are for illustration purposes only, and the image processing pipeline can be implemented on a processing subsystem that include any number of processing cores that may or may not be mentioned below.

[0048] Moreover, the examples in this specification generally discuss an image processing pipeline for processing image data, e.g., to generate processed image data from raw image data. However, the same techniques can also be applied to any other data processing pipeline for processing any other forms of data items, including text data items and audio data items, to name just a few examples. In fact, the same techniques for selecting DVFS settings can more broadly be applied to any system that has one or more components or devices with separately settable DVFS settings (which need not necessarily include processing cores).

[0049] FIG. 2 is an illustration 200 of an example image processing pipeline 210. By repeatedly executing multiple iterations of the image processing pipeline 210, the system can generate a stream of output image frames from a stream of input image frames. A ‘'stream” of output image frames refers to an ordered sequence of output image frames that are being generated at a predetermined rate (or frequency), for example every 10ms, 50ms. 100ms, or the like.

[0050] In some situations, the quality of the stream of output image frames can be measured in terms of frame drop rate. When repeatedly performing the image processing pipeline 210 to generate the stream of output image frames, if an output image frame cannot be generated at the predetermined rate, the system may not display (or omit) the output image frame. This is called frame drop. For example, when an output image frame cannot be generated at the predetermined rate, e.g., due to the latency in the current iteration of the image processing pipeline 210 that is being performed, the system might choose to omit the frame without displaying it. and proceed to generate a next frame to be displayed. The quality of output image frames is typically in inverse proportion to the frame drop rate. That is, the uality increases as frame drop rate decreases.

[0051] In the example of FIG. 2, the image processing pipeline 210 includes a total of five operations. Each operation is executed on a respective one of a plurality of independent processing cores.

[0052] The image processing pipeline includes a first operation 221 that is executed on a first processing core (“IP block #1”). For example, the first processing core can be an image signal processor (ISP). The ISP can be configured to receive an input image frame, e.g., in the form of raw image data, from an image sensor, and process the input image frame into a form that is usable by later operations of the image processing pipeline. For example, the ISP can perform various image manipulation operations such as image translation operations, horizontal and vertical scaling, color space conversion, image stabilization transformations, and so on.

[0053] The image processing pipeline includes a second operation 222 that follows the first operation and that is executed on a second processing core (“IP block #2”). For example, the second processing core can be a graphics processing unit (GPU). The GPU can include graphics processing circuitry that can be configured to execute graphics software to perform a part or all of the graphics operation, or hardware acceleration of certain graphics operations.

[0054] The image processing pipeline includes a third operation 223 that follows the second operation and that is executed on a third processing core. For example, the third processing core can be an artificial intelligence / machine learning hardware accelerator. Artificial intelligence / machine learning hardware accelerators (or “AI / ML accelerators" for short) are computing devices having specialized hardware configured to perform specialized computations including computations for a machine learning workload. When the machine learning workload is known, the length of time the AI / ML accelerator will take to execute the computations for the machine learning workload can usually be pre-calculated, and may not vary significantly from one execution to another.

[0055] The image processing pipeline includes a fourth operation 224 that follows the third operation and that is executed on a fourth processing core (“IP block #n”). For example, the fourth processing core be a central processor unit (CPU). The CPU can be implemented using any suitable instruction set architecture, and can be configured to execute any instructions defined in that instruction set architecture.

[0056] The image processing pipeline includes a fifth operation 225 that follows the fourth operation and that is executed on a fifth processing core (“IP block #n+l”). For example, the fifth processing core can be a display processing unit (DPU). The DPU can be configured to read image data and generate display signals corresponding to the image data.

[0057] In the example of FIG. 2, the third processing core, i.e., the AI / ML accelerator, will be referred to as the “target” processing core for which the operating voltage values can be selected. Correspondingly, the first operation 221 and the second operation 222 that are executed on the first and second processing cores may each be referred to as a respective preprocessing operation of the image processing pipeline.

[0058] Further, FIG. 2 illustrates that the target operation executed on the target processing core is followed by the fourth operation 224 and the fifth operation 225, each of which may be referred to as a respective post-processing operation of the image processing pipeline. Thus, in the example of FIG. 2, the output image frame is generated after the fifth operation 225. In other examples, however, the target operation might be the last operation of the image processing pipeline, such that the output image frame is generated after the target operation.

[0059] During the executing of each iteration of the image processing pipeline, the system monitors a length of time consumed in executing the first operation 221 and the second operation 222. The system determines a slack time with reference to the target operation that will be executed on the target processing core as follows:

[0060] T[slack] =

[0061] T | target pre-processing time length] - T[monitored pre-processing time length]

[0062] That is, the system determines the slack time based on (i) the monitored preprocessing time length, i.e., the actual length of time that the first and second processing cores took to complete the first operation 221 and the second operation 222, and on (ii) the target time length within which to complete the first operation 221 and the second operation 222. The target time length is defined in the target pre-processing time length data that is stored in the memory’.

[0063] In various iterations of the image processing pipeline, the slack time may be positive, indicating that the monitored pre-processing time length is shorter than the target preprocessing time length; or may alternatively be negative, indicating that the monitored preprocessing time length is longer than the target pre-processing time length.

[0064] The system determines an additional length of time Tadd[i] for each different voltage value at which the target processing core can operate. The different voltage values are defined in the voltage value data that is stored in the memory 130.

[0065] Suppose, for example, the voltage value data defines that the target processing core can operate at three different voltage values, namely a high voltage value, a middle voltage value, and a low voltage value, then the system can calculate a first additional length of time Tadd[l] for the high voltage value, a second additional length of time Tadd[2] for the middle voltage value, and a third additional length of time Tadd[3] for the low voltage value.

[0066] There are many ways in which the system can determine the additional lengths of time. For example, the system can maintain, for each of one or more of the plurality of processing cores included in the processing subsystem, a mapping that maps (i) each voltage value at which the processing core might operate to (ii) a corresponding predetermined additional length of time that is an estimation of the time that would be needed in order for the processing core to complete the corresponding operation when it is operating at the voltage value, and then obtain the additional lengths of time in accordance with the mapping.

[0067] As another example, the system can calculate the additional lengths of time during the execution of the iteration of the image processing pipeline. For example, the additional lengths of time can be calculated by using a linear (or quadratic) function which linearly (or quadratically) scales the additional lengths of time with the operating frequency values of the processing core.

[0068] Then, the system selects a voltage value that has a corresponding additional length of time which satisfies the following:

[0069] T[slack] > Tadd[i]

[0070] Continuing with the example above, the system can select, as the selected voltage value at which the target processing core will operate, the lowest value from among the three different voltage values — i.e., the lowest value within the high voltage value, the middle voltage value, and the high voltage value — that has a corresponding additional length of time which satisfies this equation. For example, if Tadd[l] = 6ms, Tadd[2] = 8ms, Tadd[3] = 10ms, and T[target preprocessing time length] = 80ms, when Tfmonitored pre-processing time length] 71ms, then T[slack] = 80ms-71ms = 9ms. Correspondingly, the system selects the middle voltage as the selected voltage value at which the target processing core will operate when executing the target operation. In particular, the high and the middle voltage values have Tadd[l] = 6ms and Tadd[2] = 8ms, respectively, that are both smaller than T[slack] = 9ms. However the middle voltage is selected because it is lower than the high voltage value. The low voltage value has Tadd[3] = 10ms which is greater than Tslack = 9ms, hence it is also not selected.

[0071] Optionally, the system adds a time margin to accommodate for any potential inaccuracies in estimating the additional length of time Tadd[i] for each different voltage value at which the target processing core can operate. When the time margin is added, the equation that is used to facilitate the selection thus becomes:

[0072] T[slack] > Tadd[i] + Tmargin

[0073] FIG. 3 is a flow diagram of an example process 300 for processing an input data item using a data processing pipeline to generate an output data item. The data processing pipeline includes multiple operations that are executed on different processing cores. The multiple operations include one or more pre-processing operations followed by a target operation.

[0074] For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a system, e.g., the system 100 depicted in FIG. I. appropriately programmed in accordance with this specification, can perform the process 300.

[0075] In general, the system obtains an input data item and performs an iteration of the process 300 to generate an output data item. The system can repeatedly perform iterations of the process 300 on different input data items obtained by the system to generate a corresponding output data item for each obtained input data item. By repeatedly performing the process 300 on a stream of input data items, the system can generate a stream of output data items.

[0076] For example, the input data item can be a raw image frame in a sequence of raw image frames, and the output data item can be a final image frame in a sequence of final image frames. The final image frame can include, for example, an array of pixel data. The array of pixel data can include a plurality of channels, e.g., a plurality of colors channels, such as RGB channels.

[0077] The system executes a sequence of one or more pre-processing operations of the data processing pipeline on the input data item to generate a pre-processed data item (step 302). The one or more pre-processing operations can be executed on one or more processing cores. The one or more pre-processing operations are followed by the target operation, which can be executed on a target processing core. The target processing core can be an AI / ML accelerator, and the target operation can be an operation for a machine learning algorithm, for example.

[0078] At least one of the pre-processing operations is a non-time-deterministic operation that may complete in a varying or nondeterministic time length, such that the lengths of time the corresponding processing core takes to complete the pre-processing operation will generally vary from one iteration of the process 300 to another. Common examples of a non- time-deterministic operation include an operation that is dependent on a size of the input data item, an operation that consumes or otherwise relies on input data that arrives in a non-time- deterministic manner, or another operation for a random processing performed by a CPU or GPU, to name just a few.

[0079] During the execution of the sequence of one or more pre-processing operations, the system monitors a length of time consumed in executing the sequence of one or more preprocessing operations (step 304). That is, the system monitors how long the one or more processing cores takes to complete the corresponding pre-processing operations that precede the target operation of the data processing pipeline.

[0080] The system determines, based on the length of time consumed in executing the sequence of one or more pre-processing operations, a voltage value of the target processing core that will execute one or more subsequent processing operations including the target operation of the data processing pipeline to process the pre-processed data item to generate a processed data item (step 306).

[0081] After having selected the voltage value, the system supplies a voltage at the selected voltage value to the target processing core, and then executes the one or more subsequent processing operations including the target operation on the pre-processed data item while the target processing core operates at an operating frequency corresponding to the selected voltage value.

[0082] A specific example of selecting voltage values has been discussed above with reference to FIG. 2, but in general, the system can make this determination based on comparing (i) the monitored length of time consumed in executing the one or more preprocessing operations to (ii) a target length of time in which the one or more pre-processing operations of the data processing pipeline should complete.

[0083] Because at least one of the pre-processing operations is a non-time-deterministic operation and, correspondingly, the monitored lengths of time can vary from one iteration of the process 300 to another, the system will generally select different voltage values at different iterations of the process 300.

[0084] For example, when the monitored length of time is less than the target length of time by a threshold amount, and hence the slack time for the target operation is large enough, the system can select, from among a predetermined set of voltage values at which the target processing core might operation, a relatively lower voltage value for the target processing core, so as to adjust the target processing core to operate at a lower frequency to reduce power consumption.

[0085] Alternatively, when the monitored length of time is only less than the target length of time by no greater than the threshold amount, and hence the slack time for the target operation is limited, the system can select, from among the predetermined set of voltage values at which the target processing core might operation, a relatively higher voltage value for the target processing core, so as to adjust the target processing core to operate at a higher frequency to improve throughput and latency, but at the cost of more power consumed.

[0086] In implementations where the target operation is the last operation of the data processing pipeline, the processed data item can be used as the final output data item. Alternatively, in other implementations where the target operation is not the last operation of the data processing pipeline, i.e., where it is followed by one or more post-processing operations, the processed data item can be subsequently processed by the one or more postprocessing operations that follow the target operation to generate the final output data item. Like the pre-processing operations, those post-processing operations can be executed on one or more processing cores.

[0087] FIG. 4 shows a quantitative example of the power consumption that can be saved by using the system described in this specification. In the example of FIG. 4, the system can select, from among a high voltage value (UD), a middle voltage value (SUD), and a low voltage value (UUD), a selected voltage value at which a target processing core will operate, when executing a target operation of an image processing pipeline.

[0088] FIG. 4 illustrates that, when executing 455 iterations of the image processing pipeline to generate a stream of 455 output image frames, by using the described techniques to dynamically adjust the voltage values for the target processing core, the described system can save as much as 18.1% of power consumption relative to conventional systems which rely on metrics such as utilization rate or throughput to select the voltage and / or frequency values at which a processing core will operate (assuming the target pre-processing time length is set to be the 99thpercentile value of the historical lengths of time the processing cores previously took to complete the pre-processing operations), and at least 12.3% of power consumption relative to the conventional systems (assuming the target pre-processing time length is set to be the 90thpercentile value of the historical lengths of time the processing cores previously took to complete the pre-processing operations).

[0089] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.

[0090] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of. data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory' device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0091] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0092] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0093] In this specification, the term ’database" is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.

[0094] Similarly, in this specification the term “engine’7is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0095] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g.. an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0096] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0097] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.

[0098] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.

[0099] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.

[0100] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework or a Jax framework. Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0101] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g.. for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0102] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0103] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0104] What is claimed is:

Claims

CLAIMS1. A method performed by a computing device, wherein the method comprises: obtaining an input data item; processing the input data item using a data processing pipeline executed on the computing device to generate an output data item, wherein processing the input data item comprises: executing a sequence of one or more pre-processing operations of the data processing pipeline on the input data item to generate a pre-processed data item; during the executing, monitoring a length of time consumed in executing the sequence of one or more pre-processing operations; and determining, based on the length of time consumed in executing the sequence of one or more pre-processing operations, a voltage value for a target processing core that will execute one or more subsequent processing operations of the data processing pipeline to process the pre-processed data item to generate a processed data item.

2. The method of claim 1 , wherein determining the voltage value comprises: selecting, as the voltage value, one of a high voltage value, a middle voltage value, or a low voltage value.

3. The method of claim 2, wherein processing the input data item further comprises: supplying a voltage at the selected voltage value to the target processing core; and executing the one or more processing operations on the pre-processed data item while the target processing core operates at an operating frequency corresponding to the selected voltage value.

4. The method any one of claims 1-3, wherein processing the input data item further comprises: executing a sequence of one or more post-processing operations of the data processing pipeline on the processed data item to generate the output data item.

5. The method of any one of claims 1-4, wherein the sequence of one or more preprocessing operations comprise at least one non-time-deterministic operation, and whereinlengths of time consumed in executing the sequence of one or more pre-processing operations on different input data items are different.

6. The method of any one of claims 1-5, wherein selecting one of a high voltage value, a middle voltage value, or a low voltage value comprises: determining that the current length of time is less than a target length of time by a threshold amount; and in response, selecting the low voltage value as the selected voltage value.

7. The method of any one of claims 1-5, wherein selecting one of a high voltage value, a middle voltage value, or a low voltage value comprises: determining that the current length of time is greater than the target length of time by the threshold amount; and in response, selecting the high voltage value as the selected voltage value.

8. The method of any one of claims 5-7, wherein determining that the current length of time is less than the target length of time by the threshold amount comprises: maintaining historical performance data specifying historical lengths of time consumed in executing the sequence of one or more pre-processing operations on history input data items; and determining, from the historical performance data, the target length of time.

9. The method of claim 8, wherein maintaining history data comprises: receiving at least a portion of the historical performance data from another computing device over a data communication network.

10. The method of any one of claims 1-9, wherein the target processing operations comprise operations for a machine learning algorithm.

11. The method of any one of claims 1-10, wherein the target processing core comprises a hardware accelerator of a system-on-chip device.

12. The method of any one of claims 1-11, wherein the input data item comprises a raw image frame in a sequence of raw image frames, and wherein the output data item comprises a final image frame in a sequence of final image frames.

13. The method of claim 12, further comprising outputting the sequence of final image frames for display on the computing device.

14. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one more computers to perform the operations of the respective method of any one of claims 1-13.

15. One or more computer storage media storing instructions that when executed by one or more computers cause the one more computers to perform the operations of the respective method of any one of claims 1-14.

Citation Information

Patent Citations

  • Power consumption adjustment method and apparatus

    EP4328712A1

  • Serialization Floors and Deadline Driven Control for Performance Optimization of Asymmetric Multiprocessor Systems

    US20210349726A1

  • Intra-frame real-time frequency control

    WO2018191086A1