Idle state management of an soc and processing-in-memory architecture in a heterogeneous computing system

WO2025189181A8PCT designated stage Publication Date: 2025-10-02GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/019156
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-08
Filing Date
2025-03-10
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing computing systems face inefficiencies in managing state transitions between execution and idle modes in heterogeneous computing environments, particularly in memory devices with Processing-in-Memory (PiM) architectures, leading to suboptimal power consumption and increased latency during workload execution.

Method used

Implementing a method for managing state transitions in SoC and PiM blocks by determining optimal idle modes based on energy models and transition latencies, using control signaling to iteratively trigger mode changes, and proactively selecting idle modes to minimize power consumption and latency.

Benefits of technology

The method reduces power consumption and transition latency by dynamically optimizing idle modes, enhancing the efficiency and performance of heterogeneous computing systems, especially in executing machine-learning workloads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025019156_02102025_PF_FP_ABST
    Figure US2025019156_02102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computational instructions / programs encoded on non-transitory computer-readable media, are disclosed for decreasing state transition latency by determining a target type of idle mode. A system determines a first sequence of mode transitions. Each mode transition in the first sequence is a transition from an execution mode of to an idle mode of the memory device. The system determines, before occurrence of each mode transition in the first sequence, a type of idle mode from multiple among multiple mode types based on a dynamically updated energy model. The system iteratively generates and pre-programs a first set of control signals based on the corresponding type of idle mode for each mode transition in the first sequence, and based on the first set of control signals, the system iteratively triggers occurrence of each mode transition in the first sequence. Each triggered occurrence causes the memory device to transition to a corresponding type of idle mode.
Need to check novelty before this filing date? Find Prior Art

Description

IDLE STATE MANAGEMENT OF AN SOC AND PROCESSING-IN-MEMORY ARCHITECTURE IN A HETEROGENEOUS COMPUTING SYSTEMBACKGROUND

[0001] This specification generally relates to memory devices used to execute computations.

[0002] Modem computing systems often incorporate a wide variety of compute processing units that each offer different computing capabilities and trade-offs. Efficient execution of a given compute job often involves parsing computations into meaningful subtasks or workloads that are mapped to available processor cores of a computing system. The computations may be parsed and mapped based on suitability criteria, such as processor capability', performance, and power. Generally, this overall process of allocating portions of a computation to appropriate processor resources is referred to as a heterogeneous computation.

[0003] At least one processor core of the computing system can be an Intellectual Property block (“IP block”) that executes a respective portion of a computational operation for different multimedia workloads. Example use cases can involve processing image or speech data captured respectively by a camera or microphone on the mobile device as well as performing computations for generative artificial intelligence (“GenAI”) applications. The system-on-chip can use a heterogeneous compute operation to process input samples derived from image data, speech data, a text corpus, or a combination of these. An example step in the heterogeneous compute operation can include processing data associated with the input samples using a memory device that provides in-memory processing or computing capabilities.SUMMARY

[0004] This specification describes hardware and software techniques that manage state transitions for data processing resources within a memory' device coupled to a system-on-chip (“SoC”). In particular, the memory device can be a dynamic random-access memory' (DRAM) device that includes a Processing-in-Memory (PiM) architecture configured to perform memory-bounded and compute-bounded computations for executing an inference workload or task in the memory device.

[0005] The PiM architecture defines one or more PiM blocks of the memory device and each PiM block includes computing resources / elements, such as a processor unit, mode registers, and one or more computational units, e.g.. arithmetic logic units (ALUs) or relatedaddition and multiplication circuitry. For example, the PiM block can include discrete processors, processor units, register devices, buffers, and / or multiply accumulators (MACs) that cooperate to form one or more PiM compute elements. The PiM blocks are used to execute computations for an example workload that originates at the SoC. The computations can be segmented into respective portions that are allocated between the SoC and the memory device that includes the PiM blocks.

[0006] Each of the SoC and the memory device includes an execution mode and an idle mode. More specifically, each of the SoC and the PiM blocks of the memory device can iterate or transition between these two modes. In the memory’ device, the execution mode is for executing portions of the computations that are allocated to its PiM blocks, whereas the idle mode is for restricting power consumption at the memory’ device when no portion of the computations are allocated to a PiM block. For example, a workload can include computations that are required for at least a first task and a second task that are processed sequentially, where the first task is allocated to a PiM block in the memory device and the second task is allocated to the SoC. When the SoC uses or commands the PiM block to execute the computations for the first task, the PiM block can transition from an idle mode to an execution mode to execute the first task. Relatedly, the PiM block can transition from an execution mode to an idle mode in response to completing the computations for the first task.

[0007] An example workload can have a corresponding timeline (e g., a duration) along which the tasks and corresponding computations are performed to execute the workload. The memory device, and ultimately the PiM block(s), can transition back and forth from an execution mode to an idle mode along this timeline. An example system that implements the idle state management techniques can be configured to determine a ty pe of idle mode for mode transitions of the memory device. The types of idle modes include varying degrees of lowering voltages to the memory' device, refraining from maintaining execution contexts, or both. Additionally, the techniques include generating control signaling, that is passed from the SoC to the memory' device, to trigger mode transitions at a PiM block of the memorydevice. In this case, the SoC and the PiM block cooperate to implement aspects of a machine-learning ("ML”) model that executes an ML workload, which includes executing computations to perform / execute one or more tasks of the workload.

[0008] One aspect of the subject matter described in this specification can be embodied in a method implemented using an integrated circuit including an SoC and a memory device coupled to the SoC. The method includes: i) determining a first sequence of mode transitions, each mode transition in the first sequence being a transition from an executionmode of the memory device to an idle mode of the memory device; and ii) before occurrence of each mode transition in the first sequence: determining, from among multiple idle mode types, a type of idle mode that the memory device will transition to after occurrence of the mode transition. The method further includes: i) iteratively generating a first set of control signals based on the corresponding ty pe of idle mode for each mode transition in the first sequence, wherein an idle mode type selection is determined using a dynamically updated energy model; and ii) based on the first set of control signals, iteratively triggering occurrence of each mode transition in the first sequence, wherein each triggered occurrence causes the memory' device to transition to a corresponding type of idle mode.

[0009] These and other implementations can each optionally include one or more of the following features. For example, in some implementations, the method further includes determining a second sequence of mode transitions, each mode transition in the second sequence being a transition from an idle mode of the memory7device to an execution mode of the memory7device; and for each ty pe of idle mode of the memory7device: determining a respective transition latency for transitioning to an execution mode of the memory device from the corresponding type of idle mode of the memory device.

[0010] The method further includes iteratively generating a second set of control signals based on the corresponding transition latency for each mode transition in the second sequence; and based on the second set of control signals, iteratively triggering occurrence of each mode transition in the second sequence. Each triggered occurrence of a mode transition in the second sequence causes the memory device to begin executing computations immediately after a period of transition latency that corresponds to the type of idle mode from which the memory device is transitioning.

[0011] In some implementations, generating the first or second set of control signals includes: iteratively generating control signals using a memory controller of the SoC. Iteratively triggering occurrence of each mode transition in the first sequence can include: transitioning to an idle mode of the memory device to restrict power consumption at the memory device by clock gating or power gating a Processing-in-Memory C iM") block of the memory device. In some implementations, iteratively triggering occurrence of each mode transition in the second sequence includes: transitioning to an execution mode of the memory device to execute computations at the memory7device using the PiM block of the memory device.

[0012] Determining a first sequence of mode transitions can include, determining a first sequence of mode transitions based on execution durations and idle durations derived from acomputational graph for a processor of the PiM block. In some implementations, the multiple idle mode types includes one or more of: i) a retention idle mode; ii) a clock-gating idle mode; and iii) a power-gating idle mode. Iteratively generating the first set of control signals can include iteratively generating the first set of control signals over a time duration that coincides with a duration for executing a machine-learning (“ML”) workload using the SoC and / or a PiM block of the memory device. In some implementations, each control signal in the first set of control signals is used to trigger occurrence of each mode transition in the first sequence of mode transitions.

[0013] During the execution mode, a Processing-in-Memory (“PiM”) block of the memory device executes computations for a machine-learning (“ML”) workload, and the ML workload can be executed using both the SoC and the PiM block. In some implementations, determining the type of idle mode includes determining the type of idle mode based on energy overhead calculations generated by one or more dynamically updated energy models of the SoC.

[0014] The one or more dynamically updated energy models include a power & energy model, the method further includes: i) estimating, by the power & energy model of the SoC. a first transition energy consumption of a transition from an execution mode to an idle mode; ii) estimating, by the power & energy model of the SoC, an energy consumption of an idle mode duration corresponding to a length of time in the idle mode after transitioning to the idle mode from the execution mode; and iii) estimating, by the power & energy model of the SoC, a second transition energy consumption of a transition from the idle mode to a subsequent execution mode.

[0015] Another aspect of the subject matter described in this specification can be embodied in an integrated circuit that includes an SoC; a memory device coupled to the SoC; and a processor and a non-transitory machine-readable storage medium of the SoC for storing instructions that are executable by the processor to cause performance of certain operations. The operations include i) determining a first sequence of mode transitions, each mode transition in the first sequence being a transition from an execution mode of the memory device to an idle mode of the memory device; and ii) before occurrence of each mode transition in the first sequence: determining, from among multiple idle mode types, a type of idle mode that the memory' device will transition to after occurrence of the mode transition. The operations further include i) iteratively generating a first set of control signals based on the corresponding type of idle mode for each mode transition in the first sequence; and ii) based on the first set of control signals, iteratively triggering occurrence of each modetransition in the first sequence, where each triggered occurrence causes the memory device to transition to a corresponding ty pe of idle mode.

[0016] Yet another aspect of the subject matter described in this specification can be embodied in a method implemented using an integrated circuit comprising an SoC and a memory7device coupled to the SoC. The method includes identifying a mapping of respective portions of computations between the SoC and the memory device, wherein the memory device is external to the SoC; and determining, based on the mapping, multiple transition points that define a sequence of transitions between each portion of the computations mapped to the SoC and each portion of the computations mapped to the memory device.

[0017] The method further includes for each transition point of the multiple transition points and before occurrence of a transition defined by the transition point: i) determining a mode the SoC will transition to after occurrence of the transition; ii) determining a mode the memory7device will transition to after occurrence of the transition; and iii) selecting a particular type of idle mode that restricts power consumption at the memory device when the memory device will transition to an idle mode after occurrence of the transition.

[0018] As described, each of the preceding different aspects of the subject matter described in this specification can include a corresponding set of optional features. Each optional feature of one aspect applies similarly to. and can be included in, each of the other aspects described above.

[0019] Other implementations of this and other aspects include corresponding systems, apparatus, and computer programs, configured to perform the actions of the methods, encoded on computer storage devices. A system of one or more computers can be so configured by virtue of software, firmware, hardware, or a combination of them installed on the system that in operation causes the system to perform the actions. One or more computer programs can be so configured by virtue of having instructions that, when executed by7a data processing apparatus, cause the apparatus to perform the actions.

[0020] The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other potential features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Fig. 1 is a block diagram of an example computing system with at least one SoC.

[0022] Fig. 2 shows an example PiM architecture with corresponding compute elements.

[0023] Fig. 3 is an example timeline that shows respective sequences of processing and idle mode transitions for a host and memory device, and corresponding examples of idle mode types.

[0024] Fig. 4 is an example timeline of transitioning to one or more mode ty pes for an example memory’ device.

[0025] Fig. 5 is an example architecture for predicting execution modes and idle modes of a host device and a corresponding memory device.

[0026] Fig. 6 shows example graphical data indicating energy' overhead for different temperatures and operating voltages.

[0027] Fig. 7 is an example process for managing mode transitions for a memory device.

[0028] Fig. 8 is another example process for managing mode transitions for a memory device.

[0029] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0030] Fig. 1 is a block diagram of an example computing system 100 that includes a system-on-chip 102 (“SoC 102”). The SoC 102 includes a central processing unit 104 (‘"CPU 104”), a memory controller 105. a shared memory 106 (“memory 106”), a resource manager 108, and an IP / circuit block 110. In some implementations, system 100 can include multiple SoCs and any descriptions for the SoC 102 will apply equally to each of the multiple SoCs that may be included at system 100.

[0031] The CPU 104 can be a general-purpose CPU (e.g.. a single or multi-core CPU). The CPU 104 generates one or more indicators, such as an app-launch indicator or a function call that is triggered in response to executing or launching an application at a user device.For example, the application can be a camera application that uses an imaging sensor to generate image data or a gaming application that requires substantial memory’ and graphics processing resources to render graphical content of the game. The CPU 104 also generates one or more application values, such as pixel values or frame rate. The application values may be associated with a function call, may be descriptive of an event that occurs during execution of the application, or both.

[0032] The memory 106 is a system memory, shared memory, or both. In the example of Fig. 1, memory’ 106 is depicted external to circuit block 1 10. However, memory 106 caninclude portions of memory that are: i) specific to circuit block 110, ii) external to circuit block 110, or iii) both. The memory 106 can be random access memory of the SoC 102. such as static random access memory (SRAM), dynamic random access memory (DRAM), a synchronous DRAM (SDRAM), or double data rate (DDR) SDRAM.

[0033] In some implementations, aspects of memory 106 are configured as a shared scratchpad memory’ that supports parallel access of its memory resources by two or more processors of the circuit 110. Memory 106 can also include various other types of memory, such as high bandwidth memory’ (HBM), narrow memoiy (e.g., for storing 8-bit values), wide memory (e.g., for storing 16-bit or 32-bit values), or a combination of these.

[0034] The resource manager 108 is implemented in hardware and software. Aspects of the resource manager 108 can be also implemented as firmware of the SoC 102 or firmware of a device of the SoC 102, such as a DRAM memoiy' device or the CPU 104. The resource manager 108 is a processor-in-memory (PiM) resource manager (“PiM resource manager 108”) that includes control logic implemented in hardware, software, or both. For example, the PiM resource manager 108 can include resources such as flip-flops, registers, buffers, cache units, or a combination of these, which may be implemented in hardware, along with related control logic (e.g., programmed code), which may be implemented in software as well as hardware.

[0035] The circuit block 110 generally includes individual IP devices such as processors, processor cores, or special-purpose processing devices. For example, the circuit block 110 can include an image signal processor (ISP) 1 12, a host processing unit (HPU) 1 14, a digital signal processor (DSP) 116, and a graphics processing unit (GPU) 118. The circuit block 110 is referred to alternatively as an IP block 110, where the IP block can include one or more proprietary hardware elements. For example, each of the ISP 112, HPU 114. DSP 116, and GPU 118 can be a respective proprietary IP block (or IP device) of a particular entity or device manufacturer.

[0036] One or more aspects of the PiM resource manager 108 can be implemented as a software routine (or module) of the CPU 104, which uses one or more hardware resources of the CPU 104, such as registers, buffers, etc. The CPU 104 can be configured as an instruction and vector data processing engine that processes data obtained from a system memory of the SoC 102, such as memoiy7106. In some implementations, each processor (e.g., ISP 112, DSP 116, HPU 114, GPU 118) of the SoC 102 includes multiple cores and the CPU 104 and / or the PiM resource manager 108 can generate control signaling 124 to manage and distribute memory intensive compute operations to a memoiy device 122 (e.g., DRAM)to minimize the processing load at each core of the processors. The control signaling 124 is routed at system 100 using an example bus 120 of the SoC 102. The control signaling 124 can include commands, requests, data, instructions, or combination of these.

[0037] The PiM resource manager 108 cooperates with the CPU 104, memory controller 105 and storage controller 107 to dynamically control and manage one or more compute-inmemory (CIM) operations. In some implementations, the CIM operations are executed at the SoC 102 in support of a heterogeneous compute operation between two or more processing units that are included among the IP block 110, the CPU 104, or both. More specifically, the PiM resource manager 108 is configured to generate control signaling 124 and use one or more discrete signal values of the control signaling 124 to manage, configure, and / or boost data access operations at the memory device 122.

[0038] In general, techniques for idle state management implemented at system 100 can include generating data and control signaling at the SoC 102, which are used to communicate with the PiM architecture and memory device 122 by way of a memory controller 105 of the SoC. For example, the data and control signaling 124 generated at the SoC 102 are passed to, and processed by, compute elements of the PiM blocks 202 within a given PiM architecture 200 to execute the computations for an example ML inference workload. In some implementations, the data / control signaling 124 are processed at the PiM architecture 200 to trigger execution of certain data processing and computing operations, including proactive idle state management, using compute intervals that represent different pipeline stages of the PiM blocks 202(1) and 202(2). Control signals can also be generated locally at the memory device 122, which may be external to, or separate from, the SoC 102.

[0039] The system 100 includes an example memory device 122. The memory' device 122 can include multiple memory dies. For example, the memory device 122 can include N memory die, where N is an integer greater than 1. The memory device 122 can be a dynamic random-access memory (DRAM) or Double Data Rate (DDR) synchronous DRAM (SDRAM). The memory' device 122 is configured to perform or support various types of PiM operations, CiM operations, and memory-near-computing operations (“MnC operations’7). The memory device 122 performs or supports these operations using its multiple PiM compute elements, which are described below with reference to Fig. 2.

[0040] The SoC 102 cooperates with the memory' device 122 to perform computations across one or more bank groups of the memory device 122. The computations can be for operations or workloads that involve one or more of the processors at IP block 110. Additionally, the computations can be for a heterogenous operation that spans multipleprocessors of IP block 110, multiple IP blocks 110, or both. In at least one example the memory device 122 may be external to the SoC 102, whereas in another example the memory device 122 may be internal to the SoC 102. In some implementations, the SoC 102 is integrated (or co-located) with the memory device 122, for example, as distinct circuit die(s) that are co-located in a single integrated circuit package.

[0041] In the example of Fig. 1, system 100 and the SoC 102 is an integrated circuit of an example user / chent device 130, consumer electronic device, or mobile device, where each of these devices can include items such as a smartphone 130a, tablet 130b, laptop 130c, smartwatch or wearable device 130d. The devices 130 may also include other items such as an eNotebook, Netbook, smart speaker, or mobile computer. In some implementations, the system 100 and the SoC 102 are integrated circuits of a desktop computer, network server, or related cloud-based asset.

[0042] Fig. 2 shows an example processor-in-memory (PiM) architecture 200 for boosting PiM data access performance based on control signals generated using the SoC 102, the memory device 122, or both. In the example of Fig. 2, the memory device 122 includes a first memory die-1 with a first bank group that has multiple memory banks, where each memory bank includes one or more memory arrays and a second memory die-2 with a second bank group that has multiple memory' banks, where each memory' bank includes one or more memory arrays. In some implementations, the PiM architecture 200 includes multiple bank groups, multiple memory dies, or both. For example, a single memory die can include multiple bank groups and / or multiple bank groups can be distributed across multiple memory dies.

[0043] The PiM architecture 200 includes multiple PiM blocks 202, where each PiM block 202 includes multiple compute elements. For example, a first PiM block 202 of PiM architecture 200 includes mode register 204-1 and process unit 206-1, whereas a second, different PiM block 202 of PiM architecture 200 includes mode register 204-2 and process unit 206-2. Each process unit 206-1, 206-2 can include a processor, a processor unit, or a processor core, such as a CPU. Each process unit 206-1, 206-2 can also include an example computation unit such as an arithmetic logic unit (ALU) or multiply-accumulate cell (MAC).

[0044] In some implementations, the PiM architecture 200 is included in the memory device 122 as multiple discrete integrated circuits, where each integrated circuit is local to a given memory die (e.g., die-1 and die-2) and interacts or communicates with arrays of memory cells at that memory die. For example, the PiM architecture 200 can include compute elements that are replicated and distributed across each of the memory die in thememory device 122. In some other implementations, the PiM architecture 200 is included in the memory device 122 as a single integrated circuit that interacts or communicates with each memory die of the memory device 122, including the arrays of memory cells at each memory die.

[0045] To boost PiM data access performance as described in this specification, the PiM blocks 202 or process units in the PiM architecture 200 are located within the memory device 122 but outside of a section of the memory device 122 that includes the bank groups. The section may be defined as a discrete memory die or defined in some other way (e.g., a portion of a memory die). Irrespective of the hardware configuration or layout of PiM architecture 200, the PiM blocks 202 are sufficiently external to the bank groups such that the PiM blocks 202 communicate with the bank groups based on a particular timing constraint that can be leveraged to boost PiM data access performance with cross bank group data aggregation.

[0046] The PiM operations can include standard CPU functions, whereas the CiM operations and MnC operations can include standard arithmetic operations, such as computations normally performed by an ALU or MAC. The CiM operations and MnC operations can also include computational functions of a HPU 114. such as multiplication and addition operations for matrix math, vector computations, linear algebra, and dot-product accumulations. In some implementations, each of the PiM operations, CiM operations, and MnC operations are performed in support of ML computations, neural network computations, or both.

[0047] In some implementations, the PiM operations are an extension of the computational functions of the HPU 114. For example, a PiM block 202 can generate accumulated values from sets of weight values / inputs and activation inputs obtained from memory banks of different bank groups based on a particular timing constraint. The timing constraint can be a minimum delay required between successive column commands issued to different bank groups within a DRAM device. The accumulated values are generated based on neural network computations performed using a computational array of the PiM block 202. The computational array can be a matrix multiplication unit with compute cells that are arranged as a systolic array. The accumulated values can be dot products of the sets of weight values and the activation inputs. That is, for a set of weights, the PiM block 202 multiplies each weight with each activation input and sums the products together to form an accumulated value.

[0048] The PiM architecture 200 can include a register or other portion of memory for storing data for a respective memory die or group of memory banks. For example, the datacan be mode / configuration values. The data can also describe errors that occurred during a compute operation at a corresponding PiM block 202 of the memory device 122, or both. In some implementations, the register or other portion of memory is used to store configuration information, or associated instructions, for configuring aspects of a PiM block, or respective memory die, group of memory banks, or a combination of these.

[0049] For example, the mode registers 204-1, 204-2 can be used to control or trigger selection of a particular mode in a PiM architecture, such as an error-capture mode, interleave configuration mode, multi-batch processing mode, etc. In some implementations, a particular mode is selected based on bit values of the mode registers 204-1, 204-2. For example, to trigger or select an interleave configuration mode(s) or multi-batch processing mode(s), a single bit, or a sequence of bits, can be defined for use in the mode register.

[0050] Fig. 3 is an example timeline of transitioning to one or more mode types for example memory' devices. In the example of Fig. 3, the timeline 300 represents an operation timeline for a host device of the SoC 102 and a corresponding PiM block 202 that may be coupled to the host device by way of the memory device 122. as described above with reference to Fig. 1 and system 100. In some implementations, the host is the SoC 102 or any device of the SoC 102. For example, the host device can be one or more of the proprietary7devices of IP block 110, such as the HPU 114.

[0051] The timeline 300 includes time durations corresponding to an idle mode 302 and execution (or processing) mode 304. The time durations corresponding to idle mode 302 indicate that the host or the PiM block 202 is not actively performing a data processing function, such as executing computations using various sets of operands or executing memory7bound tasks in support of a workload. The ty pes of idle modes can include a clock gating idle mode 314, a state retention idle mode 316. and a power gating idle mode 318, each of which is described below. Other types of idle modes are also within the scope of this disclosure. In some implementations, the PiM block 202 transitions to a restricted power consumption idle mode based on an expected duration of the idle mode, as described below.

[0052] Each of modes 314, 316, and 318 are power reduction measures that are employed at system 100 to minimize power consumption during a non-execution mode of the system, where data processing functions and computations are not actively being executed. Each of modes 314, 316, and 318 can provide a corresponding measure (or amount) of power reduction. For example, the clock-gating idle mode 314 can provide a first measure of power savings, the state retention idle mode 316 can provide a second measure of power savings that is more than the power savings provided by the first measure, and the power gating idlemode 318 can provide a third measure of power savings that is more than the power savings provided by the first or second measures.

[0053] The clock-gating idle mode 314 is an idle mode type in which the system 100 gates one or more clock signals generated at a PiM block, decreases a level of voltage applied to components of a corresponding device, or both. For example, the system 100 can disable (e.g., clock-gate) a clock signal(s) for a particular circuit or portion of circuitry at a HPU 114 or the PiM block 202. In general, clock gating is a technique used in integrated circuit design to reduce power consumption by selectively disabling clock signals to unused circuitry of the integrated circuit. The approaches for clock gating and other power reduction techniques disclosed in this specification can be leveraged to significantly improve the power consumption, performance, and overall reliability of integrated circuit designs of SoCs and integrated memory units for edge devices.

[0054] In some implementations, the system 100 (e.g., the SoC 102 or a host device) is configured to maintain an operating context of an execution mode of a PiM block 202 during an example clock-gating idle mode 314. For example, a processor (or processing device) of the PiM block 202 can receive a threshold level of voltage that enables the PiM block 202 to maintain an operating context of an execution mode for the PiM block 202. In this example, the minimum threshold voltage for maintaining an operating context is lower (e.g., substantially lower) than the nominal voltage levels applied during an execution mode of the PiM block 202. In some implementations, one or more operating contexts of a PiM block are maintained concurrent with clock -gating a portion of circuitry in a PiM block 202 as a power reduction measure of the clock-gating idle mode 314. The operating context that is maintained can be a context for a corresponding execution mode that precedes occurrence of the clock-gating idle mode 314.

[0055] Each of modes 314, 316, and 318 has a respective corresponding latency penalty, which is an amount of time that is required to enter and exit that idle mode. For example, the clock gating idle mode 314 includes an entry' time duration 306 and an exit time duration 308. As indicated, collectively the entry’ and exit time durations 306. 308 indicate a duration of time that is required for a host at the SoC 102 or the memory device 122 (e.g., the PiM block 202) to enter the clock-gating idle mode 314 and to exit the clock-gating idle mode 314. In some implementations, the respective time that elapses for each of the entry and exit time durations 306, 308 of the clock-gating idle mode 314 is a latency penalty incurred at system 100 in response to transitioning to the clock-gating idle mode 314.

[0056] The state retention idle mode 316 is an idle mode type where system 100 implements state retention features to enable a previously executing device, such as the PiM block 202, to maintain (or retain) an operating state / context associated with its execution environment. In some implementations, the power savings of the state retention idle mode 316 are realized at least by decreasing the level of voltage that is applied at a device relative to the voltage levels applied during an execution mode of the device.

[0057] In particular, the system 100 (or SoC 102) can employ additional clock-gating functions to reduce the switching frequency at a particular circuit or portion of circuitry in, for example, the HPU 114, the PiM block 202, or both. The system 100 can decrease the applied voltage levels and reduce the switching frequency while also maintaining a minimum threshold level of power required to retain the operating context. These aspects of the state retention idle mode 316 can be leveraged to achieve a higher amount of power savings relative to the clock gating idle mode 314. The state retention idle mode 316 incurs an entry time duration 310 and an exit time duration 312 that are relatively longer than the entry time duration 306 and the exit time duration 308 for clock-gating idle mode 314.

[0058] The power gating idle mode is a type of idle mode where the system 100 provides no voltage and no current to the corresponding device such as the PiM block 202 and does not maintain the operating context of the execution for the PiM block 202. In particular, the system 100 disables (e g., power-gates) power to the PiM block 202 or to certain compute elements of the PiM block 202, resulting in a relatively high amount of power savings relative to clock gating idle mode 314 and retention idle mode 314. The power gating idle mode 318 incurs an entry time duration 320 and an exit time duration 322 that are relatively longer than the entry time duration 310 and the exit time duration 312 for state retention idle mode 316.

[0059] As such, transitioning from the power gating idle mode 318 to the processing mode 304 incurs a greater latency penalty than transitioning from the state retention idle mode 316 to the processing mode 304. In some implementations, transitioning from the power gating idle mode 318 incurs a greater latency penalty because system 100 must supply power to the PiM block 202 and re-establish certain power states and operating contexts to actively perform tasks during the execution mode 304.

[0060] The system 100 can determine a sequence of mode transitions for the PiM block 202 based on the w orkload for the PiM block 202. In particular, the sequence of mode transitions includes one or more idle mode time durations 302 interspersed with multiple processing mode (e.g., execution mode) time durations 304, where each of the processingmode time durations 304 represent a time duration for the PiM block 202 to execute actions associated with the workload. For example, the workload can be a ML workload executed using a large language model (LLM) for a GenAI application. The ML workload is sufficiently deterministic such that the system 100 can accurately determine an optimal sequence of mode transitions for a device such as the PiM block 202, prior to the PiM block 202 being used to execute the workload.

[0061] In some implementations, the system 100 can determine the processing mode time durations 304 and the idle mode time durations 302 based on a computational graph that indicates the type and number of computations the PiM block 202 will be requested to perform for a given ML workload. The computational graph can be based on a simulation or an analytical model that predicts and evaluates computational sequences employed by a neural network ML model to execute different types of ML workloads. The computational graph can indicate an execution time duration, computational intensity at different layers of the neural network, or both.

[0062] Based on determining the sequence of mode transitions, the system 100 determines an optimal idle mode type from the types of idle modes (e.g.. clock gating idle mode 314, state retention idle mode 316, or power gating idle mode 318) for each of the mode transitions. In particular, the system 100 determines an optimal idle mode type for the PiM block 202 to transition to after each processing mode time duration 304 based on the length of the subsequent idle mode time duration 302. The system 100 then generates a set of control signals for each mode transition that indicates the optimal idle mode type, and the system 100 triggers each mode transition for the PiM block 202 corresponding to the respective optimal idle mode ty pe to decrease transition latency from idle mode to execution mode, as described in further detail below with reference to Fig. 4.

[0063] Fig. 4 is an example timeline of transitioning to one or more mode types for an example memory device. In the example of Fig. 4, the timeline 402 includes a conventional operation timeline for an example PiM block 202, which includes time durations 406 of an execution mode 418 (e.g., an active execution mode) that is part of the processing mode 302 for the PiM block 202. time durations 408 of a clock gating idle mode 314. time durations of a state retention idle mode 316, and time durations 414 of a power gating idle mode 318. Each of the time durations for each of the modes indicate a particular timeframe in which the PiM block 202 is operating in the particular mode. Additionally, the timeline 402 includes timeout time durations 410 for a particular mode to transition to another mode, whereas the timeline 402 includes transition time durations 412 that indicate the mode transitions.

[0064] In conventional systems, the system 100 implements timeout mechanisms (e.g., timeout time durations 410) to trigger a memory’ device 122, such as a PiM block 202, to transition to a type of idle mode. In particular, the PiM block 202 can be constantly operating between execution mode 418 and clock gating idle mode 314 (e.g., a relatively’ low latency idle mode) for time durations 406 and 408, respectively, to perform a task. The PiM block 202 can then be operating in the clock gating idle mode 314 for a time duration greater than a threshold (e.g.. the timeout time duration 410-A), resulting in the system 100 to implement the timeout mechanism. In this case, the system 100 triggers the PiM block 202 to transition from the clock gating idle mode 314 to the state retention idle mode 316 during the transition time duration 412-A. In particular, the system 100 can trigger the mode transition by generating a control signal and transmitting the control signal to the PiM block 202, where the control signal indicates for the PiM block 202 to transition to a particular ty pe of idle mode.

[0065] Conventionally, the system 100 triggers the PiM block 202 to enter idle modes in a consecutive manner (e.g., from clock gating mode 314 to retention mode to power gating mode 318) based on the timeout time durations. For example, the PiM block 202 may then be operating in the state retention idle mode 316 for a time duration greater than a threshold (e.g., the timeout time duration 410-B), resulting in the system 100 to again implement the timeout mechanism by triggering the PiM block to transition from the state retention idle mode 316 to the power gating idle mode 318 during the transition time duration 412-B. By transitioning to idle mode types, the system 100 conserves power by clock gating and power gating the PiM block 202 based on the type of idle mode.

[0066] However, transitioning from relatively deeper idle modes, such as retention idle mode 316 and power gating mode 318, to execution mode 418 can result in high levels of transition latency and overall system latency. For example, the PiM block 202 can be operating in the power gating idle mode 318 for the time duration 414, and the system 100 may then trigger the PiM block 202 to re-enter the execution mode 418 to perform a task. Because of the exit time duration (e.g., exit time duration 322) associated with the power gating idle mode 318, along with having to re-establish the operating context for the PiM block 202, the transition time duration 412-C is relatively long, resulting in increased latency in the PiM block 202 returning to execution mode 418 and performing the task.

[0067] In contrast, the timeline 404 illustrates an example timeline implementation for the system 100 to efficiently manage mode transitions of the memory device 122 from an idle mode to an execution mode. Based on a particular task, the system 100 can decrease latencyfor mode transitions by determining a sequence of mode transitions by selecting an optimal idle mode type for an executing device (e.g., PiM block 202) that transitions from an execution mode to the idle mode. The system 100 can anticipate an upcoming transition from a power gating idle mode 318 to an execution mode and proactively trigger aspects of the transition to decrease the overall latency penalty when a device transitions from the selected idle mode to the execution mode.

[0068] In the example of Fig. 4, the timeline 404 includes an improved operation timeline for an example PiM block 202, which includes time durations 406 of the execution mode 418 for the PiM block 202, time durations 408 of the clock gating idle mode 314, and time duration 414 of a power gating idle mode 318. Each of the time durations for each of the modes indicate a particular timeframe in which the PiM block 202 is operating in the particular mode. Additionally, the timeline 404 includes transition time durations 412 that indicate the system 100 triggering a mode transition of the PiM block 202, and the timeline 404 includes a decreased latency time duration 416 in comparison to the timeline 402.

[0069] Prior to entering the execution mode 418, the system 100 can determine the sequence of mode transitions for performing a workload of a particular task. For example, the task can include executing a ML workload using the SoC and the PiM block 202, and the system 100 can determine the sequence of mode transitions from analysis performed on a computational graph associated with execution of the workload. The sequence of mode transitions can include a first sequence of transitions from the execution mode 418 to a type of idle mode and a second sequence of transitions from a type of idle mode to the execution mode 418.

[0070] For each of the mode transitions, the SoC 102 (or host device) can select an optimal type of idle mode based on a determined transition latency. The SoC 102 can then iteratively generate control signals based on the selected type of idle mode. Using the control signals, the system 100 can iteratively trigger occurrence of each mode transition.Additionally, based on the determined sequence of mode transitions, the SoC 102 can preemptively transition from the idle mode to the execution mode 418 to decrease a total latency required for the PiM block 202 to reenter the execution mode 418.

[0071] For example, the system 100 can determine that the PiM block 202 must operate between execution mode 418 and clock gating idle mode 314 (e.g., a relatively low latency idle mode) for time durations 406 and 408, respectively, to perform the task. The system 100 can also determine that the PiM block 202 can enter a deep idle mode (e.g., power gating idle mode 318) for a longer time duration (e.g., time duration 414) because the PiM block 202 isnot required to operate or perform computations for the task during the time duration 414. In particular, the system 100 can trigger the mode transition 412-D from the execution mode 418 to the power gating mode 318 by generating a control signal and transmitting the control signal to the PiM block 202 to enter the corresponding idle mode.

[0072] Additionally, the system 100 can determine that the PiM block 202 must transition from the power gating idle mode 318 to execution mode 418 to perform a particular task, and the system 100 can determine the transition latency (e.g., decreased latency time duration 416). Transition latency can be represented using clock cycles, time durations, or both. For example, the resource manager 108 and / or HPU 114 can determine that the transition latency from power gating idle mode 318 to execution mode is 418 is approximately 200 ms. In some implementations, proactive transition can occur -100 ms early to reduce the prior transition latency by about (or at least) 50%, e.g., from -200 ms to -100 ms.

[0073] Control logic in system 100 for proactive idle state management can proactively trigger re-establishing a prior operating context, powering up certain logic gates, and / or initializing control circuitry required by execution mode operations to reduce the overall transition latency, e.g.. from 200 ms to 100 ms or from 200 ms to between 50 ms and 100 ms. In some implementations, for a portion of circuitry that is power-gated or powered-down, the transition latency is reduced in proportion to the time required to initialize or activate that portion of circuitry, such that it can be immediately used execute an computation or operation used in generating an ML inference output.

[0074] The transition latency can be determined using the resource manager 108, the HPU 114, the memory7controller 105, or a combination of these. In some implementations, the transition latency is determined from encoded transition latency values generated during compile time by an example compiler software stack of system 100. For example, the compiler can determine the transition latencies based on idle mode instances and execution mode operations that are assigned to the SoC 102 and a PiM block 202. In some implementations, the idle mode instances and execution mode operations assigned to the SoC 102 and a PiM block 202 are obtained by the compiler during compile-time.

[0075] For example, the compiler obtains this mode information based on a data flow graph of computations that indicates a set of deterministic operations and computations that are required to execute an ML model using the HPU 114 and / or the PiM block 202. For instance, using an example performance model of a compiler software stack, the system 100 can determine, e.g., from the ML computation graphs, that a particular sequence of mode transitions 330 will occur for an example set of computations at the PiM block, including thecorresponding transition latencies for each mode transition. This is described in more detail below with reference to Fig. 5.

[0076] Based on the determined transition latency, the system 100 can pre-emptively transmit a control signal to the PiM block 202 to trigger the mode transition 412-E at a time duration (e.g., the time duration of the decreased latency time duration 416) prior to performing the particular task. In some examples, the system 100 generates the control signals using memory controller 105 of the SoC 102. Thus, by implementing the described techniques, the system 100 can decrease the overall system latency by managing mode transitions in order to perform a particular task.

[0077] Fig. 5 is an example architecture for predicting execution modes and idle modes of a host device 514 of the SoC 102 and a corresponding PiM block 202 of the memory device 122. In the example of Fig. 5, the architecture 500 includes a compiler software stack 520, a host 514, and the memory device 122. In this example, memory device 122 can be a low-power DDR (“LPDDR’ ) random access memory device. The compiler software stack 520 is configured to generate power mode and transition information 522, 524 for the host device 514 and PiM block 202, respectively (descnbed below). The compiler software stack 520 includes a ML model network architecture 502, a performance model 504, a power and energy model 506, model binary / instructions 508, and host firmware 510. The host 514 includes a host power management unit 516, an SoC power management unit 518, and a memory controller 105, and the LPDDR 122 includes a PiM block 202.

[0078] The system 100 uses the architecture 500 to manage mode transitions at the host device 514, the PiM block 202, or both. For example, each of the PiM block 202 and a host device, such as the HPU 114, can be tasked with executing an ML workload using a neural network ML model that is based on a neural network implemented on the HPU 114. The ML model network architecture 502 is representative of an example architecture of the neural network implemented on the HPU 114. For example, the neural network can be a convolutional neural network (“CNN’') and an architecture of the CNN can be 3 layers comprising an input layer, a hidden layer, and an output layer. In some implementations, the ML model neural network architecture 502 is generated using an example ML platform such as TensorFlow® or TensorFlowlite®.

[0079] The performance model 504 generates the performance data by modeling or simulating dataflow and ML computation graphs for the ML model neural network architecture 502. Likewise, the power and energy model 506 generates the corresponding energy consumption estimates by modeling or simulating power / energy profiles for the MLmodel neural network architecture 502 being implemented in hardware. For example, the power and energy model 506 can generate the corresponding energy consumption estimates with reference to an integrated circuit or hardware accelerator that implements the underlying neural network of the ML model.

[0080] Using the performance model 504, the system 100 can determine, e.g., from the ML computation graphs, that a particular sequence of mode transitions 330 will occur for an example set of computations at the PiM block. For each idle period 302 of the sequence, the power and energy model 506 computes energy consumption estimates that indicate a respective energy overhead for the sequence of mode transitions 330. In some implementations, and for each idle period 302 of the sequence, the power and energy model 506 calculates energy overhead for the different idle mode types 314. 316, and 318, and for different combinations of idle mode types that may be selected for the different idle states / periods 302 of the mode transitions 330.

[0081] The system 100 determines, from the energy overhead, optimal low power modes for idle durations on a host device (e.g., HPU or tensor processor) and the PiM block to minimize energy overhead for the sequence of mode transitions and set of computations. The system 100 can also compute or otherwise determine a corresponding latency overhead. For example, the system 100 can compute or otherwise determine transition latencies that are required for each mode transition in the sequence of mode transitions and set of computations. The system 100 can encode transition latency values for different mode transitions or types of mode transitions as control values in a model binary / instruction set for executing an ML / NN model on a special-purpose hardware integrated circuit, such as an HPU or neural processor, of the SoC 102.

[0082] In some cases, the performance model 504 infers performance attributes of one or more PiM blocks 202 of the memory device 122, and system 100 uses the performance attributes to predict how a PiM block 202 will perform when executing computations for a ML workload. The system 100 uses the performance model 504 to generate similar inferences and predictions for the host 514. Relatedly, the power and energy model 506 infers the power and energy utilization of the one or more PiM blocks 202. and system 100 uses this utilization data to predict energy output of the PiM blocks 202 when the blocks execute the computations for the ML workload. The system 100 uses the performance model 504 to generate similar inferences and predictions for the host 514.

[0083] The compiler software stack 520 includes an example compiler that generates executable code for a model binary 508 in response to compiling source code for the neuralnetwork ML model. The model binary 508 that is used to execute the neural network ML model, for example, using an application processor of the host device (e.g.. HPU 114). In some implementations, the compiler software stack 520 generates a model file that combines the model binary 508 and metadata indicative of the power mode and transition information 522, 524. For example, the metadata can be generated in response to analyzing data / parameter values of the performance data and corresponding energy consumption estimates.

[0084] The system 100 loads the model binary 508 on the host device 514, which generates instructions for triggering one or more mode transitions in response to executing the model binary 508. In some implementations, the system 100 communicates the model binary 508 using control signals that are transmitted wirelessly to the host 514. In response to executing the model binary 508, the host 514 can generate instructions for triggering one or more mode transitions at the PiM block 202. The host 514 uses the memory controller 105 to pass these mode transition instructions to the PiM block 202.

[0085] For example, the host 514 may use the SoC power management unit 518 to pass these mode transition instructions as pre-programmed instructions or control signals to the PiM block 202. In some implementations, the SoC power management unit 518 passes mode transition instructions along with (or as) memory transaction requests or commands that are generated by the memory controller 105. The PiM block 202 processes the instructions to efficiently transition to, and from, an optimal idle mode when performing ML / neural network computations to execute an ML workload. In some implementations, each of the host 514 and the SoC 102 is configured to send mode transition instructions independent of the memory' controller 105.

[0086] The system 100 uses the compiler software stack 520 to determine power mode and transition information 522, 524 for the host device 514 and PiM block 202, respectively. In some implementations, a power mode controller of the host firmware 510 determines the power mode and transition information 522, 524 in response to executing the model binary 508 at a target processor of the host 514. The power mode and transition information 522 includes target power modes and respective mode transition start times for the host 514, whereas the power mode and transition information 524 includes target power modes and respective mode transition start times for the PiM block 202.

[0087] Fig. 6 shows example graphical data for energy graphs 602-A, 602-B, which indicate energy overhead for different temperatures and operating voltages.

[0088] In the example of Fig. 6, the energy' graph 602-A shows the estimated energy consumption, indicated as energy overhead 604, corresponding to active mode 418 and the multiple idle mode types 314, 316, 318 for a “low” temperature (e.g., < 100° C) and a “low” operating voltage (e.g., < 3.3V) of the SoC 102 or host device, whereas the energy graph 602- B shows the estimated energy consumption corresponding to active mode 418 and the multiple idle mode types 314, 316, 318 for a “high” temperature (e.g., > 100° C) and a “high” operating voltage (e.g., > 3.3V) of the host 514. Other examples for low / high temperature and voltage are also within the scope of this disclosure. The energy' overhead 604 can indicate a level of energy' that corresponds to the mode type for a particular time duration.

[0089] The system 100 includes idle state management logic (e g., of the SoC 102) that is configured to determine a particular idle mode type 314, 316. 318 for an idle mode at the host or the PiM block 202. The idle state management logic makes these determinations from results of energy' overhead calculations that are performed and / or generated by one or more dynamically updated energy models of system 100. In some implementations, the idle state management logic is executed at the host power management unit 516, the SoC power management unit 518. or both. The SoC 102 (or host device) uses the power and energy model 506 to periodically generate and / or update energy overhead estimates for various mode transitions.

[0090] The energy overhead estimates are generated and / or updated in accordance with observed conditions of the SoC 102. the host device, or both. The observed conditions include temperature and voltage, for example, of the host 514 or memory' device 122. The observed conditions can also current, frequency, and / or compute intensity' for operations at the host 514 or memory device 122. In some implementations, the power and energy' model 506 dynamically updates the corresponding energy consumption estimates during execution of the ML model based on the noted system conditions, such as operating voltage and temperature. The energy consumption estimates can also include leakage estimates for each corresponding power mode.

[0091] The SoC 102 (or host device) can include a power mode selector 606 that is configured to adjust the thresholds 610 of idle duration to achieve the best power models that demonstrates lowest minimums for energy' consumption requirements. Prior to entering the execution mode 418, the system 100 can determine the sequence of mode transitions 330 for executing a ML workload using a host device of the SoC 102 and the PiM block 202.

[0092] The sequence of mode transitions can include transitions from processing state 304 to an idle state 302, where the idle state 302 triggers one or more idle modes, such asclock-gating mode 314, retention mode 316, and power-gating mode 318. In particular, for each of the mode transitions, the power mode selector 606 can select an optimal idle mode type that minimizes power consumption and transition latency from idle 302 to processing 304, while also maximizing compute performance ML workload execution.

[0093] For example, as shown in graph 602- A / B, the SoC 102 (or host device) can determine the sequence of mode transitions based on the estimated energy overhead associated with each mode according to system conditions such as low / high temperature and low / high operating voltage. The SoC 102 can determine that, for a given workload type or compute sequence, the host 514 or PiM block will experience a specific idle mode duration (e.g., 2.5 milliseconds (ms)).

[0094] More specifically, the idle state management logic can reference the idle mode duration and analyze any associated energy consumption estimates to determine that clockgating mode 314 is the optimal idle mode to minimize power consumption for that specific idle mode duration of 2.5ms. The pow er mode selector 606 can receive and / or process signal indications about the specific idle mode duration and select the clock-gating mode 314 is the optimal idle mode.

[0095] As indicated above, the power mode selector 606 can adjust thresholds 610 of idle duration to achieve the best power modes / models that have the minimum energy7consumption. In the example of Fig. 6, the thresholds correspond at least to Tl, T2, and T3. For each of energy graphs 602- A / B. if the predicted idle duration is between Tl and T2, then the target power model (or mode) during the corresponding idle state 302 is ‘Clock -gating’ 314. If the predicted idle duration is betw een T2 and T3, then the target pow er model (or mode) during the corresponding idle state 302 is ‘Retention’ 316. Additionally , if the predicted idle duration is longer than T3, then the target power model (or mode) during the corresponding idle state 302 is ‘Power-gating’ 318.

[0096] In some implementations, the SoC 102 can determine that the host 514 is required to transition from execution mode 418 to clock-gating mode relatively quickly to minimize the energy overhead 604 for ML computations that span a sequence of mode transitions 330. The system 100 can also determine to transition from a first idle mode (e.g., clock-gating mode 314) to other relatively deeper types of idle modes (e.g., retention mode 316 and pow er-gating mode 318) based on the energy overhead estimates. Due to the low7operating temperature and the low operating voltage, the system 100 can determine to perform these transitions to other idle modes relatively slowly.

[0097] Fig. 7 is an example process for managing mode transitions for a memory device 122. Process 700 is implemented or executed at system 100 using at least the SoC 102 and memory device 122 described above. Hence, descriptions of process 700 will reference the above-mentioned computing resources of system 100. In some examples, the steps or actions of process 700 are enabled by programmed software instructions, firmware instructions, or both. Each t pe of instruction may be stored in a non-transitory machine-readable storage device and is executable by one or more of the processors or other resources described in this specification.

[0098] Referring again to process 700, the system 100 determines a first sequence of mode transitions, each mode transition in the first sequence being a transition from an execution mode of the memory device 122 to an idle mode of the memory device 122 (702). For example, the SoC 102 can determine the first sequence of mode transitions for the memory device 122 to transition from the execution mode to the idle mode based on execution durations and idle durations derived from a computational graph for a processor of the memory device 122, such as the processor of the PiM block 202.

[0099] Before each occurrence of each mode transition in the first sequence, the system 100 determines, from among multiple idle mode types, a type of idle mode that the memory device 122 will transition to after occurrence of the mode transition (704). For example, the SoC 102 can determine a type of idle mode for the memory device 122 to transition to after transitioning from the execution mode to the idle mode. The type of idle mode can be a clock -gating idle mode, a retention idle, or a power-gating idle mode.

[0100] The system 100 iteratively generates a first set of control signals based on the corresponding type of idle mode for each mode transition in the first sequence (706). For example, the SoC 102 can iteratively generate and transmit control signals to the memory device 122 using a control channel that couples the memoiy device 122 and the SoC 102. In some examples, the SoC 102 can iteratively generate and transmit the control signals to the memoiy device 122 over a time duration that coincides with a duration for executing a ML workload.

[0101] Based on the first set of control signals, the system 100 iteratively triggers occurrence of each mode transition in the first sequence (708). In some implementations, each triggered occurrence causes the memory device 122 to transition to a corresponding idle mode type. For example, each control signal can be used to trigger the occurrence of each mode transition in the first sequence of mode transitions. In some implementations, the SoC 102 or host device triggers each mode transition at the PiM block 202 via control signals thatare passed to the memory' device 122, e.g., via the memory controller 105 of the SoC 102, to cause the PiM block 202 to transition to the idle mode to restrict power consumption at the memory device 122.

[0102] Fig. 8 is another example process 800 for managing mode transitions for a memory7device. Like process 700, process 800 is also implemented or executed at system 100 using at least the SoC 102 and memory device 122 described above. In some examples, the steps or actions of process 7600 are enabled by programmed software instructions, firmware instructions, or both. Each type of instruction may be stored in a non-transitory machine-readable storage device and is executable by one or more of the processors or other resources described in this specification.

[0103] Process 800 includes predicting each time duration of operation processing 304 and idle 302 in a host processor and PiM block based on operation parameters (802). The system 100 estimates energy' overhead of mode transitions and leakage for each idle mode ty pe 314, 316, 318 according to SoC conditions, PiM block conditions, or both (804). The system 100 determines the target idle mode for minimal energy overhead (806). For example, the SoC 102 can determine the target idle mode from among the idle mode types of clock-gating mode 314, retention mode 316, and power-gating mode 318.

[0104] The estimated energy' overhead of mode transitions for a particular idle mode may include a first transition energy consumption of a transition from an execution mode to the particular idle mode, and a second transition energy consumption of a transition from the particular idle mode to a subsequent execution mode. The leakage for each idle mode type may be considered as an estimated energy' consumption of an idle mode duration corresponding to a length of time in the respective idle mode after transitioning to the respective idle mode from the execution mode. The energy consumptions described above may be estimated by the power & energy model 506 of the SoC 102. An idle mode which has the minimal energy' overhead (e.g., a sum of the first transition energy consumption, the estimated energy consumption of the idle mode duration, and the second transition energy consumption being the lowest) may be determined by the SoC 102 as the target idle mode.

[0105] The system 100 schedules a wake-up (exit) transition from an idle mode to minimize the mode transition latency to the next processing operation (808). For example, the SoC 102 is configured to determine, select, and schedule a respective wake-up (exit) for transitioning from each target idle mode. The SoC 102 selects the wake-up to minimize a corresponding mode transition latency when the host device or the PiM block 202 transitions from an idle mode to the next processing operation.

[0106] The respective steps of process 700 and 800 can be performed at a hardware integrated circuit as part of a larger compute operation to generate a machine-learning output, including an output for a neural network layer of a neural network that implements one or more ML models. For example, the output can be a portion of a computation for a ML task or inference workload to generate an image processing, speech processing, or image recognition output.

[0107] As indicated above, a portion of the integrated circuit can include a specialpurpose neural network processor or hardware ML accelerator configured to accelerate computations for generating different ty pes of data processing outputs. In some implementations, one or more of the PiM operations, CiM operations, or MnC operations are performed by the memory device 122 to support or enable accelerating computations for generating different types of data processing outputs.

[0108] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory program carrier for execution by, or to control the operation of. data processing apparatus.

[0109] Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0110] The term “computing system” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0111] A computer program (which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0112] A computer program may. but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0113] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array), an ASIC (application specific integrated circuit), or a GPGPU (General purpose graphics processing unit).

[0114] Computers suitable for the execution of a computer program include, by way of example, can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. Some elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0115] Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory7devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0116] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.

[0117] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.

[0118] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0119] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can alsobe implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0120] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0121] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

What is claimed is:

1. A method implemented using an integrated circuit comprising a System-on-Chip (“SoC") and a memory device coupled to the SoC, the method comprising: determining a first sequence of mode transitions, each mode transition in the first sequence being a transition from an execution mode of the memory device to an idle mode of the memory device; before occurrence of each mode transition in the first sequence: determining, from among a plurality of idle mode types, a type of idle mode that the memory device will transition to after occurrence of the mode transition; iteratively generating a first set of control signals based on the corresponding type of idle mode for each mode transition in the first sequence; and based on the first set of control signals, iteratively triggering occurrence of each mode transition in the first sequence, wherein each triggered occurrence causes the memory device to transition to a corresponding type of idle mode.

2. The method of claim 1, further comprising: determining a second sequence of mode transitions, each mode transition in the second sequence being a transition from an idle mode of the memory device to an execution mode of the memory device; and for each type of idle mode of the memory device: determining a respective transition latency for transitioning to an execution mode of the memory device from the corresponding type of idle mode of the memory device.

3. The method of claim 2, further comprising: iteratively generating a second set of control signals based on the corresponding transition latency for each mode transition in the second sequence; and based on the second set of control signals, iteratively triggering occurrence of each mode transition in the second sequence.

4. The method of claim 3, wherein each triggered occurrence of a mode transition in the second sequence causes the memory device to begin executing computations immediatelyafter a period of transition latency that corresponds to the type of idle mode from which the memory device is transitioning.

5. The method of any preceding claim, wherein generating the first or second set of control signals comprises: iteratively generating control signals using a memory controller of the SoC.

6. The method of any preceding claim, wherein iteratively triggering occurrence of each mode transition in the first sequence comprises: transitioning to an idle mode of the memory device to restrict power consumption at the memory device by clock gating or power gating a Processing-in-Memory C‘PiM’?) block of the memory device.

7. The method of claim 6 as dependent from claim 3, wherein iteratively triggering occurrence of each mode transition in the second sequence comprises: transitioning to an execution mode of the memory device to execute computations at the memory device using the PiM block of the memory device.

8. The method of claim 6 or 7, wherein determining a first sequence of mode transitions comprises, determining a first sequence of mode transitions based on: execution durations and idle durations derived from a computational graph for a processor of the PiM block.

9. The method of any preceding claim, wherein the plurality of idle mode types comprises one or more of: i) a retention idle mode; ii) a clock-gating idle mode; and iii) a power-gating idle mode.

10. The method of any preceding claim, wherein iteratively generating the first set of control signals comprises: iteratively generating the first set of control signals over a time duration that coincides with a duration for executing a machine-learning (“ML’?) workload using the SoC and / or a PiM block of the memory device.

11. The method of any preceding claim, wherein each control signal in the first set of control signals is used to trigger occurrence of each mode transition in the first sequence of mode transitions.

12. The method of any preceding claim, wherein during the execution mode, a Processing-in-Memory (“PiM”) block of the memory device executes computations for a machine-learning ("MI.") workload, and the MT workload is executed using both the SoC and the PiM block.

13. The method of any preceding claim, wherein determining the type of idle mode comprises: determining the type of idle mode based on energy overhead calculations generated by one or more dynamically updated energy models of the SoC.

14. The method of claim 13. wherein the one or more dynamically updated energy models comprise a power & energy model, the method further comprising: estimating, by the power & energy' model of the SoC, a first transition energy' consumption of a transition from an execution mode to an idle mode; estimating, by the power & energy model of the SoC, an energy consumption of an idle mode duration corresponding to a length of time in the idle mode after transitioning to the idle mode from the execution mode; and estimating, by the power & energy model of the SoC, a second transition energy consumption of a transition from the idle mode to a subsequent execution mode.

15. The method of claim 13 or 14, wherein the energy' model is dynamically updated: i) during an active state of the SoC, while the SoC is running or executing an ML workload; and ii) based on system conditions comprising an operating voltage of the SoC and a temperature measured at the SoC or at the PiM block of the memory device.

16. An integrated circuit comprising: a System-on-Chip (“SoC’); a memory device coupled to the SoC; anda processor and a non-transitory machine-readable storage medium of the SoC for storing instructions that are executable by the processor to cause performance of operations comprising: determining a first sequence of mode transitions, each mode transition in the first sequence being a transition from an execution mode of the memory device to an idle mode of the memory device; before occurrence of each mode transition in the first sequence: determining, from among a plurality of idle mode ty pes, a type of idle mode that the memory device will transition to after occurrence of the mode transition; iteratively generating a first set of control signals based on the corresponding type of idle mode for each mode transition in the first sequence: and based on the first set of control signals, iteratively triggering occurrence of each mode transition in the first sequence, where each triggered occurrence causes the memory' device to transition to a corresponding type of idle mode.

17. A method implemented using an integrated circuit comprising a System-on-Chip ’SoC”) and a memory device coupled to the SoC, the method comprising: identifying a mapping of respective portions of computations between the SoC and the memory device, wherein the memory device is external to the SoC; determining, based on the mapping, a plurality of transition points that define a sequence of transitions between each portion of the computations mapped to the SoC and each portion of the computations mapped to the memory device; for each transition point of the plurality of transition points and before occurrence of a transition defined by the transition point: i) determining a mode the SoC will transition to after occurrence of the transition; ii) determining a mode the memory device will transition to after occurrence of the transition; and iii) selecting a particular type of idle mode that restricts power consumption at the memory device when the memory device will transition to an idle mode after occurrence of the transition.