Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

35 results about "Memory hierarchy" patented technology

In computer architecture, the memory hierarchy separates computer storage into a hierarchy based on response time. Since response time, complexity, and capacity are related, the levels may also be distinguished by their performance and controlling technologies. Memory hierarchy affects performance in computer architectural design, algorithm predictions, and lower level programming constructs involving locality of reference.

Cache writeback circuit

A cache writeback circuit is disclosed for writing back, without invalidating, dirty cache lines. The cache writeback circuit is configured to enter an active state based on detecting a trigger condition indicative of cache misses to a memory cache circuit within a memory hierarchy of a computer memory subsystem causing cache line eviction activity. During the active state, the cache writeback circuit is configured to identify a set of dirty cache lines in the memory cache circuit, and write back, without invalidating, cache lines of the identified set of dirty cache lines from the memory cache circuit to a memory circuit within a lower level of the memory hierarchy, such as DRAM. The cache writeback circuit may further be configured to identify the dirty cache lines via a cache walk operation, which can be suspended, for example, when a higher-priority cache operation occurs.
Owner:APPLE INC

Non-blocking vector instruction dispatch with micro-operations

A processor core is coupled to a memory hierarchy. The processor core is configured to execute vector instructions, scalar instructions, and micro-operations. A dispatch unit within the processor core receives a vector memory operation. The dispatch unit sends the vector memory operation to a first vector input queue of multiple vector input queues. The sending is based on the memory addressing mode. A micro-operation sequencer splits the vector memory operation into one or more memory micro-operations, which includes forwarding each micro-operation within the one or more micro-operations to a first memory queue within multiple memory queues. A memory operation is then issued to a load-store unit within the processor core. The issuing includes selecting, from the multiple memory queues, the memory operation. The vector memory operation comprises either a vector load operation or a vector store operation.
Owner:AKEANA INC

Reducing DC power consumption using dynamic memory prefetching for processing engines

Aspects of the disclosure are directed to dynamic memory prefetching. In accordance with one aspect, the disclosure includes a main memory configured to store data; a memory hierarchy coupled to the main memory, the memory hierarchy configured to augment the main memory; and a trusted zone element coupled to the memory hierarchy, the trusted zone element configured to read a shared memory to determine a specific control register and one or more prefetch values based on a secure monitor call (SMC). In another aspect, the disclosure includes reading a shared memory to determine a specific control register and one or more prefetch values based on a secure monitor call (SMC); and writing the one or more prefetch values to an exception level register.
Owner:QUALCOMM INC

Hardware-optimized recurrent neural network system

A system includes a machine-learning model implemented on a data processing apparatus, which features a parallel processor with a memory hierarchy. The machine-learning model is a recurrent neural network (RNN) with a multi-head architecture, comprising multiple sub- vectors that process parallel data streams. The RNN's weight matrix is structured as a block-diagonal matrix, allowing for parallel processing of the sub-vectors. A fused computational kernel executes an entire time-series processing loop for the multi-head RNN, maintaining the weight matrix blocks in on-chip memory and performing matrix multiplications and element-wise operations for each sub-vector in a single kernel execution.
Owner:NXAI GMBH

Dynamic ai model selection and pre-loading based on data temperature scoring and next prompt prediction

PendingUS20260178343A1Resource allocationWeb data indexingModel managementMemory hierarchy
A system and method are provided for artificial intelligence (AI) model selection and loading. The method includes receiving an input data stream for an AI application and analyzing the stream to determine temperature scores based on topic frequency, importance weightage, and access frequency. The method also includes mapping these temperature scores to relevant expert AI models and predicting the temperature and context of the next prompt using historical data patterns through a sliding window that smooths temporary anomalies. The system dynamically selects AI models based on temperature scores and predicted contexts, then pre-loads these models into a memory hierarchy where higher-temperature models are placed in faster memory. The method processes incoming prompts using these pre-loaded models and outputs responses, improving processing efficiency through predictive model management. The system addresses challenges of managing large-scale AI models by implementing intelligent pre-loading mechanisms that anticipate and prepare for upcoming processing needs.
Owner:SK HYNIX NAND PRODUCT SOLUTIONS CORP

Compression Metadata Value Induced Read and Write Operation Conservation

PendingUS20260088831A1Code conversionMemory hierarchyData compression
In disclosed embodiments, compression circuitry may compress data blocks and compressed write data blocks and corresponding compression metadata to a certain level in a cache / memory hierarchy. In some embodiments, the compression circuitry is configured to detect a pre-determined set of data values in a block of data to be compressed. In response, the compression circuitry may write a special metadata value for the data block to data storage circuitry to indicate compression of the data block, without writing a compressed version of the data block to the data storage circuitry. This may advantageously improve performance and reduce power consumption for data blocks with certain data values (e.g., having uniform pixel / texel values).
Owner:APPLE INC

Vector Scatter and Gather with Single Memory Access

PendingUS20260119176A1Register arrangementsMemory hierarchyComputer architecture
Disclosed embodiments provide techniques for improved performance in processing vector instructions. A processor core is accessed. The processor core is coupled to a memory hierarchy, and the processor core includes one or more vector execution units (VUs), and one or more load store units (LSUs). The processor core includes a vector register file (VRF). The VRF includes multiple vector registers, and each vector register includes multiple vector elements. Vector elements that have a source or destination in contiguous memory are identified. Load store units (LSUs) take advantage of the contiguous memory condition by executing a vector load or vector store operation as a single memory access, requiring a reduced number of clock cycles. The single memory access satisfies each memory operation for each vector element within the vector register file.
Owner:AKEANA INC

Dirty tracking bit compression

PendingUS20260079840A1Memory systemsMemory hierarchyDirty data
A cache controller of a cache assigns a dirty tracking bit for each dirty byte of a cache line. Once a predetermined interval has elapsed without any accesses to the cache line or to a cache set that includes the cache line, the cache controller compresses contiguous dirty tracking bits for each portion of the cache line. Compressing the dirty tracking bits for contiguous dirty portions of the cache line allows the cache to store more dirty data using fewer dirty tracking bits, reducing area cost and bandwidth among levels of a memory hierarchy.
Owner:ADVANCED MICRO DEVICES INC

SYSTEM AND METHOD FOR ADAPTIVE HYBRID Hardware PREFETCHING

An apparatus includes a processor core and a memory hierarchy. The memory hierarchy includes a main memory and one or more caches between the main memory and the processor core. A plurality of hardware prefetchers are coupled to the memory hierarchy, and a prefetch control circuit is coupled to the plurality of hardware prefetchers. The prefetch control circuit is configured to compare changes in the one or more cache performance metrics within the two or more sampling intervals and control operation of the plurality of hardware prefetchers in response to changes in the one or more performance metrics between at least a first sampling interval and a second sampling interval.
Owner:HUAWEI TECH CO LTD

Non-blocking unit stride vector instruction dispatch with micro-operations

Disclosed techniques enable vector instruction processing. A processor core is accessed. The processor core is coupled to a memory hierarchy, and is configured to execute vector operations, scalar operations, and micro-operations. A decode unit decodes a vector memory operation. The vector memory operation is associated with a unit stride addressing mode. The decoding includes dividing the vector memory operation into one or more vector memory micro-operations. A dispatch unit sends at least one vector micro-operation within the one or more vector micro-operations to a scalar request queue within a plurality of request queues. The at least one vector micro-operation is issued to a load-store unit within the processor core. The issuing includes selecting, from the plurality of request queues, the at least one vector memory micro-operation.
Owner:AKEANA INC

Method and system for one-dimensional signal extraction for various computer processors

Methods and systems for extracting a one-dimensional (1D) signal from a two-dimensional (2D) digital image along a projection line are provided herein. The method and system store the digital image in a memory hierarchy, where a non-blocking prefetch operation may fetch pixels from a main memory to a data cache. In response to the direction of the projection line, a prefetch plan, a pixel processing plan, and a prefetch distance are selected. The prefetch plan uses a first address order designed to facilitate efficient extraction of pixels from the main memory to the data cache for a given direction. The pixel processing plan uses a second address order designed to facilitate calculation of a one-dimensional signal along the projection line. A pixel processing plan is used in coordination with a prefetch plan to compute the one-dimensional signal such that the pixel is fetched from the main memory to the data cache in advance in response to an amount of time of the prefetch distance before the pixel operation uses the pixel.
Owner:COGNEX CORP

Dynamic ai model selection and pre-loading based on data temperature scoring and next prompt prediction

PCT designated stageWO2026143127A1Memory hierarchyModel management
A system and method are provided for artificial intelligence (Al) model selection and loading. The method includes receiving an input data stream for an Al application and analyzing the stream to determine temperature scores based on topic frequency, importance weightage, and access frequency. The method also includes mapping these temperature scores to relevant expert Al models and predicting the temperature and context of the next prompt using historical data patterns through a sliding window that smooths temporary anomalies. The system dynamically selects Al models based on temperature scores and predicted contexts, then pre-loads these models into a memory hierarchy where higher-temperature models are placed in faster memory. The method processes incoming prompts using these pre-loaded models and outputs responses, improving processing efficiency through predictive model management. The system addresses challenges of managing large-scale Al models by implementing intelligent pre-loading mechanisms that anticipate and prepare for upcoming processing needs.
Owner:SK HYNIX NAND PRODUCT SOLUTIONS CORP

Intelligent agent interaction memory cooperation method and device based on memory grading and medium

PendingCN121765041AAchieve efficient reuseImplement persistence managementDatabase distribution/replicationExecution paradigmsPersonalizationEngineering
The invention discloses an agent interaction memory cooperation method and device based on memory grading and a medium, and the method comprises the steps: obtaining the original interaction data of a user, and carrying out the context interaction analysis of the original interaction data, so as to obtain an interaction state variable; based on the interaction state variable, determining user interaction multiplexing state data through cross-session multiplexing of the user unique identifier; obtaining an iteratable shared knowledge base through high-frequency multiplexing knowledge precipitation according to the user interaction multiplexing state data; performing data interaction control on the interaction state variable, the user interaction multiplexing state data and the iteratable shared knowledge base to determine a directional calling state of the memory data; and based on the directional calling state of the memory data, through multi-level cache matching of agent demand analysis, obtaining agent interactive memory collaborative storage data. Through the method, the technical problem that an artificial intelligence interaction memory function cannot meet real-time interaction dynamic and personalized requirements in the prior art is solved.
Owner:浪潮智慧科技有限公司 +2

Cache control to preserve register data

Techniques are disclosed relating to eviction control for cache lines that store register data. In some embodiments, memory hierarchy circuitry is configured to provide memory backing for register operand data in one or more cache circuits. Lock circuitry may control a first set of lock indicators for a set of registers for a first thread, including to assert one or more lock indicators for registers that are indicated, by decode circuitry, as being utilized by decoded instructions of the first thread. The lock circuitry may preserve register operand data in the one or more cache circuits, including to prevent eviction of a given cache line from a cache circuit based on an asserted lock indicator. The lock circuitry may clear the first set of lock indicators in response to a reset event. Disclosed techniques may advantageously retain relevant register information in the cache with limited control circuit area.
Owner:APPLE INC

CACHE STORAGE ACCESS

ActiveDE112017001959B4Data processing systemMemory hierarchy
A method of data processing in a multiprocessor data processing system (200) comprising several vertical cache memory hierarchies supporting a plurality of processor cores (102a, 102b), a system memory (132), and a system connection connected to the system memory (132) and the several vertical cache memory hierarchies, wherein the method comprises: in response to receiving a request, loading and reserving from a first processor core (102a, 102b), outputting through a first cache memory in a first vertical cache memory hierarchy supporting the first processor core (102a, 102b), on a system connection, a memory access request for a target cache memory row of the request to load and reserve;In response to the memory access request and prior to receiving a system-wide coherence response for the memory access request, the first cache memory receives the target cache memory row from a second cache memory in a second vertical cache memory hierarchy through cache-to-cache intervention and an early specification of the system-wide coherence response for the memory access request; and in response to the early specification of the system-wide coherence response and prior to receiving the system-wide coherence response, the first cache memory initiates processing to update the target cache memory row in the first cache memory.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Weight-stationary matrix multiply accelerator with tightly coupled l2 cache

PendingUS20260037599A1Complex mathematical operationsComputational scienceMemory hierarchy
An accelerator is accessed. The accelerator includes a weight-stationary systolic array of one or more multiply-accumulate units. The accelerator is coupled to a memory hierarchy and a processor core. The processor core sends a work request to the accelerator. The work request is based on execution of a machine learning model and an activation matrix. In response to the work request, the accelerator loads a weight matrix and the activation matrix. The loading uses the memory hierarchy. The accelerator multiplies the weight matrix by the activation matrix. The multiplication results in an answer matrix. The accelerator stores the answer matrix in the memory hierarchy. The processor core obtains the answer matrix that was stored. The machine learning model is trained. The training produces the weight matrix, which is transposed and saved to the memory hierarchy.
Owner:AKEANA INC

Atomic polymerization

Techniques related to polymeric atomic manipulation are disclosed. In some embodiments, a cache control circuit caches data values in a cache storage circuit, and receives a plurality of requests to atomically update the cached data values according to one or more arithmetic operations. The control circuit may, in response to determining that the one or more arithmetic operations meet the one or more criteria, perform updates to the cached data values based on the plurality of requests, and store operational information indicating a recently requested atomic arithmetic operation for the updated data values. The control circuit may, in response to an event, refresh both: the updated data value, and the operational information, to a higher level in the memory hierarchy that includes the cache storage circuit. This may advantageously aggregate atomic operations at the cache and reduce operations on higher level caches or memories, which may be the actual point of coherency of atomic requests.
Owner:APPLE INC

Cache control for storing registry data

Techniques for swapping out cache rows that store register data are disclosed. In some embodiments, the memory hierarchy switching logic is configured to provide memory support for register operand data in one or more cache circuits. Lock switching logic can control a first set of lock indicators for a set of registers for a first thread, including acknowledging one or more lock indicators for registers marked by decode switching logic as being used by decoded instructions of the first thread. The lock switching logic can retain register operand data in one or more cache circuits, including preventing the swapping out of a given cache row from a cache circuit based on an enabled lock indicator. The lock switching logic can clear the first set of lock indicators in response to a reset event.The disclosed techniques can advantageously retain relevant register information in the cache with restricted control.
Owner:APPLE INC

Vector floating-point flag update with micro-operations

A processor core is coupled to a memory hierarchy. The processor core is configured to execute vector floating-point instructions and micro-operations. A vector floating-point instruction is decoded. The decoding includes replacing the vector floating-point instruction with one or more vector floating-point micro-operations (VFPMs). A reorder buffer assigns a reorder buffer ID (ROBID) to each of the one or more VFPMs, in which the assigning includes a micro-sequencer ID (MSID). The processor core executes the one or more VFPMs. The executing includes requiring, by a first VFPM within the one or more VFPMs, a first update to an architectural floating-point flag. The architectural floating-point flag is set, based on the first update. The setting occurs after the one or more VFPMs have been committed by the processor core. A temporary floating-point flag is revised. The revising is based on the first update.
Owner:AKEANA INC

High-bandwidth computing architecture based on storage hierarchy reconstruction and data processing method thereof

The application discloses a high-bandwidth, large-capacity and high-reliability computing architecture based on storage hierarchy reconstruction, a data processing method and system. It belongs to the field of computer architecture, high-performance computing (HPC) and artificial intelligence hardware acceleration. The architecture adopts an innovative design of "small node, large cluster", and is composed of a large cluster which can be infinitely expanded by standardized storage and computing integrated small nodes. The core innovation lies in: reconstructing the storage hierarchy, using SSD as main memory, and directly transmitting data between SSD and cache in the processor chip, eliminating multi-level data copying; and realizing strict matching of full-link bandwidth, so that GPU computing power and Ethernet bandwidth grow synchronously, ensuring that the bandwidth is aligned within the node, the chassis, the rack and across the data center. The application can make the computing power utilization rate and bandwidth utilization rate reach 100%; reduce the system construction cost to less than 1 / 10 of the traditional scheme, reduce the power consumption to 1 / 9, increase the single-node main memory capacity by 20 times to 340,000 times, and significantly shorten the large model training period.
Owner:林沧

Tensor slicing implementation method on tpu and tpu

ActiveCN122332132BMemory hierarchyAlgorithm
The application provides a tensor slicing implementation method on TPU and TPU, and relates to the technical field of TPU. The method comprises the following steps: obtaining slicing parameters corresponding to a to-be-sliced tensor; determining a target execution branch from a plurality of preset execution branches according to the slicing parameters; wherein for each preset execution branch, at least one of a storage level, a data transmission mode and a data processing mode adopted by the branch is different from those of other branches; and performing a slicing operation on the to-be-sliced tensor on a TPU chip according to the target execution branch to obtain a target tensor. The application can select a corresponding tensor slicing processing flow according to the slicing parameters to adapt to the hardware characteristics of the TPU and give full play to the advantages of the TPU architecture.
Owner:ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD

Event-driven asynchronous graph neural network FPGA accelerator for real-time edge vision

The invention belongs to the technical field of artificial intelligence chips and reconfigurable computing, and particularly relates to an event-driven asynchronous graph neural network FPGA accelerator for real-time edge vision. Aiming at the problems of low storage utilization rate, obvious data access bottleneck, limited degree of parallelism, too high calculation redundancy and the like of a traditional GNN accelerator in an event vision scene, the invention provides a novel architecture combining efficient graph feature storage, hierarchical graph construction and redundancy elimination convolution calculation. By introducing parallel read-write optimization, low-dependence graph structure generation and a reusable computing cache mechanism, the computing delay is remarkably reduced while resource consumption is kept controllable, the overall throughput rate is increased, and therefore the sub-microsecond real-time reasoning capacity is achieved. The method is suitable for unmanned driving, intelligent monitoring, robot navigation and other application scenes needing low-delay visual calculation on edge equipment.
Owner:SHANGHAI TECH UNIV

A hardware-software co-optimization method for a hybrid in-memory architecture

The application discloses a kind of hardware and software collaborative optimization methods of hybrid in-memory architecture, comprising the following steps: the joint feature representation of target AI algorithm and in-memory computing architecture is carried out, context features are extracted, and the in-memory computing unit configuration of in-memory computing architecture, neural network processor pipeline, multi-core interconnection topology and storage hierarchy interface are parameterized;Offline benchmark dataset is constructed and predictive agent model is trained, the model uses embedding method to process discrete architecture parameters, uses encoder with self-attention mechanism to learn feature association, and realizes the joint prediction of multidimensional PPA index through parallel prediction network, and introduces feasible region constraint learning mechanism in training;The trained agent model is embedded in multi-objective evolutionary algorithm, with energy efficiency, average computing power utilization, model execution delay as optimization target, while meeting chip area efficiency and power consumption constraints, searching for Pareto optimal in-memory computing architecture configuration set.
Owner:SOUTH CHINA UNIV OF TECH

Pre-installing page table cache rows on a virtual machine

Procedures, exhibiting: Running an initial operating system (OS), corresponding to an initial virtual machine (VM), using a processor for an initial predefined period; Pre-installing page table entries before the first predefined period expires, corresponding to a second OS of a second VM, by moving the page table entries from a lower level to an upper level of a memory hierarchy, the upper level being closer to the processor in the memory hierarchy than the lower level; and After the first predefined period has expired, the second OS, corresponding to the second VM, is executed using the processor for a second predefined period.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Variable dispatch walk for successive cache accesses

A processing system is configured to translate a first cache access pattern of a dispatch of work items to a cache access pattern that facilitates consumption of data stored at a cache of a parallel processing unit by a subsequent access before the data is evicted to a more remote level of the memory hierarchy. For consecutive cache accesses having read-after-read data locality, in some embodiments the processing system translates the first cache access pattern to a space-filling curve. In some embodiments, for consecutive accesses having read-after-write data locality, the processing system translates a first typewriter cache access pattern that proceeds in ascending order for a first access to a reverse typewriter cache access pattern that proceeds in descending order for a subsequent cache access. By translating the cache access pattern based on data locality, the processing system increases the hit rate of the cache.
Owner:ATI TECHNOLOGIES ULC +1

Technique for handling prefetching using trigger circuitry

ActiveUS12541457B2Memory systemsComputer hardwareMemory hierarchy
An apparatus has cache circuitry providing a cache storage to store data for access by processing circuitry, and request handling circuitry arranged to process requests, each request providing an address indication for associated data. The request handling circuitry determines with reference to the address indication whether the associated data is available in the cache circuitry. The cache circuitry forms a given level of a multi-level memory hierarchy, and the request handling circuitry is responsive to determining that the associated data is unavailable in the cache circuitry to issue an onward request to cause the associated data to be retrieved into the cache circuitry from a lower level of the multi-level memory hierarchy than the given level. Prefetch circuitry issues, as one type of request to be handled by the request handling circuitry, prefetch requests, and the request handling circuitry is arranged in response to a given prefetch request to retrieve into the cache circuitry the associated data in anticipation of that associated data being requested by the processing circuitry. In addition, trigger circuitry, responsive to a specified condition being detected in respect of the given prefetch request, issues a prefetch trigger signal for receipt by control circuitry associated with further cache circuitry at a higher level of the multi-level memory hierarchy, to cause a higher level prefetch procedure to be triggered by the control circuitry to retrieve the associated data from the cache circuitry into the further cache circuitry.
Owner:ARM LTD

A method for accelerating Scloud+ algorithm based on multi-thread matrix multiplication

A method for accelerating Scloud+ algorithm based on multi-thread matrix multiplication. S1, a coarse-grained parallel architecture is constructed; S2, the AES-128 encryption module of the matrix A generation link in Scloud+.KEM is optimized by GPU native; S3, a local variable storage reuse strategy is adopted, and the local matrix variables with non-overlapping life cycle are stored in the storage space; S4, a pre-computation and persistence mechanism is introduced, and the pre-computation of the index data required for the generation of the matrix A is completed in the program initialization stage; S5, instruction level deep optimization; S6, a warp level fine-grained synchronization mechanism is constructed; S7, a double-thread cooperation fine-grained parallel implementation based on the coarse-grained architecture. Through the multi-dimensional collaborative innovation of architecture, memory hierarchy and instruction set, the computing, storage and communication bottlenecks of the post-quantum cryptography algorithm Scloud+ on the general parallel computing architecture are systematically solved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Tensor slicing implementation method on tpu and tpu

PendingCN122332132AMemory hierarchyAlgorithm
This invention provides a method for implementing tensor slicing on a TPU and a TPU in the field of TPU technology. The method includes: obtaining slicing parameters corresponding to the tensor to be sliced; determining a target execution branch from multiple preset execution branches based on the slicing parameters; wherein, for each preset execution branch, at least one of the following—storage hierarchy, data transmission mode, and data processing method—is different from other branches; and performing a slicing operation on the tensor to be sliced ​​on the TPU chip according to the target execution branch to obtain the target tensor. This invention can select the corresponding tensor slicing processing flow based on the slicing parameters to adapt to the hardware characteristics of the TPU and fully leverage the advantages of the TPU architecture.
Owner:ZHONGHAO XINYING (HANGZHOU) TECH CO LTD

Large language model kernel generation method and system and text generation method

The invention provides a large language model kernel generation method and system and a text generation method, and the method comprises the steps: firstly obtaining a hardware constraint parameter, and constructing a resource constraint perception rule comprising a degree of parallelism and a cyclic partitioning rule according to the hardware constraint parameter; then, space exploration is performed based on the rules and a sequence decision algorithm to generate a target kernel. The sequence decision algorithm sequentially traverses according to the memory hierarchy and the cycle dimension, the parallelism degree of each dimension and the feasible interval of the block size are determined, memory hierarchy configuration parameters are obtained step by step and finally summarized into target kernel configuration parameters, and a target large language model kernel is generated according to the target kernel configuration parameters. According to the large language model kernel generation method provided by the invention, the resource constraint perception rule is dynamically constructed based on the hardware constraint parameter of the target large language model, so that the high-performance target large language model kernel is quickly generated, the kernel generation time is shortened, and the kernel generation efficiency is improved. And the calculation efficiency and the resource utilization rate of the large language model on special hardware are improved.
Owner:SUN YAT SEN UNIV

Memory hierarchy power management

Some embodiments include a system, apparatus, method, and computer program product for memory hierarchy power management. Some embodiments include a performance controller that balances memory hierarchy power and compute power to maintain package-level power efficiency of a systems-on-a-chip (SoC)-memory package. The performance controller can determine a ratio of memory hierarchy power to compute agent power, compare the ratio against a threshold value, and based on the comparison, determine how to manage memory hierarchy power. When the energy costs of the memory hierarchy power are large relative to the energy costs of the compute agent power, some embodiments include changing a performance state of a fabric and / or memory to increase the power efficiency of the overall SoC-memory package, even though a number of memory stall cycles experienced by the compute agent may increase.
Owner:APPLE INC