Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

60 results about "Memory hierarchy" patented technology

In computer architecture, the memory hierarchy separates computer storage into a hierarchy based on response time. Since response time, complexity, and capacity are related, the levels may also be distinguished by their performance and controlling technologies. Memory hierarchy affects performance in computer architectural design, algorithm predictions, and lower level programming constructs involving locality of reference.

Distributed time sequence library data management method supporting cold and hot data level-to-level management

The invention discloses a distributed time sequence library data management method supporting cold and hot data level-to-level management, and relates to the technical field of computer databases, comprising: receiving a write-in request of time sequence data, and executing preliminary write-in data aggregation and sorting; dynamically identifying data cold and hot attributes based on a multi-dimensional data cold and hot degree calculation model in combination with a write-in behavior and an access behavior of time series data; according to the cold and hot attribute recognition result, in combination with the hierarchy boundary of self-adaptive division, automatic hierarchical storage of the time series data is executed; based on a cold and hot data dynamic migration and scheduling mechanism, according to the access frequency and time decay characteristics of the time series data, dynamically adjusting the storage hierarchy of the time series data, and executing a data migration task; and constructing a hierarchical index system adaptive to different cold and hot attribute time series data, and combining query frequency dynamic identification and index structure automatic upgrading. By adopting a cold and hot data automatic identification and hierarchical storage mechanism, the overall performance and the resource utilization rate of the system are improved.
Owner:GUODIAN NANJING AUTOMATION

Method, device and equipment for optimizing reasoning operator of large model based on mercuric chloride chip

The technical scheme can be applied to the field of financial science and technology / medical health. The invention discloses a mercuric chloride chip-based large model reasoning operator optimization method, device and equipment, and the method comprises the steps: carrying out the blocking processing of original input data according to the parallel calculation capability of a mercuric chloride chip, and generating the blocking data of an adaptive chip calculation unit; optimizing a memory access path of a matrix multiplication operator by combining a memory hierarchical structure and a calculation core type of a mercuric chloride chip based on the block data, and adjusting a sliding step length and a filling mode of convolution operation; a plurality of operators continuously executed in the large model are fused into a composite operator, the data storage process of the composite operator is optimized, and collaborative execution is achieved by dynamically allocating computing resources; and integrating and decoding block calculation results after collaborative optimization, and adjusting an input data block strategy and operator execution parameters through a verification feedback mechanism to form closed-loop optimization. According to the technical scheme, the reasoning efficiency of a large model can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Cache writeback circuit

A cache writeback circuit is disclosed for writing back, without invalidating, dirty cache lines. The cache writeback circuit is configured to enter an active state based on detecting a trigger condition indicative of cache misses to a memory cache circuit within a memory hierarchy of a computer memory subsystem causing cache line eviction activity. During the active state, the cache writeback circuit is configured to identify a set of dirty cache lines in the memory cache circuit, and write back, without invalidating, cache lines of the identified set of dirty cache lines from the memory cache circuit to a memory circuit within a lower level of the memory hierarchy, such as DRAM. The cache writeback circuit may further be configured to identify the dirty cache lines via a cache walk operation, which can be suspended, for example, when a higher-priority cache operation occurs.
Owner:APPLE INC

Software Runtime Assisted Co-Processing Acceleration with a Memory Hierarchy Augmented with Compute Elements

The concepts and technologies disclosed herein are directed to software runtime assisted co-processing acceleration with a memory hierarchy augmented with compute elements. An example system disclosed herein includes one or more switches and a plurality of hardware compute nodes connected via the one or more switches. Each hardware compute node of the plurality of hardware compute nodes includes an in-memory compute (IMC) element configured to perform in-memory processing operations on data, such as graph data. The system also includes a near-memory compute (NMC) element configured to perform near-memory processing operations on the data. The system also includes a far-memory compute (FMC) element configured to perform far-memory processing operations on the data.
Owner:ADVANCED MICRO DEVICES INC

Non-blocking vector instruction dispatch with micro-operations

A processor core is coupled to a memory hierarchy. The processor core is configured to execute vector instructions, scalar instructions, and micro-operations. A dispatch unit within the processor core receives a vector memory operation. The dispatch unit sends the vector memory operation to a first vector input queue of multiple vector input queues. The sending is based on the memory addressing mode. A micro-operation sequencer splits the vector memory operation into one or more memory micro-operations, which includes forwarding each micro-operation within the one or more micro-operations to a first memory queue within multiple memory queues. A memory operation is then issued to a load-store unit within the processor core. The issuing includes selecting, from the multiple memory queues, the memory operation. The vector memory operation comprises either a vector load operation or a vector store operation.
Owner:AKEANA INC

Self-healing virtualized file server

In one embodiment, a system for managing a virtualization environment comprises a plurality of host machines, one or more virtual disks comprising a plurality of storage devices, a virtualized file server (VFS) comprising a plurality of file server virtual machines (FSVMs), wherein each of the FSVMs is running on one of the host machines and conducts I / O transactions with the one or more virtual disks, and a virtualized file server self-healing system configured to identify one or more corrupt units of stored data at one or more levels of a storage hierarchy associated with the storage devices, wherein the levels comprise one or more of file level, filesystem level, and storage level, and when data corruption is detected, cause each FSVM on which at least a portion of the unit of stored data is located to recover the unit of stored data.
Owner:NUTANIX INC

Software and hardware collaborative optimization method of hybrid in-memory architecture

The invention discloses a software and hardware collaborative optimization method for a hybrid in-memory architecture, which comprises the following steps of: performing joint feature representation on a target AI algorithm and an in-memory computing architecture, extracting context features, and parameterizing in-memory computing unit configuration, a neural network processor assembly line, a multi-core interconnection topology and a storage level interface of the in-memory computing architecture; an off-line reference data set is constructed, a predictive agent model is trained, discrete architecture parameters are processed by the model by adopting an embedding method, feature association is learned by applying an encoder with a self-attention mechanism, joint prediction of multi-dimensional PPA indexes is realized through a parallel prediction network, and a feasible region constraint learning mechanism is introduced in training; the trained agent model is embedded into a multi-objective evolutionary algorithm, the energy efficiency ratio, the average computing power utilization rate, the model execution delay and the like serve as optimization objectives, chip area efficiency and power consumption constraints are met at the same time, and a Pareto optimal in-memory computing architecture configuration set is searched.
Owner:SOUTH CHINA UNIV OF TECH

System using transformer architecture with quantization-aware non-linear approximation and near-memory computing

This invention proposes a GQA-LUT method, utilizing a genetic algorithm and LUT-based circuit to efficiently approximate non-linear operators in Transformers. It adaptively finds optimal solutions for various non-linear functions, outperforming conventional neural network methods. A novel rounding mutation (RM) algorithm enhances approximation accuracy during quantization, improving low-bit integer precision. The invention also introduces a LayerNorm folding strategy as a near-memory computing principle, reducing IO and energy overheads with a two-stage memory hierarchy. Additionally, an additive partial sum quantization method is proposed to reduce energy consumption by quantizing accumulated PSUMs in matrix multiplication, alongside a PSQ-APSQ grouping strategy and floating-point regularization.
Owner:THE HONG KONG UNIV OF SCI & TECH +1

Reducing DC power consumption using dynamic memory prefetching for processing engines

Aspects of the disclosure are directed to dynamic memory prefetching. In accordance with one aspect, the disclosure includes a main memory configured to store data; a memory hierarchy coupled to the main memory, the memory hierarchy configured to augment the main memory; and a trusted zone element coupled to the memory hierarchy, the trusted zone element configured to read a shared memory to determine a specific control register and one or more prefetch values based on a secure monitor call (SMC). In another aspect, the disclosure includes reading a shared memory to determine a specific control register and one or more prefetch values based on a secure monitor call (SMC); and writing the one or more prefetch values to an exception level register.
Owner:QUALCOMM INC

Hardware-optimized recurrent neural network system

A system includes a machine-learning model implemented on a data processing apparatus, which features a parallel processor with a memory hierarchy. The machine-learning model is a recurrent neural network (RNN) with a multi-head architecture, comprising multiple sub- vectors that process parallel data streams. The RNN's weight matrix is structured as a block-diagonal matrix, allowing for parallel processing of the sub-vectors. A fused computational kernel executes an entire time-series processing loop for the multi-head RNN, maintaining the weight matrix blocks in on-chip memory and performing matrix multiplications and element-wise operations for each sub-vector in a single kernel execution.
Owner:NXAI GMBH

Vector quantization-based encoding cache methods, systems, electronic devices, and media

This application provides a method, system, electronic device, and medium for caching codebooks based on vector quantization. The method includes: acquiring a codebook to be processed; sorting the codebook to be processed to obtain a target codebook index; obtaining an index boundary based on the target codebook index and a codebook cache; the codebook cache being a cache for storing the codebook to be processed; comparing the index boundary and the target codebook index to obtain a judgment result; and storing the codebook to be processed in the corresponding cache within the codebook cache based on the judgment result. This application improves memory performance and execution efficiency by placing codebook entries at different locations in the GPU's memory hierarchy based on the codebook's usage frequency, thus solving the problems of low efficiency in shared memory and global memory.
Owner:SHANGHAI JIAOTONG UNIV +1

Dynamic ai model selection and pre-loading based on data temperature scoring and next prompt prediction

A system and method are provided for artificial intelligence (AI) model selection and loading. The method includes receiving an input data stream for an AI application and analyzing the stream to determine temperature scores based on topic frequency, importance weightage, and access frequency. The method also includes mapping these temperature scores to relevant expert AI models and predicting the temperature and context of the next prompt using historical data patterns through a sliding window that smooths temporary anomalies. The system dynamically selects AI models based on temperature scores and predicted contexts, then pre-loads these models into a memory hierarchy where higher-temperature models are placed in faster memory. The method processes incoming prompts using these pre-loaded models and outputs responses, improving processing efficiency through predictive model management. The system addresses challenges of managing large-scale AI models by implementing intelligent pre-loading mechanisms that anticipate and prepare for upcoming processing needs.
Owner:SK HYNIX NAND PRODUCT SOLUTIONS CORP

Compression Metadata Value Induced Read and Write Operation Conservation

In disclosed embodiments, compression circuitry may compress data blocks and compressed write data blocks and corresponding compression metadata to a certain level in a cache / memory hierarchy. In some embodiments, the compression circuitry is configured to detect a pre-determined set of data values in a block of data to be compressed. In response, the compression circuitry may write a special metadata value for the data block to data storage circuitry to indicate compression of the data block, without writing a compressed version of the data block to the data storage circuitry. This may advantageously improve performance and reduce power consumption for data blocks with certain data values (e.g., having uniform pixel / texel values).
Owner:APPLE INC

Vector Scatter and Gather with Single Memory Access

Disclosed embodiments provide techniques for improved performance in processing vector instructions. A processor core is accessed. The processor core is coupled to a memory hierarchy, and the processor core includes one or more vector execution units (VUs), and one or more load store units (LSUs). The processor core includes a vector register file (VRF). The VRF includes multiple vector registers, and each vector register includes multiple vector elements. Vector elements that have a source or destination in contiguous memory are identified. Load store units (LSUs) take advantage of the contiguous memory condition by executing a vector load or vector store operation as a single memory access, requiring a reduced number of clock cycles. The single memory access satisfies each memory operation for each vector element within the vector register file.
Owner:AKEANA INC

Methods and apparatus to access main memory

Systems, apparatus, articles of manufacture, and methods are disclosed. An example apparatus includes main memory, a memory hierarchy in circuit with the memory, buffer circuitry in circuit with the memory, and control circuitry to selectively couple programmable circuitry to the main memory via the cache circuitry or to the main memory via the buffer circuitry.
Owner:OPENCHIP & SOFTWARE TECHNOLOGIES SL

Dirty tracking bit compression

A cache controller of a cache assigns a dirty tracking bit for each dirty byte of a cache line. Once a predetermined interval has elapsed without any accesses to the cache line or to a cache set that includes the cache line, the cache controller compresses contiguous dirty tracking bits for each portion of the cache line. Compressing the dirty tracking bits for contiguous dirty portions of the cache line allows the cache to store more dirty data using fewer dirty tracking bits, reducing area cost and bandwidth among levels of a memory hierarchy.
Owner:ADVANCED MICRO DEVICES INC

SYSTEM AND METHOD FOR ADAPTIVE HYBRID Hardware PREFETCHING

An apparatus includes a processor core and a memory hierarchy. The memory hierarchy includes a main memory and one or more caches between the main memory and the processor core. A plurality of hardware prefetchers are coupled to the memory hierarchy, and a prefetch control circuit is coupled to the plurality of hardware prefetchers. The prefetch control circuit is configured to compare changes in the one or more cache performance metrics within the two or more sampling intervals and control operation of the plurality of hardware prefetchers in response to changes in the one or more performance metrics between at least a first sampling interval and a second sampling interval.
Owner:HUAWEI TECH CO LTD

Non-blocking unit stride vector instruction dispatch with micro-operations

Disclosed techniques enable vector instruction processing. A processor core is accessed. The processor core is coupled to a memory hierarchy, and is configured to execute vector operations, scalar operations, and micro-operations. A decode unit decodes a vector memory operation. The vector memory operation is associated with a unit stride addressing mode. The decoding includes dividing the vector memory operation into one or more vector memory micro-operations. A dispatch unit sends at least one vector micro-operation within the one or more vector micro-operations to a scalar request queue within a plurality of request queues. The at least one vector micro-operation is issued to a load-store unit within the processor core. The issuing includes selecting, from the plurality of request queues, the at least one vector memory micro-operation.
Owner:AKEANA INC

Method and system for one-dimensional signal extraction for various computer processors

Methods and systems for extracting a one-dimensional (1D) signal from a two-dimensional (2D) digital image along a projection line are provided herein. The method and system store the digital image in a memory hierarchy, where a non-blocking prefetch operation may fetch pixels from a main memory to a data cache. In response to the direction of the projection line, a prefetch plan, a pixel processing plan, and a prefetch distance are selected. The prefetch plan uses a first address order designed to facilitate efficient extraction of pixels from the main memory to the data cache for a given direction. The pixel processing plan uses a second address order designed to facilitate calculation of a one-dimensional signal along the projection line. A pixel processing plan is used in coordination with a prefetch plan to compute the one-dimensional signal such that the pixel is fetched from the main memory to the data cache in advance in response to an amount of time of the prefetch distance before the pixel operation uses the pixel.
Owner:COGNEX CORP

Dynamic ai model selection and pre-loading based on data temperature scoring and next prompt prediction

PCT designated stageWO2026143127A1Memory hierarchyModel management
A system and method are provided for artificial intelligence (Al) model selection and loading. The method includes receiving an input data stream for an Al application and analyzing the stream to determine temperature scores based on topic frequency, importance weightage, and access frequency. The method also includes mapping these temperature scores to relevant expert Al models and predicting the temperature and context of the next prompt using historical data patterns through a sliding window that smooths temporary anomalies. The system dynamically selects Al models based on temperature scores and predicted contexts, then pre-loads these models into a memory hierarchy where higher-temperature models are placed in faster memory. The method processes incoming prompts using these pre-loaded models and outputs responses, improving processing efficiency through predictive model management. The system addresses challenges of managing large-scale Al models by implementing intelligent pre-loading mechanisms that anticipate and prepare for upcoming processing needs.
Owner:SK HYNIX NAND PRODUCT SOLUTIONS CORP

Intelligent agent interaction memory cooperation method and device based on memory grading and medium

PendingCN121765041AAchieve efficient reuseImplement persistence managementDatabase distribution/replicationExecution paradigmsPersonalizationEngineering
The invention discloses an agent interaction memory cooperation method and device based on memory grading and a medium, and the method comprises the steps: obtaining the original interaction data of a user, and carrying out the context interaction analysis of the original interaction data, so as to obtain an interaction state variable; based on the interaction state variable, determining user interaction multiplexing state data through cross-session multiplexing of the user unique identifier; obtaining an iteratable shared knowledge base through high-frequency multiplexing knowledge precipitation according to the user interaction multiplexing state data; performing data interaction control on the interaction state variable, the user interaction multiplexing state data and the iteratable shared knowledge base to determine a directional calling state of the memory data; and based on the directional calling state of the memory data, through multi-level cache matching of agent demand analysis, obtaining agent interactive memory collaborative storage data. Through the method, the technical problem that an artificial intelligence interaction memory function cannot meet real-time interaction dynamic and personalized requirements in the prior art is solved.
Owner:浪潮智慧科技有限公司 +2

Cache control to save register data

Techniques related to eviction control for cache lines that store register data are disclosed. In some embodiments, a memory hierarchy circuit is configured to provide memory support for register operand data in one or more cache circuits. The lock circuitry may control a first set of lock indicators for a set of registers of a first thread, including one or more lock indicators to assert registers indicated by the decoding circuitry to be utilized by decoded instructions of the first thread. The lock circuit may hold register operand data in the one or more cache circuits, including to prevent eviction of a given cache line from the cache circuit with an assertion-based lock indicator. The lock circuit may clear the first set of lock indicators in response to a reset event. The disclosed techniques may advantageously retain relevant register information in a cache having a limited control circuit area.
Owner:APPLE INC

Blockchain-based two-dimensional code data security detection method and system

The application discloses a blockchain-based two-dimensional code data security detection method and system, relates to the technical field of data security detection, and is used for solving the problems that there is no fine storage hierarchy and dynamic management of access frequency in the prior art, data compliance and integrity verification usually depend on a single node or a centralized mode, the dynamic adaptation capability to data access frequency and risk is lacked, and the abnormal detection and response mechanism of most systems is passive; through a blockchain consensus mechanism and hash verification, the data is ensured to be non-tamperable and consistent, a hierarchical storage architecture and dynamic adjustment of a smart contract are used, the real-time detection and automatic response capability to potential threats is improved by combining an isolated forest algorithm with the smart contract, and the system can dynamically adjust strategies according to data access frequency and risk, and adapt to different demand changes.
Owner:NANJING SANLONG PACKAGING CO LTD

Memory scheduling method and device, electronic equipment, medium and program product

The invention provides a memory scheduling method and device, electronic equipment, a medium and a program product, and relates to the technical field of computers. The memory scheduling method comprises the steps of dividing memory hierarchies for the hybrid memory system on the basis of a performance index of a far-end memory; dynamically scanning the container based on the attribute of the memory data in the container to obtain a container access heat map; configuring a cold and hot threshold value of the container for the importance weight allocated to the container and the memory hierarchy based on the container access heat map; and performing memory page scheduling on the container based on the cold and hot threshold and the container access heat map. According to the technical scheme, scheduling of the memory page in the container gives consideration to data heat, container importance and dynamic data features, and rationality of cold and hot scheduling of the memory page is improved while global scheduling of the memory page is achieved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Cache control to preserve register data

Techniques are disclosed relating to eviction control for cache lines that store register data. In some embodiments, memory hierarchy circuitry is configured to provide memory backing for register operand data in one or more cache circuits. Lock circuitry may control a first set of lock indicators for a set of registers for a first thread, including to assert one or more lock indicators for registers that are indicated, by decode circuitry, as being utilized by decoded instructions of the first thread. The lock circuitry may preserve register operand data in the one or more cache circuits, including to prevent eviction of a given cache line from a cache circuit based on an asserted lock indicator. The lock circuitry may clear the first set of lock indicators in response to a reset event. Disclosed techniques may advantageously retain relevant register information in the cache with limited control circuit area.
Owner:APPLE INC

Systems, methods, and apparatuses for selecting devices in a tiered memory

A method can include receiving a request for a memory page in a memory hierarchy including a first memory device and a second memory device, wherein the first memory device has a first parameter, the second memory device has a second parameter, selecting the first memory device based on the first parameter and the second parameter, and allocating the memory page from the first memory device based on the request based on the selection. The selection can include determining a first result based on the first parameter, determining a second result based on the second parameter, and comparing the first result and the second result. Determining the first result can include combining the first parameter with a first weight. The first weight can include a first scaling factor, and combining the first parameter with the first weight can include multiplying the first parameter and the first scaling factor.
Owner:SAMSUNG ELECTRONICS CO LTD

CACHE STORAGE ACCESS

A method of data processing in a multiprocessor data processing system (200) comprising several vertical cache memory hierarchies supporting a plurality of processor cores (102a, 102b), a system memory (132), and a system connection connected to the system memory (132) and the several vertical cache memory hierarchies, wherein the method comprises: in response to receiving a request, loading and reserving from a first processor core (102a, 102b), outputting through a first cache memory in a first vertical cache memory hierarchy supporting the first processor core (102a, 102b), on a system connection, a memory access request for a target cache memory row of the request to load and reserve;In response to the memory access request and prior to receiving a system-wide coherence response for the memory access request, the first cache memory receives the target cache memory row from a second cache memory in a second vertical cache memory hierarchy through cache-to-cache intervention and an early specification of the system-wide coherence response for the memory access request; and in response to the early specification of the system-wide coherence response and prior to receiving the system-wide coherence response, the first cache memory initiates processing to update the target cache memory row in the first cache memory.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Cryptographic computing with context information for transient side channel security

In one embodiment, a processor includes a memory hierarchy that stores encrypted data, tracking circuitry that tracks an execution context for instructions executed by the processor, and cryptographic computing circuitry to encrypt / decrypt data that is stored in the memory hierarchy. The cryptographic computing circuitry obtains context information from the tracking circuitry for a load instruction to be executed by the processor, where the context information indicates information about branch predictions made by a branch prediction unit of the processor, and decrypts the encrypted data using a key and the context information as a tweak input to the decryption.
Owner:INTEL CORP

Weight-stationary matrix multiply accelerator with tightly coupled l2 cache

An accelerator is accessed. The accelerator includes a weight-stationary systolic array of one or more multiply-accumulate units. The accelerator is coupled to a memory hierarchy and a processor core. The processor core sends a work request to the accelerator. The work request is based on execution of a machine learning model and an activation matrix. In response to the work request, the accelerator loads a weight matrix and the activation matrix. The loading uses the memory hierarchy. The accelerator multiplies the weight matrix by the activation matrix. The multiplication results in an answer matrix. The accelerator stores the answer matrix in the memory hierarchy. The processor core obtains the answer matrix that was stored. The machine learning model is trained. The training produces the weight matrix, which is transposed and saved to the memory hierarchy.
Owner:AKEANA INC

Data stream mapping search method, system, electronic device and storage medium

The embodiments of the present application provide a data flow mapping search method, system, electronic device and storage medium, which obtain data types and corresponding divided dimensional data from the pulse neural network convolution process, and form multiple data flow mapping sets; obtain the number of cluster nodes and the number of processor cores, determine the total number of parallel threads of the data cluster, and create parallel threads; obtain the data flow mapping search number of each parallel thread based on the data flow mapping set and the total number of parallel threads, assign a data flow mapping to each parallel thread, and screen the data flow mapping according to the preset storage rules of each storage level to obtain multiple candidate data flow mappings; in each parallel thread, perform energy calculation on each candidate data flow mapping, and compare the energy calculation result with the historical energy calculation result. If the energy calculation result is lower than the historical energy calculation result, the candidate data flow mapping corresponding to the energy calculation result is used as the target data flow mapping.
Owner:PENG CHENG LAB