Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

267 results about "Memory interface" patented technology

Tensor Memory Accelerator Enhancements

One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster including a plurality of graphics cores and tensor processing circuitry. The tensor processing circuitry includes a local memory, a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply and accumulate operation, and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory coupled to the memory interface and the local memory. The tensor data movement accelerator includes circuitry configured to translate the tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Rendering equipment based on three-dimensional Gaussian sputtering

The invention discloses rendering equipment based on three-dimensional Gaussian sputtering. The hardware architecture includes a memory interface, a pre-processing engine, a cardinality ordering engine, a rasterization engine, and the like. The preprocessing engine converts the three-dimensional Gaussian point into a two-dimensional representation and generates a key value pair containing a tile identifier and depth information; the cardinal number sorting engine performs staged efficient sorting through a parallel processing mechanism and a BRAM alternate working mode; and the rasterization engine performs rendering processing by adopting a full-pipeline design and an early-stage stopping mechanism. According to the method, parallel execution of depth sequencing and rasterization is achieved, memory access and calculation redundancy are reduced through optimized data flow design, an efficient and low-power-consumption hardware acceleration solution is provided for three-dimensional Gaussian sputtering rendering, and the method is particularly suitable for application scenes such as virtual reality, games and scientific visualization requiring real-time rendering.
Owner:TSINGHUA UNIVERSITY

Combined MX and sparsity representation

One embodiment provides a graphics processor comprising a memory interface and a processing cluster array including a plurality of processing resources interconnected via a switched interconnect network, at least one of the plurality of processing resources including a matrix accelerator configured to execute an instruction to perform a multi-dimensional sparse matrix multiply and accumulate operation having input in a sparse microscaling format including merged sparsity and scaling metadata.
Owner:INTEL CORP

Tensor data moving accelerator

The invention relates to a tensor data movement accelerator. One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster comprising a plurality of graphics cores and tensor processing circuitry. The tensor processing circuit includes: a local memory; a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply-accumulate operation; and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory and a local memory coupled to the memory interface. The tensor data moving accelerator includes circuitry configured to convert tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Speculative execution of kernel programs in a chiplet based architecture

One embodiment provides a multi-chiplet graphics processor comprising a plurality of chiplets, where a chiplet of the plurality of chiplets comprise a memory interface, processing resources configured to execute threads of a kernel, and thread dispatch circuitry to facilitate dispatch of threads of the kernel to the processing resources. The processing resources are configured to execute threads of a first kernel, receive dispatch of threads of a second kernel for execution before completion of the first kernel as threads of the first kernel retire, execute a first phase of the second kernel during completion of execution of the first kernel, via a thread of the first kernel, signal an event via an uncached write to a global memory, and execute a second phase of the second kernel based on detection of the event via an uncached read from the global memory.
Owner:INTEL CORP

Memory module and electronic device

The present disclosure belongs to the technical field of storage chip designs. Disclosed are a memory module and an electronic device. The memory module includes at least two memory expander controller (MXC) chips, a peripheral component Interconnect express (PCIe) golden finger, a multichannel input / output (MCIO) connector, and a dual inline memory module (DIMM), wherein the MXC chips are connected to the DIMM through a double data rate 5 (DDR5) synchronous dynamic random access memory (DRAM) interface controller port, a compute express link (CXL) port of at least one of the MXC chips interacts with an external device through the PCIe golden finger, and the CXL port of at least one of the MXC chips interacts with the external device through the MCIO connector. The present disclosure may improve the memory capacity of the memory module.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Memory system with processor in memory (PIM)

A memory system includes a memory interface including a first sub-channel interface associated with a first plurality of memory banks and a first plurality of processor in memory (PIM) blocks and further including a second sub-channel interface associated with a second plurality of memory banks and a second plurality of PIM blocks. During a particular mode of operation associated with the memory system, the first sub-channel interface is configured to communicate, with a host device, one or more memory access commands associated with the first plurality of memory banks. During the particular mode of operation, the second plurality of PIM blocks are configured to perform, concurrently with communication of the one or more memory access commands, one or more PIM operations associated with the second plurality of memory banks. The second sub-channel interface is configured to be disabled during the particular mode of operation.
Owner:QUALCOMM INC

AI accelerator integrated circuit chip with integrated cell-based fabric adapter

An integrated circuit formed on (i) a single semiconductor die or (ii) a plurality semiconductor dies that are integrated into a single package. The integrated circuit may include a communication interface including a serializer / deserializer (SerDes) interface; a fabric adapter communicatively coupled to the communication interface; a plurality of inference engine clusters, each inference engine cluster including a respective memory element and / or memory interface; and a data interconnect communicatively coupling each respective memory element and / or memory interfaces of the plurality of inference engine clusters to the fabric adapter. The fabric adapter may be configured to facilitate remote direct memory access (RDMA) read and write services and / or datagram communication over a cell-based switch fabric to and from the respective memory elements and / or memory interfaces of the plurality of inference engine clusters via the data interconnect.
Owner:TENSORDYNE INC

Speculative execution of kernel programs in chiplet-based architectures

The invention relates to speculative execution of kernel programs in a chiplet-based architecture. One embodiment provides a multi-chiplet graphics processor comprising a plurality of chiplets, where a chiplet of the plurality of chiplets comprises a memory interface; a processing resource configured to execute a thread of the kernel; and a thread dispatch circuitry module to facilitate dispatch of threads of the kernel to the processing resources. The processing resource is configured to: execute a thread of the first core; when the threads of the first core are revolved, receiving the dispatch of the threads of the second core for execution before the first core is completed; executing the first stage of the second core during execution completion of the first core; signaling, via a thread of the first core, an event via an uncached write to the global memory; and performing a second stage of the second kernel based on the detection of the event via the uncached read from the global memory.
Owner:INTEL CORP

Aggregation of multiple memory modules for a system-on-chip

A system includes a substrate comprising a first circuit. The system also includes an integrated circuit formed in a first die disposed on the substrate. The integrated circuit includes at least a processor, a controller, and a first memory interface. The first memory interface is located in a first edge of the first die and is configured to couple to the first circuit. The system also includes a first buffer circuit formed in a second die disposed on the interposer substrate adjacent to the first edge of the first die. The first buffer circuit includes a second memory interface configured to couple to the first connection circuit. The system further includes multiple memory modules disposed on the second die. Each of the multiple memory modules at least partially share the second memory interface to communicate with the integrated circuit.
Owner:AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD

Systems and methods for memory replay protection

Systems and methods for memory replay protection are disclosed. In an example, a system in the form of a device includes a processor, and a replay protection circuit coupled to the processor and including a memory interface circuit, wherein the replay protection circuit is configured to generate a message authentication code (MAC) from at least one block of data and to access the at least one of block of data via the memory interface circuit using ciphertext that is generated by the replay protection circuit in response to a plaintext memory address received from the processor.
Owner:PENSANDO SYSTEMS INC

AI accelerator integrated circuit chip with integrated cell-based fabric adapter

An integrated circuit formed on (i) a single semiconductor die or (ii) a plurality semiconductor dies that are integrated into a single package. The integrated circuit may include a communication interface including a serializer / deserializer (SerDes) interface; a fabric adapter communicatively coupled to the communication interface; a plurality of inference engine clusters, each inference engine cluster including a respective memory element and / or memory interface; and a data interconnect communicatively coupling each respective memory element and / or memory interfaces of the plurality of inference engine clusters to the fabric adapter. The fabric adapter may be configured to facilitate remote direct memory access (RDMA) read and write services and / or datagram communication over a cell-based switch fabric to and from the respective memory elements and / or memory interfaces of the plurality of inference engine clusters via the data interconnect.
Owner:TENSORDYNE INC

Configurable fabric bandwidth throttling for GPU virtualized workloads

One embodiment provides a graphics processor comprising a memory interface, a graphics core cluster including a plurality of graphics cores, and an interconnect fabric to interconnect a plurality of hardware clients including the plurality of graphics cores. The interconnect fabric include a plurality of fabric ports coupled with the plurality of graphics cores. A fabric port is configured to limit bandwidth available to an associated graphics core via a bandwidth throttler circuit coupled with the fabric port.
Owner:INTEL CORP

Combining MX and sparsity representations

The name of the invention is combining MX and sparsity representation. One embodiment provides a graphics processor comprising a memory interface and a processing cluster array comprising a plurality of processing resources interconnected via a switched interconnect network, at least one of the plurality of processing resources comprising a matrix accelerator, the matrix accelerator is configured to execute instructions to perform multi-dimensional sparse matrix multiplication and accumulation operations with inputs in a sparse microscaling format that includes merged sparsity and scaling metadata.
Owner:INTEL CORP

Mixed format hardware and instruction

One embodiment provides a graphics processor comprising a memory interface and a processing resource coupled with the memory interface. The processing resource including a mixed format functional unit configured to receive an integer input and a floating-point input, perform a fused format conversion multiply operation on the integer input and the floating-point input to generate a floating-point result, and output the floating-point result.
Owner:INTEL CORP

Multi-core processor chip and storage access method and device of multi-core processor chip

The invention provides a multi-core processor chip and a storage access method and device of the multi-core processor chip, and relates to the technical field of computer processors, the multi-core processor chip comprises at least one processor core, the processor core comprises a general processor core, a first address space mapping module and a first core interconnection interface, the universal processor core supports high-speed cache consistency; at least one input / output core grain, wherein the input / output core grain comprises a memory interface, a directory, a second address space mapping module and a second core grain interconnection interface; both the first address space mapping module and the second address space mapping module store an address space mapping table, and the address space mapping table stores a mapping relationship between a memory address accessed by a processor core grain and a memory interface of an input / output core grain; the directory stores cache consistency information of the cache blocks corresponding to the memory interfaces of the input / output core grain and other input / output core grains corresponding to the directory.
Owner:BEIJING VCORE TECH CO LTD

AI processor and method based on storage and calculation integration, three-dimensional integration and operator separation

The invention relates to an AI processor and method based on storage and calculation integration, three-dimensional integration and operator separation. The AI processor comprises a storage layer; the calculation layer and the storage layer are stacked in the vertical direction through a three-dimensional integrated connection structure, and a three-dimensional integrated high-speed channel is generated and used for executing calculation tasks in a large language model; the standardized memory interface is used for being connected with an external main control chip; the calculation task comprises a pre-filling stage and a decoding stage; the scheduling module configures the main control chip to interact data with the storage layer through the standardized memory interface so as to execute the pre-filling stage; the scheduling module configures the computing layer to interact data with the storage layer through the three-dimensional integrated high-speed channel so as to execute the decoding stage, the pre-filling stage is processed by utilizing the large computing power advantage of a main control chip, and the decoding stage is processed by utilizing the three-dimensional stacked high-bandwidth advantage, so that accurate matching of the computing power and the bandwidth is realized; and the large model reasoning efficiency is obviously improved.
Owner:WUXI MICRONANO CORE ELECTRONIC TECH CO LTD +1

Vector and matrix calculation-oriented memory access system

The invention provides a memory access system oriented to vector and matrix calculation, the system comprises a vector memory access unit, a matrix memory access unit, a vector register group and a matrix register group, the vector memory access unit is connected with a memory interface and the vector register group, reads elements of a one-dimensional data structure or a two-dimensional data structure from a memory, and stores the elements of the one-dimensional data structure or the two-dimensional data structure; vector data are generated through data reorganization operation, a matrix access unit is connected with a memory interface and a matrix register set, matrix block data of a two-dimensional data structure are read from a memory together with the matrix access unit, matrix data are generated after data reorganization, and the matrix data are broadcasted to one or more computing units according to rows or columns. And performing calculation on the matrix data and the vector data. In the memory access system, the vector memory access unit and the matrix memory access unit can load data in parallel, and the utilization rate is improved through a plurality of computing units, so that the problems of low memory access efficiency and low memory bandwidth utilization rate are solved.
Owner:NANJING UNIV

Adaptive virtualization of GPU cores and engine based virtualization

One embodiment provides a graphics processor comprising a memory interface, a plurality of interfaces to a plurality of compute engines, a processing resource cluster including a plurality of processing resources, the plurality of processing resources configured to execute instructions on behalf of the plurality of compute engines, and virtualization circuitry configured to enable time-sliced virtualization of the plurality of processing resources via the plurality of compute engines, wherein the virtualization circuitry to concurrently process workloads from a plurality of guest software environments during a time-slice via dynamic assignment of the workloads to the plurality of interfaces to the plurality of compute engines.
Owner:INTEL CORP

Hardware accelerated random number generation

One embodiment provides a graphics processor comprising a memory interface and a processing resource coupled with the memory interface. The processing resource includes multiple processing lanes, each of the multiple processing lanes including circuitry dedicated to generation of one or more randomized numbers. The processing resource configured to receive an instruction to generate a two-dimensional matrix of randomized numbers, generate one or more hardware generated seed values for use by each of the multiple processing lanes, generate one or more randomized numbers at each of the multiple processing lanes based on the one or more hardware generated seed values and output the two-dimensional matrix of randomized numbers to a destination register.
Owner:INTEL CORP

Early exit based on activation of RELU

The name of the invention is' early exit of RELU-based activation '. One embodiment provides a graphics processor comprising a memory interface and a processing resource coupled with the memory interface. The processing resources include circuit modules configured to perform operations fused with rectified linear unit operations. The circuit module is configured to detect a negative output of the operation prior to completion of the operation, clock-gate a portion of the circuit module, and operate an output zero for the rectified linear unit.
Owner:INTEL CORP

Adaptive virtualization of GPU cores, and engine-based virtualization

One embodiment provides a graphics processor comprising: a memory interface; a plurality of interfaces to the plurality of compute engines; a cluster of processing resources, the cluster of processing resources comprising a plurality of processing resources configured to execute instructions on behalf of the plurality of compute engines; and virtualization circuitry configured to enable time division virtualization of the plurality of processing resources via the plurality of computing engines, the virtualization circuitry is to concurrently process workloads during a time slice via dynamic assignment of workloads from a plurality of guest machine software environments to a plurality of interfaces to a plurality of compute engines.
Owner:INTEL CORP

Remote access solution for compute express link (CXL) memory interface configured to send information regarding a capacity, a latency, or bandwidth of memory via interface

A system with an interface for remote memory. In some embodiments, the system includes: an interface circuit having: a first interface, configured to be connected to a processing circuit; and a second interface, configured to be connected to memory, the first interface including a cache coherent interface, and the second interface being different from the first interface.
Owner:SAMSUNG ELECTRONICS CO LTD

Memory time sequence parameter adjusting and optimizing method and device, computer equipment and storage medium

The invention discloses a memory time sequence parameter tuning method and device, computer equipment and a storage medium, relates to the technical field of server memories, and dynamically relates to time sequence parameter combination and memory interface signal quality by establishing an automatic closed-loop process of parameter configuration, signal measurement, condition verification and parameter optimization. Parameter suitability is accurately judged based on an actual measurement result, and a parameter optimization mechanism is driven to automatically iterate to generate a new candidate parameter combination, so that dependence on artificial experience is eliminated, and the limitation that a traditional method is low in efficiency and difficult to deal with dynamic changes in a complex parameter space is overcome. Therefore, the technical problems that in the prior art, memory time sequence parameter tuning is low in efficiency, poor in adaptability and insufficient in stability are solved, and the technical effects of remarkably improving the integrity of memory signals, enhancing the running stability of a server and improving the energy efficiency ratio of a memory subsystem are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Verification system of FSMC multi-memory interface based on UVM

The invention discloses a verification system of an FSMC multi-memory interface based on a UVM, and belongs to the technical field of chip verification. The verification system of the FSMC multi-memory interface based on the UVM comprises a test layer, a simulation instruction definition layer, an AHB bus interface component, an assertion inspection component and an external memory interface component. The problems that in the prior art, a single memory is mainly connected for verification, only a single protocol can be verified, the verification means is limited, and the efficiency and the coverage rate are not high are solved, the transmission type of the AHB is defined in the environment configuration component, therefore, the verification diversity is improved, the verification requirements for verifying different memory interfaces are well met, and the verification efficiency is improved. According to the interface protocol verification method and device, the randomized excitation and self-checking function conforming to the protocol can be generated, the verification comprehensiveness is greatly improved, the pertinence and correctness of the verified module are greatly improved, the verification efficiency is improved, the code cost is reduced, and the accuracy of interface protocol verification is improved.
Owner:HITENX (WUXI) TECH CO LTD +1

Early exit for relu-based activation

One embodiment provides a graphics processor comprising a memory interface and a processing resource coupled with the memory interface. The processing resource including circuitry configured to perform an operation fused with a rectified linear unit operation. The circuitry is configured to detect a negative output of the operation before completion of the operation, clock gate a portion of the circuitry, and output a zero value for the rectified linear unit operation.
Owner:INTEL CORP

Hybrid format hardware and instructions

The name of the invention is hybrid format hardware and instructions. One embodiment provides a graphics processor including a memory interface and a processing resource coupled with the memory interface. The processing resource includes a mixed format functional unit configured to receive an integer input and a floating point input; performing a fused format conversion multiplication operation on the integer input and the floating point input to generate a floating point result; and outputting a floating point result.
Owner:INTEL CORP

Multi-chip module (MCM) with multi-port unified memory

Semiconductor devices, packaging architectures and associated methods are disclosed. In one embodiment, an integrated circuit (IC) base die is disclosed. The IC base die is configured to couple to a stack of memory die and includes a first port including a die-to-die (D2D) interface to couple to an IC device. A second port includes a memory interface to access a memory other than the stack of memory die. Memory control circuitry controls memory access operations directed to the memory other than the stack of memory die.
Owner:ELIYAN CORP

Memory module and method for writing data thereto

PendingUS20260186990A1Memory interfaceData memory
Various aspects relate to a memory module including: a memory interface; a non-volatile storage device for persistently storing data; a non-volatile memory device providing a write first in first out, FIFO, buffer in hardware, the write FIFO buffer being configured to receive, via the memory interface, and to store packets associated with one or more atomic transactions, wherein each of the one or more atomic transactions includes a respective plurality of packets and indicates corresponding data to be written to the memory module; a memory controller configured to write the corresponding data of an atomic transaction of the one or more atomic transactions to the memory module in the case that the write FIFO buffer stores the respective plurality of packets of the atomic transaction.
Owner:FERROELECTRIC MEMORY GMBH

Reference signals in active TCI switching

Techniques discussed herein can facilitate the use of reference signals in active TCI switching. One example aspect is a user equipment (UE), comprising: a memory interface; and processing circuitry communicatively coupled to the memory interface and configured to: receive a physical downlink shared channel (PDSCH) message that includes an activation command indicating a target transmission configuration indicator (TCI) state; decode the activation command within a decoding period, where the decoding period is a time period allocated for the UE to decode the activation command; receive a TCI resource where the TCI resource is a reference signal; perform time and frequency tracking associated with the target TCI state according the TCI resource; and switch to the target TCI state after performing time and frequency tracking.
Owner:APPLE INC