Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

11results about "Single instruction multiple data multiprocessors" patented technology

Vertical and horizontal broadcast of shared operands

PendingJP2026094416ASingle instruction multiple data multiprocessorsConcurrent instruction execution
The present invention provides an apparatus, method, and storage medium for reducing the bandwidth of a memory fabric and increasing its efficiency in an array processor system. [Solution] The array processor 300 includes processor element arrays 311 to 384 distributed in rows and columns and performing operations on parameter values; a memory interface that broadcasts a set of parameter values ​​to mutually exclusive subsets of rows and columns of the processor element arrays; SIMD units 310 to 380 that include a subset of the processor element array for the corresponding row; and DMA engines 301 to 304 interconnected with mutually exclusive subsets of the processor element arrays 311 to 384.
Owner:ADVANCED MICRO DEVICES INC

Flexibly programmable processing unit

PendingEP4764822A1Single instruction multiple data multiprocessorsDigital data processing details
A digital data processing unit comprises: inputs for receiving input data; outputs for providing output data; one or more processing elements, and one or more switches. Each processing element have one or more inputs for receiving input data and one or more outputs for providing output data. Each processing element is configured to apply, during a time slot, a mathematical operation to its input data so as to generate output data. Each switch has outputs and at least one input for receiving a respective input value, wherein each switch is operated during a time slot based on control data to provide each received input value to one of its outputs selected based on the control data; wherein the one or more processing elements and the one or more switches are interconnected to form data paths between the inputs of the processing unit and the outputs of the processing unit, wherein at least one of the data paths is configured to provide, at a corresponding output of the processing unit, a final mathematical result from intermediate mathematical results generated by the one or more processing elements in the considered data path.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Calculation device and data transfer method

PendingEP4679280A4Single instruction multiple data multiprocessorsElectric digital data processing
An accelerator (10) includes multiple PEs (12). The multiple PEs (12) are represented by a coordinate system, which includes two dimensions of X direction and Y direction and two or more different dimensions indicating different directions. Each PE (12) is capable of performing data input or data output with another PE (12) adjacent in the X direction or the Y direction, and is also capable of performing data input or data output with another PE (12) adjacent in different dimension.
Owner:DENSO CORP

Wide key hash table for a graphics processing unit

ActiveEP3743821B1Memory architecture accessing/allocationSingle instruction multiple data multiprocessorsGraphicsProcessing element
A wide hash key, that exceeds the word size of a GPU memory, is used to perform a key-value mapping by using paired hash tables configured in a multi-level tree configuration. The wide hash key is partitioned into segments, where each segment is used as a key into a respective paired hash table. The paired hash table has one hash table that stores an upper portion of an address and another hash table that stores the lower portion of the address. The upper and lower portions are combined to generate either an address to a paired hash table at the next level in the multi-level tree configuration or the address to the location of the value associated with the wide hash key.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Neural network scheduling mechanism

PendingUS20260170596A1Single instruction multiple data multiprocessorsResource allocation
An apparatus to facilitate workload scheduling is disclosed. The apparatus includes one or more clients, one or more processing units to processes workloads received from the one or more clients, including hardware resources and scheduling logic to schedule direct access of the hardware resources to the one or more clients to process the workloads.
Owner:INTEL CORP

Neural processing device and load / store method of neural processing device

ActiveUS12657130B2Memory architecture accessing/allocationSingle instruction multiple data multiprocessors
A neural processing device is provided. The neural processing device comprises: a processing unit configured to receive an input activation and a weight and perform a two-dimensional matrix calculation with the input activation and the weight to generate an output activation, a first memory, and a load-store unit (LSU) configured to perform memory access operations between the first memory and a second memory. The memory access operations include a main memory access operation for a current processing operation that is performed by the processing unit, and a standby memory access operation for a standby processing operation that is performed by the processing unit after the current processing operation. A level of the first memory is equal to a level of the processing unit, and a level of the second memory is different from the level of the first memory.
Owner:NABTESCO CORP

Graph streaming neural network processing system and method thereof

PendingUS20260140768A1Program initiation/switchingSingle instruction multiple data multiprocessorsData packThread scheduling
Disclosed herein is a graph streaming neural network processing system comprising a first processor array, a second processor, and a thread scheduler. The thread scheduler dispatches a thread of a first node to the first processor array or the second processor, wherein the thread is executed to generate output data comprising a data unit stored in a private data buffer of the second processor. The thread scheduler determines that the data unit is sufficient for executing a thread of a second node. The second node is dependent on the output data generated by execution of a plurality of threads of the first node. Upon determining that the data unit is sufficient, the thread scheduler dispatches the thread of the second node. The thread scheduler determines to dispatch a subsequent thread of the first node for execution when a predefined threshold buffer size is available on the private data buffer.
Owner:BLAIZE INC

Artificial intelligence core, artificial intelligence core system, and loading / storing method for the artificial intelligence core system

InactiveJP7875536B2Memory architecture accessing/allocationSingle instruction multiple data multiprocessors
The present invention relates to an artificial intelligence core, an artificial intelligence core system, and a load / store method for an artificial intelligence core system. The artificial intelligence core includes: a process unit that receives input activations and weights and generates output activations through two-dimensional matrix calculations; and a load / store unit that performs load / store operations to transfer programs and input data received via an external interface to an on-chip buffer and transfer output data from the on-chip buffer to the external interface, the load / store operations including a main load / store operation for a currently executed operation performed by the process unit and a standby load / store operation for a standby execution operation performed by the process unit after the currently executed operation.
Owner:REBELLIONS INC

Multiple segments for a memory unit in a reconfigurable data processor

A non-transitory computer readable medium having instructions encoded thereon datapath configuring solutions for reconfigurable dataflow computing systems comprises a coarse-grained reconfigurable (CGR) processor a compiler configured to generate one or more configuration files for an application for execution on the CGR processor. The CGR processor includes an array of pattern compute units (PCUs) and pattern memory units (PMUs) configured to execute a dataflow graph. A PMU is coupled to a PCU via a multi-segment datapath pipeline The configuration file includes a portion of operation-specific data corresponding to an operation in the PMU. The CGR processor configures a configurable field in a segment of the multi-segment datapath pipeline. A PMU context including a set of configuration bits activates the segment corresponding to a portion of the operation-specific data, the PMU communicates the portion of the operation-specific data the to the PCU via the activated segment.
Owner:SAMBANOVA SYSTEMS INC