Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

95results about "Single instruction multiple data multiprocessors" patented technology

Method and system for in-line data conversion outside of a machine learning hardware

A system includes a component configured to send data in a first data format. The system includes a direct memory access (DMA) engine configured to receive the data in the first data format and convert the first data format to a second data format, wherein the second data format is associated with a data format of a machine learning (ML) hardware, wherein the second data format is different from the first data format. The ML hardware is configured to receive the data in the second format and perform at least one ML operation on the received data in the second format. The received data in the second data format is stored on an on-chip memory (OCM) of the ML hardware.
Owner:MARVELL ASIA PTE LTD

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Integrated circuit and operating method thereof

The embodiment of the invention provides an integrated circuit and an operation method thereof. An integrated circuit may include a plurality of in-memory computing (CIM) circuits physically formed on a substrate. Each of the plurality of CIM circuits may include: an input circuit configured to receive a plurality of first data elements; a memory array coupled to the input circuit and configured to store the plurality of first data elements; a data multiplexer configured to output the plurality of first data elements through a first data path or through a second data path; and a plurality of arithmetic units coupled to the data multiplexer and configured to perform a multiply-accumulate (MAC) operation on the plurality of first data elements and the plurality of second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

In-memory computing accelerator using high-density operation circuit and low-power sense amplifier as peripheral circuit

An in-memory computing (IMC) accelerator using a high-density operation circuit and a low-power sense amplifier as a peripheral circuit includes a plurality of dynamic random-access memory (DRAM) banks each including a pair of cell arrays, a data supply logic, a memory, and a controller for IMC, a global SRAM, and a top-level controller, wherein the cell array includes a plurality of subarrays, each of the subarrays includes a DRAM array including a big array and a little array, and an arithmetic circuit configured to perform an operation, and the arithmetic circuit includes a sense amplifier configured to amplify a bit line voltage difference, and a compact multiply-accumulate (MAC)-single instruction multiple data (SIMD) unit (CMSU) for an MAC operation and an SIMD operation, so that functionality of an in-memory operation is diversified.
Owner:KOREA ADVANCED INST OF SCI & TECH

Configurable wavefront parallel processor

An apparatus comprising: at least one processing element configured to process a data flow in at least one direction of a plurality of directions; a configuration register comprising at least one setting that determines the processing of the data flow with the at least one processing element; and a shift register configured to select data of the at least one processing element from the at least one direction, and to provide at least one shifted data sample to a plurality of slices configured to perform at least one arithmetic operation with the data flow; wherein at least one slice of the plurality of slices is configured with the at least one setting of the configuration register.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Vertical and horizontal broadcast of shared operands

The present invention provides an apparatus, method, and storage medium for reducing the bandwidth of a memory fabric and increasing its efficiency in an array processor system. [Solution] The array processor 300 includes processor element arrays 311 to 384 distributed in rows and columns and performing operations on parameter values; a memory interface that broadcasts a set of parameter values ​​to mutually exclusive subsets of rows and columns of the processor element arrays; SIMD units 310 to 380 that include a subset of the processor element array for the corresponding row; and DMA engines 301 to 304 interconnected with mutually exclusive subsets of the processor element arrays 311 to 384.
Owner:ADVANCED MICRO DEVICES INC

Memory circuits and methods for encoder / decoder dual mode for compute-in-memory

An integrated circuit may comprise a plurality of compute-in-memory (CIM) circuits physically formed on a substrate. Each of the plurality of CIM circuits may comprise: an input circuit configured to receive a plurality of first data elements; a memory array coupled to the input circuit and configured to store the plurality of first data elements; a data multiplexer configured to output the plurality of first data elements through a first data path or through a second data path; and a plurality of computing cells coupled to the data multiplexer and configured to perform multiply-accumulate (MAC) operations on the plurality of first data elements and a plurality of second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Processing data using accelerators in a system on a chip

In various examples, systems and methods are disclosed that relate to processing data using accelerators in a system on a chip. For example, a plurality of processing elements (PEs) can interconnect to form a processing engine, and a control system can control operation of the PEs based at least on the connections between the PEs. In some examples, the PEs can receive sub-inputs and transfer the sub-inputs to one or more other PEs to enable performance of the instructed operations. In examples, once the PEs complete the instructed operations, the sub-inputs can be transferred out of the processing engine.
Owner:NVIDIA CORP

Flexibly programmable processing unit

A digital data processing unit comprises: inputs for receiving input data; outputs for providing output data; one or more processing elements, and one or more switches. Each processing element have one or more inputs for receiving input data and one or more outputs for providing output data. Each processing element is configured to apply, during a time slot, a mathematical operation to its input data so as to generate output data. Each switch has outputs and at least one input for receiving a respective input value, wherein each switch is operated during a time slot based on control data to provide each received input value to one of its outputs selected based on the control data; wherein the one or more processing elements and the one or more switches are interconnected to form data paths between the inputs of the processing unit and the outputs of the processing unit, wherein at least one of the data paths is configured to provide, at a corresponding output of the processing unit, a final mathematical result from intermediate mathematical results generated by the one or more processing elements in the considered data path.
Owner:NOKIA SOLUTIONS & NETWORKS OY

MEMORY CIRCUITS AND METHOD FOR ENCODER / DECODER DUAL MODE FOR COMPUTE-IN-MEMORY

An integrated circuit can have multiple compute-in-memory (CIM) circuits physically implemented on a substrate. Each of the multiple CIM circuits can have: an input circuit configured to receive multiple first data elements; a memory array coupled to the input circuit and configured to store the multiple first data elements; a data multiplexer configured to output the multiple first data elements via a first data path or a second data path; and multiple arithmetic cells coupled to the data multiplexer and configured to perform multiplication-accumulation (MAC) operations on the multiple first data elements and multiple second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Calculation device and data transfer method

An accelerator (10) includes multiple PEs (12). The multiple PEs (12) are represented by a coordinate system, which includes two dimensions of X direction and Y direction and two or more different dimensions indicating different directions. Each PE (12) is capable of performing data input or data output with another PE (12) adjacent in the X direction or the Y direction, and is also capable of performing data input or data output with another PE (12) adjacent in different dimension.
Owner:DENSO CORP

Quantum block cipher resisting coprocessor for computer monitoring system and operation method

The invention discloses an anti-quantum block cipher coprocessor operation method for a computer monitoring system, and the method comprises the following steps: S1, loading data to generate a random number and a secret key, and splicing the random number and the secret key with a preset initialization vector to generate an input state; s2, receiving the associated parameters and the plaintext, and performing preprocessing to generate grouping mask shares; s3, performing parallel encryption on the grouped mask shares based on an input state to generate a plurality of mask ciphertexts and mask authentication tags; and S4, performing linear reconstruction on the mask ciphertext to finally run to obtain a complete ciphertext, decomposing plaintext data into a plurality of mask shares for parallel processing, effectively diluting data structure features, and fundamentally cutting off a way of acquiring key information by an attacker through power consumption analysis. Compared with the prior art, the parallel computing advantage of reconfigurable hardware is fully utilized, efficient execution of cryptographic operation is achieved, the whole data processing process is completed in the register, and the influence of memory access delay on system performance is remarkably reduced.
Owner:HANGZHOU HUADIAN BANSHAN POWER GENERATION

Processor system for performing neural network computations

A processor system (1) is disclosed herein for performing neural network computations. The processor system (1) comprises a plurality of synchronously operating processor clusters arranged in a two-dimensional array, wherein a first plurality of processor clusters in a first column of the array are logically coupled to a second plurality of processor clusters in a second, directly adjacent column by at least one shared external bus. Each processor cluster comprises at least one processor unit including a programmable processor and an instruction memory, and further comprises a data memory. The processor system is configured to allow a first programmable processor of one of the first plurality of processor clusters to write a value to the at least one shared external bus and to allow a plurality of other programmable processors comprised in the second plurality of processor clusters to concurrently read the value from the at least one external bus.
Owner:EUCLYD BV

Method, accelerator, and electronic device with tensor processing

A processor-implemented tensor processing method includes: receiving a request to process a neural network including a normalization layer by an accelerator; and generating an instruction executable by the accelerator in response to the request, wherein, by executing the instruction, the accelerator is configured to determine an intermediate tensor corresponding to a result of performing a portion of operations included in the normalization layer, by performing, in a channel axis direction, a convolution based on: a target tensor on which the portion of operations is to be performed; and a kernel having a number of input channels and a number of output channels determined based on the target tensor and including elements of scaling values determined based on the target tensor.
Owner:SAMSUNG ELECTRONICS CO LTD

Interconnect device, operation method of interconnect device, and artificial intelligence (AI) accelerator system including interconnect device

An interconnect device may include one or more hardware-implemented modules configured to: receive a command from a processing core; perform, based on the received command, an operation including either one or both of an accumulation operation on sets of data stored in a memory and an aggregation operation on results processed by the processing core; and provide a result of the performing of the operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Apparatus, method, and computer program for collecting diagnostic information

The device has a processing circuit for performing Single Instruction Multiple Data (SIMD) processing on an array having multiple data items, and the processing circuit supports SIMD processing for multiple array sizes. The processing circuit can select the array size for which to perform SIMD processing based on SIMD processing configuration information. A diagnostic information acquisition circuit is provided to collect diagnostic information about the software running on the processing circuit, and the diagnostic information acquisition circuit filters the collection of diagnostic information based on the SIMD processing configuration information.
Owner:ARM LTD

Rearranging data among processing elements of computational memory

An array of interconnected processing elements is modelled as a graph of nodes. Each layer of the graph represents a possible arrangement of data elements within the array of interconnected processing elements. An edge between nodes of adjacent layers of the graph represents a movement of a data element between the nodes. Constraints are set for a starting arrangement of the data elements stored in the array, an ending arrangement of the data elements stored in the array, and a limit for each node of the graph to have one input edge from a previous layer and one output edge to a subsequent layer. The model and constraints are processed with an integer programming solver to obtain a program of movements of data elements among the interconnected processing elements. The program implements a rearrangement of the data elements from the starting arrangement to the ending arrangement.
Owner:AT-MEMORY COMPUTING LP

Gather accelerated address space

In a system including a processing unit and a set of one or more stacked memory chips, the processing unit can request data. When the data is distributed such that there is at least one non-contiguous memory sector in the smallest unit of memory segments usable by the system, then a gather operation can be utilized to instruct the set of one or more stacked memory chips to gather the requested data into a virtual address space, e.g., a gather accelerated address space. The requested data can be aligned to the byte chunk size used by the processing unit and at least some of the unneeded memory segments can be skipped, e.g., not copied into the virtual address space. The requested data in the virtual address space can be communicated to the processing unit using less bandwidth resources than when not using the gather operation.
Owner:NVIDIA CORP

Vertical and horizontal broadcast of shared operands

The array processor includes an array of processor elements distributed in rows and columns. The processor element array performs operations on parameter values. The array processor also includes a memory interface that broadcasts a set of parameter values ​​to mutually exclusive subsets of the rows and columns of the processor element array. In some cases, the array processor includes a single instruction, multiple data (SIMD) unit that includes a subset of the processor element array of a corresponding row, a workgroup processor (WGP) that includes a subset of the SIMD units, and a memory fabric configured to interconnect with an external memory that stores the parameter values. The memory interface broadcasts the parameter values ​​to the SIMD unit that includes the row of processor element array associated with the memory interface and the column of processor element array implemented across the SIMD units in the WGP. The memory interface accesses the parameter values ​​from the external memory via the memory fabric.
Owner:ADVANCED MICRO DEVICES INC

Calculation device and data transfer method

An accelerator (10) includes multiple PEs (12). The multiple PEs (12) are represented by a coordinate system, which includes two dimensions of X direction and Y direction and two or more different dimensions indicating different directions. Each PE (12) is capable of performing data input or data output with another PE (12) adjacent in the X direction or the Y direction, and is also capable of performing data input or data output with another PE (12) adjacent in different dimension.
Owner:DENSO CORP

Data processing devices and methods

A data processor is suggested comprising at least an instruction issue stage issuing instructions, a number of processing elements to which at least some of the instructions are issued and which receive operand data, generate result data in accordance with the instructions received and transmit their result data to other processing elements for use as new operands; and a bus system for these transmissions, wherein the instruction issue stage is adapted to issue instructions to a group of processing elements to operate them in at least two different modes, namely an out-of-order mode wherein instructions may be executed out of order and a loop acceleration mode wherein loops can be executed efficiently, and wherein the bus system comprises an arbiter operative in the out-of-order mode to arbitrate access of the group of processing elements to at least a part of the bus system and inoperative in the loop acceleration mode.
Owner:UBITIUM GMBH

Use of a virtual mirror for a vehicle

This invention relates to the use of automotive virtual mirrors. A system, method, and computer-readable medium may include techniques for optimal use of automotive virtual mirrors. A gaze detector monitors the driver's eyes to determine whether the driver is looking in the direction of the virtual mirror. If the driver is not looking in the direction of the virtual mirror, all virtual mirrors are placed in a low operating mode. If the driver is looking in the direction of one of the virtual mirrors, the virtual mirror being viewed is placed in a high operating mode and all other virtual mirrors are placed in a low operating mode.
Owner:INTEL CORP

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Wide key hash table for a graphics processing unit

A wide hash key, that exceeds the word size of a GPU memory, is used to perform a key-value mapping by using paired hash tables configured in a multi-level tree configuration. The wide hash key is partitioned into segments, where each segment is used as a key into a respective paired hash table. The paired hash table has one hash table that stores an upper portion of an address and another hash table that stores the lower portion of the address. The upper and lower portions are combined to generate either an address to a paired hash table at the next level in the multi-level tree configuration or the address to the location of the value associated with the wide hash key.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Apparatus, method, and computer program for collecting diagnostic information

An apparatus has processing circuitry to perform single instruction, multiple data (SIMD) processing on an array with a plurality of data items and the processing circuitry supports the SIMD processing for a plurality of array sizes. The processing circuitry is able to select an array size with which to perform the SIMD processing based on SIMD processing configuration information. Diagnostic information collection circuitry is provided to collect diagnostic information about software executing on the processing circuitry and the diagnostic information collection circuitry filters collection of the diagnostic information based on the SIMD processing configuration information.
Owner:ARM LTD