Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

151results about "Single instruction multiple data multiprocessors" patented technology

Method and apparatus for supporting distributed graphics and compute engines and synchronization in multi-dielet parallel processor architectures

This disclosure describes supporting distributed graphics and compute engines in a multi-dielet processor, such as, for example, a multi-dielet graphics processing unit (GPU), architectures and synchronization in such architectures. Each multi-dielet processor includes a hardware-implemented remapping capability and / or a hardware-implemented memory barrier capability.
Owner:NVIDIA CORP

Processor core, processor, and method for processor

The embodiment of the invention provides a processor core, a processor and a method for the processor. The processor core includes a processing pipeline configured to rename and execute a first instruction including at least one of a first architectural register and a second architectural register, and includes: a first type of physical register configured to be mapped by the first architectural register and configured to store a first type of data; the first type of physical register is configured to be mapped by the first architecture register and configured to store data of a first type, the second type of physical register is configured to be mapped by the second architecture register and configured to store data of a second type, the first architecture register comprises a vector architecture register, the data of the first type stored by the first type of physical register comprises vector data, and the second architecture register comprises a mask architecture register; the second type of data stored in the second type of physical register comprises mask data. The processor core may mitigate processing pipeline stagnation and increase less area.
Owner:HYGON INFORMATION TECH CO LTD

Data processing method and device, electronic equipment and storage medium

Embodiments of the invention provide a data processing method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a unit vector operation width supported by hardware; obtaining a source address to be aligned and the number of object data to be operated; in response to the fact that the width of the object data is larger than or equal to a first threshold value, the remainder obtained after the to-be-aligned source address is aligned with the unit vector operation width is calculated so as to obtain the data width of which the addresses are not aligned in the object data, and the width of the object data is equal to the product of the unit data width of the object data and the number of the object data; the width of the address misalignment data is equal to the difference value between the unit vector operation width and the remainder; determining the number of the address misalignment data according to the unit data width of the object data and the width of the address misalignment data, and determining first mask information based on the number of the address misalignment data and the unit vector operation width; and according to the first mask information, loading and processing the data of which the addresses are not aligned, and the method improves the data processing performance.
Owner:HYGON INFORMATION TECH CO LTD

Method and system for in-line data conversion outside of a machine learning hardware

A system includes a component configured to send data in a first data format. The system includes a direct memory access (DMA) engine configured to receive the data in the first data format and convert the first data format to a second data format, wherein the second data format is associated with a data format of a machine learning (ML) hardware, wherein the second data format is different from the first data format. The ML hardware is configured to receive the data in the second format and perform at least one ML operation on the received data in the second format. The received data in the second data format is stored on an on-chip memory (OCM) of the ML hardware.
Owner:MARVELL ASIA PTE LTD

Instruction execution device and method, electronic equipment and storage medium

To provide an instruction execution device configured to implement high computational throughput and energy efficiency, a method, electronic equipment, and a storage medium.SOLUTION: An instruction execution device 100 includes: a dispatching unit configured to respond to determining that at least one source register used for an instruction to be executed corresponds to at least one uniform register in a plurality of uniform registers, the instruction to be executed serving as a uniform instruction including a plurality of uniform computing operations to be dispatched to a first computing unit; and the first computing unit configured to execute uniform calculation operation to obtain a uniform calculation result, the uniform calculation result being written into at least one available uniform register in the plurality of uniform registers.SELECTED DRAWING: Figure 1
Owner:KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Integrated circuit and operating method thereof

The embodiment of the invention provides an integrated circuit and an operation method thereof. An integrated circuit may include a plurality of in-memory computing (CIM) circuits physically formed on a substrate. Each of the plurality of CIM circuits may include: an input circuit configured to receive a plurality of first data elements; a memory array coupled to the input circuit and configured to store the plurality of first data elements; a data multiplexer configured to output the plurality of first data elements through a first data path or through a second data path; and a plurality of arithmetic units coupled to the data multiplexer and configured to perform a multiply-accumulate (MAC) operation on the plurality of first data elements and the plurality of second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

In-memory computing accelerator using high-density operation circuit and low-power sense amplifier as peripheral circuit

An in-memory computing (IMC) accelerator using a high-density operation circuit and a low-power sense amplifier as a peripheral circuit includes a plurality of dynamic random-access memory (DRAM) banks each including a pair of cell arrays, a data supply logic, a memory, and a controller for IMC, a global SRAM, and a top-level controller, wherein the cell array includes a plurality of subarrays, each of the subarrays includes a DRAM array including a big array and a little array, and an arithmetic circuit configured to perform an operation, and the arithmetic circuit includes a sense amplifier configured to amplify a bit line voltage difference, and a compact multiply-accumulate (MAC)-single instruction multiple data (SIMD) unit (CMSU) for an MAC operation and an SIMD operation, so that functionality of an in-memory operation is diversified.
Owner:KOREA ADVANCED INST OF SCI & TECH

Configurable wavefront parallel processor

An apparatus comprising: at least one processing element configured to process a data flow in at least one direction of a plurality of directions; a configuration register comprising at least one setting that determines the processing of the data flow with the at least one processing element; and a shift register configured to select data of the at least one processing element from the at least one direction, and to provide at least one shifted data sample to a plurality of slices configured to perform at least one arithmetic operation with the data flow; wherein at least one slice of the plurality of slices is configured with the at least one setting of the configuration register.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Vertical and horizontal broadcast of shared operands

The present invention provides an apparatus, method, and storage medium for reducing the bandwidth of a memory fabric and increasing its efficiency in an array processor system. [Solution] The array processor 300 includes processor element arrays 311 to 384 distributed in rows and columns and performing operations on parameter values; a memory interface that broadcasts a set of parameter values ​​to mutually exclusive subsets of rows and columns of the processor element arrays; SIMD units 310 to 380 that include a subset of the processor element array for the corresponding row; and DMA engines 301 to 304 interconnected with mutually exclusive subsets of the processor element arrays 311 to 384.
Owner:ADVANCED MICRO DEVICES INC

Memory circuits and methods for encoder / decoder dual mode for compute-in-memory

An integrated circuit may comprise a plurality of compute-in-memory (CIM) circuits physically formed on a substrate. Each of the plurality of CIM circuits may comprise: an input circuit configured to receive a plurality of first data elements; a memory array coupled to the input circuit and configured to store the plurality of first data elements; a data multiplexer configured to output the plurality of first data elements through a first data path or through a second data path; and a plurality of computing cells coupled to the data multiplexer and configured to perform multiply-accumulate (MAC) operations on the plurality of first data elements and a plurality of second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Generation and use of memory access instruction order encodings

Apparatus and methods are disclosed for controlling execution of memory access instructions in a block-based processor architecture using a hardware structure that indicates a relative ordering of memory access instruction in an instruction block. In one example of the disclosed technology, a method of executing an instruction block having a plurality of memory load and / or memory store instructions includes selecting a next memory load or memory store instruction to execute based on dependencies encoded within the block, and on a store vector that stores data indicating which memory load and memory store instructions in the instruction block have executed. The store vector can be masked using a store mask. The store mask can be generated when decoding the instruction block, or copied from an instruction block header. Based on the encoded dependencies and the masked store vector, the next instruction can issue when its dependencies are available.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Processing data using accelerators in a system on a chip

In various examples, systems and methods are disclosed that relate to processing data using accelerators in a system on a chip. For example, a plurality of processing elements (PEs) can interconnect to form a processing engine, and a control system can control operation of the PEs based at least on the connections between the PEs. In some examples, the PEs can receive sub-inputs and transfer the sub-inputs to one or more other PEs to enable performance of the instructed operations. In examples, once the PEs complete the instructed operations, the sub-inputs can be transferred out of the processing engine.
Owner:NVIDIA CORP

Streaming engine for machine learning architecture

A programmable hardware system for machine learning (ML) includes a core and a streaming engine. The core receives a plurality of commands and a plurality of data from a host to be analyzed and inferred via machine learning. The core transmits a first subset of commands of the plurality of commands that is performance-critical operations and associated data thereof of the plurality of data for efficient processing thereof. The first subset of commands and the associated data are passed through via a function call. The streaming engine is coupled to the core and receives the first subset of commands and the associated data from the core. The streaming engine streams a second subset of commands of the first subset of commands and its associated data to an inference engine by executing a single instruction.
Owner:MARVELL ASIA PTE LTD

Architecture for block sparse operations on a systolic array

Embodiments described herein include, software, firmware, and hardware logic that provides techniques to perform arithmetic on sparse data via a systolic processing unit. One embodiment provides an accelerator device comprising a memory and a compute cluster including multiple processing resources coupled with the memory. The multiple processing resources are coupled with a data crossbar to facilitate data exchange between the multiple processing resources. Respective processing resources of the multiple processing resources include a matrix accelerator configured to: perform a dot product operation on elements of a sparse first matrix and a second matrix in response to a sparse dot product instruction, elements of the sparse first matrix are compacted into a compressed representation including a non-zero value element and an indication of the non-zero value element, and the sparse dot product instruction is to cause the matrix accelerator to skip computations associated with input including a zero value element; and write output of the dot product operation to the memory.
Owner:INTEL CORP

Flexibly programmable processing unit

A digital data processing unit comprises: inputs for receiving input data; outputs for providing output data; one or more processing elements, and one or more switches. Each processing element have one or more inputs for receiving input data and one or more outputs for providing output data. Each processing element is configured to apply, during a time slot, a mathematical operation to its input data so as to generate output data. Each switch has outputs and at least one input for receiving a respective input value, wherein each switch is operated during a time slot based on control data to provide each received input value to one of its outputs selected based on the control data; wherein the one or more processing elements and the one or more switches are interconnected to form data paths between the inputs of the processing unit and the outputs of the processing unit, wherein at least one of the data paths is configured to provide, at a corresponding output of the processing unit, a final mathematical result from intermediate mathematical results generated by the one or more processing elements in the considered data path.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Depthwise-convolution implementation on a neural processing core

A core of neural processing units is configured to efficiently process a depthwise convolution by maximizing spatial feature-map locality using adder trees. Data paths of activations and weights are inverted, and 2-to-1 multiplexers are every 2 / 9 multipliers along a row of multipliers. During a depthwise convolution operation, the core is operated using a RSxHW dataflow to maximize the locality of feature maps. For a normal convolution operation, the data paths of activations and weights may be configured for a normal convolution configuration and in which multiplexers are idle.
Owner:SAMSUNG ELECTRONICS CO LTD

Computing devices with transpose operations

The present specification discloses a computing device designed for efficiently transposing an N×N matrix using a spatial architecture of processing elements (PEs). The device can include a plurality of PEs arranged, literally or representatively, in a two-dimensional grid, each PE configured to store a single element of the matrix. A controller initializes the matrix across the PEs and sequentially increases the sub-matrix size, performing boundary calculations and data swaps within the PEs. The controller utilizes hardware primitives to facilitate parallel processing and lower compute cycles. The device adapts to various matrix shapes by padding them to the nearest N×N configuration, optimizing data distribution across the PEs.
Owner:AT-MEMORY COMPUTING LP

MEMORY CIRCUITS AND METHOD FOR ENCODER / DECODER DUAL MODE FOR COMPUTE-IN-MEMORY

An integrated circuit can have multiple compute-in-memory (CIM) circuits physically implemented on a substrate. Each of the multiple CIM circuits can have: an input circuit configured to receive multiple first data elements; a memory array coupled to the input circuit and configured to store the multiple first data elements; a data multiplexer configured to output the multiple first data elements via a first data path or a second data path; and multiple arithmetic cells coupled to the data multiplexer and configured to perform multiplication-accumulation (MAC) operations on the multiple first data elements and multiple second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Calculation device and data moving method

A calculation device multiple processors. The multiple processors are represented by a coordinate system, which includes two dimensions indicating X direction and Y direction and two or more different dimensions indicating different directions. Each processor is configured to perform data input or data output with the processor adjacent in the X direction or the Y direction, and is further configured to perform data input or data output with the processor adjacent in the different dimension.
Owner:DENSO CORP

Calculation device and data transfer method

An accelerator (10) includes multiple PEs (12). The multiple PEs (12) are represented by a coordinate system, which includes two dimensions of X direction and Y direction and two or more different dimensions indicating different directions. Each PE (12) is capable of performing data input or data output with another PE (12) adjacent in the X direction or the Y direction, and is also capable of performing data input or data output with another PE (12) adjacent in different dimension.
Owner:DENSO CORP

Multiple contexts for a memory unit in a reconfigurable data processor

A system includes a coarse-grained reconfigurable (CGR) processor and a compiler configured to generate one or more configuration files for an application for execution on the CGR processor including an array of pattern compute units (PCUs) and pattern memory units (PMUs). A PCU is configured to perform an operation. A PMU comprises a plurality of data structures including a plurality of portions of operation-specific data related to the operation. The PMU is coupled to the PCU via a multi-segment datapath pipeline. The CGR processor is coupled to configure a segment of the datapath pipeline using a set of configurations bits corresponding to a portion of the operation-specific data related to the operation to activate to the segment, to further communicate the operation-specific data to the PCU via the activated segment. The CGR processor is coupled to switch among multiple PMU contexts in various segments sequentially to concurrently.
Owner:SAMBANOVA SYSTEMS INC

Quantum block cipher resisting coprocessor for computer monitoring system and operation method

The invention discloses an anti-quantum block cipher coprocessor operation method for a computer monitoring system, and the method comprises the following steps: S1, loading data to generate a random number and a secret key, and splicing the random number and the secret key with a preset initialization vector to generate an input state; s2, receiving the associated parameters and the plaintext, and performing preprocessing to generate grouping mask shares; s3, performing parallel encryption on the grouped mask shares based on an input state to generate a plurality of mask ciphertexts and mask authentication tags; and S4, performing linear reconstruction on the mask ciphertext to finally run to obtain a complete ciphertext, decomposing plaintext data into a plurality of mask shares for parallel processing, effectively diluting data structure features, and fundamentally cutting off a way of acquiring key information by an attacker through power consumption analysis. Compared with the prior art, the parallel computing advantage of reconfigurable hardware is fully utilized, efficient execution of cryptographic operation is achieved, the whole data processing process is completed in the register, and the influence of memory access delay on system performance is remarkably reduced.
Owner:HANGZHOU HUADIAN BANSHAN POWER GENERATION

Processor system for performing neural network computations

A processor system (1) is disclosed herein for performing neural network computations. The processor system (1) comprises a plurality of synchronously operating processor clusters arranged in a two-dimensional array, wherein a first plurality of processor clusters in a first column of the array are logically coupled to a second plurality of processor clusters in a second, directly adjacent column by at least one shared external bus. Each processor cluster comprises at least one processor unit including a programmable processor and an instruction memory, and further comprises a data memory. The processor system is configured to allow a first programmable processor of one of the first plurality of processor clusters to write a value to the at least one shared external bus and to allow a plurality of other programmable processors comprised in the second plurality of processor clusters to concurrently read the value from the at least one external bus.
Owner:EUCLYD BV

Data processing method, apparatus and medium applied to a distributed system

The present disclosure provides a data processing method, an apparatus and a medium applied to a distributed system, relates to the field of computer technology, and particularly to the field of chip technology and distributed data processing technology. The implementation is: for each of a plurality of data blocks included in data to be processed, dividing the data block into a plurality of first sub-blocks corresponding to a plurality of computing unit groups respectively; for each first sub-block of the plurality of first sub-blocks, dividing the first sub-block, based on a target computing unit group corresponding to the first sub-block, into a plurality of second sub-blocks corresponding to the plurality of computing units in the target computing unit group, respectively; determining, by processing each second sub-block utilizing the corresponding computing unit of the second sub-block , a plurality of first processing results output by the plurality of computing units respectively; and determining a processing result of the first sub-block by performing a data reduction operation on the plurality of first processing results utilizing the target computing unit group.
Owner:KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD

Method, accelerator, and electronic device with tensor processing

A processor-implemented tensor processing method includes: receiving a request to process a neural network including a normalization layer by an accelerator; and generating an instruction executable by the accelerator in response to the request, wherein, by executing the instruction, the accelerator is configured to determine an intermediate tensor corresponding to a result of performing a portion of operations included in the normalization layer, by performing, in a channel axis direction, a convolution based on: a target tensor on which the portion of operations is to be performed; and a kernel having a number of input channels and a number of output channels determined based on the target tensor and including elements of scaling values determined based on the target tensor.
Owner:SAMSUNG ELECTRONICS CO LTD

Interconnect device, operation method of interconnect device, and artificial intelligence (AI) accelerator system including interconnect device

An interconnect device may include one or more hardware-implemented modules configured to: receive a command from a processing core; perform, based on the received command, an operation including either one or both of an accumulation operation on sets of data stored in a memory and an aggregation operation on results processed by the processing core; and provide a result of the performing of the operation.
Owner:SAMSUNG ELECTRONICS CO LTD