Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

117results about "Single instruction multiple data multiprocessors" patented technology

Data processing method and device, electronic equipment and storage medium

Embodiments of the invention provide a data processing method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a unit vector operation width supported by hardware; obtaining a source address to be aligned and the number of object data to be operated; in response to the fact that the width of the object data is larger than or equal to a first threshold value, the remainder obtained after the to-be-aligned source address is aligned with the unit vector operation width is calculated so as to obtain the data width of which the addresses are not aligned in the object data, and the width of the object data is equal to the product of the unit data width of the object data and the number of the object data; the width of the address misalignment data is equal to the difference value between the unit vector operation width and the remainder; determining the number of the address misalignment data according to the unit data width of the object data and the width of the address misalignment data, and determining first mask information based on the number of the address misalignment data and the unit vector operation width; and according to the first mask information, loading and processing the data of which the addresses are not aligned, and the method improves the data processing performance.
Owner:HYGON INFORMATION TECH CO LTD

Method and system for in-line data conversion outside of a machine learning hardware

A system includes a component configured to send data in a first data format. The system includes a direct memory access (DMA) engine configured to receive the data in the first data format and convert the first data format to a second data format, wherein the second data format is associated with a data format of a machine learning (ML) hardware, wherein the second data format is different from the first data format. The ML hardware is configured to receive the data in the second format and perform at least one ML operation on the received data in the second format. The received data in the second data format is stored on an on-chip memory (OCM) of the ML hardware.
Owner:MARVELL ASIA PTE LTD

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Integrated circuit and operating method thereof

The embodiment of the invention provides an integrated circuit and an operation method thereof. An integrated circuit may include a plurality of in-memory computing (CIM) circuits physically formed on a substrate. Each of the plurality of CIM circuits may include: an input circuit configured to receive a plurality of first data elements; a memory array coupled to the input circuit and configured to store the plurality of first data elements; a data multiplexer configured to output the plurality of first data elements through a first data path or through a second data path; and a plurality of arithmetic units coupled to the data multiplexer and configured to perform a multiply-accumulate (MAC) operation on the plurality of first data elements and the plurality of second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

In-memory computing accelerator using high-density operation circuit and low-power sense amplifier as peripheral circuit

An in-memory computing (IMC) accelerator using a high-density operation circuit and a low-power sense amplifier as a peripheral circuit includes a plurality of dynamic random-access memory (DRAM) banks each including a pair of cell arrays, a data supply logic, a memory, and a controller for IMC, a global SRAM, and a top-level controller, wherein the cell array includes a plurality of subarrays, each of the subarrays includes a DRAM array including a big array and a little array, and an arithmetic circuit configured to perform an operation, and the arithmetic circuit includes a sense amplifier configured to amplify a bit line voltage difference, and a compact multiply-accumulate (MAC)-single instruction multiple data (SIMD) unit (CMSU) for an MAC operation and an SIMD operation, so that functionality of an in-memory operation is diversified.
Owner:KOREA ADVANCED INST OF SCI & TECH

Configurable wavefront parallel processor

An apparatus comprising: at least one processing element configured to process a data flow in at least one direction of a plurality of directions; a configuration register comprising at least one setting that determines the processing of the data flow with the at least one processing element; and a shift register configured to select data of the at least one processing element from the at least one direction, and to provide at least one shifted data sample to a plurality of slices configured to perform at least one arithmetic operation with the data flow; wherein at least one slice of the plurality of slices is configured with the at least one setting of the configuration register.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Vertical and horizontal broadcast of shared operands

The present invention provides an apparatus, method, and storage medium for reducing the bandwidth of a memory fabric and increasing its efficiency in an array processor system. [Solution] The array processor 300 includes processor element arrays 311 to 384 distributed in rows and columns and performing operations on parameter values; a memory interface that broadcasts a set of parameter values ​​to mutually exclusive subsets of rows and columns of the processor element arrays; SIMD units 310 to 380 that include a subset of the processor element array for the corresponding row; and DMA engines 301 to 304 interconnected with mutually exclusive subsets of the processor element arrays 311 to 384.
Owner:ADVANCED MICRO DEVICES INC

Memory circuits and methods for encoder / decoder dual mode for compute-in-memory

An integrated circuit may comprise a plurality of compute-in-memory (CIM) circuits physically formed on a substrate. Each of the plurality of CIM circuits may comprise: an input circuit configured to receive a plurality of first data elements; a memory array coupled to the input circuit and configured to store the plurality of first data elements; a data multiplexer configured to output the plurality of first data elements through a first data path or through a second data path; and a plurality of computing cells coupled to the data multiplexer and configured to perform multiply-accumulate (MAC) operations on the plurality of first data elements and a plurality of second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Generation and use of memory access instruction order encodings

Apparatus and methods are disclosed for controlling execution of memory access instructions in a block-based processor architecture using a hardware structure that indicates a relative ordering of memory access instruction in an instruction block. In one example of the disclosed technology, a method of executing an instruction block having a plurality of memory load and / or memory store instructions includes selecting a next memory load or memory store instruction to execute based on dependencies encoded within the block, and on a store vector that stores data indicating which memory load and memory store instructions in the instruction block have executed. The store vector can be masked using a store mask. The store mask can be generated when decoding the instruction block, or copied from an instruction block header. Based on the encoded dependencies and the masked store vector, the next instruction can issue when its dependencies are available.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Processing data using accelerators in a system on a chip

In various examples, systems and methods are disclosed that relate to processing data using accelerators in a system on a chip. For example, a plurality of processing elements (PEs) can interconnect to form a processing engine, and a control system can control operation of the PEs based at least on the connections between the PEs. In some examples, the PEs can receive sub-inputs and transfer the sub-inputs to one or more other PEs to enable performance of the instructed operations. In examples, once the PEs complete the instructed operations, the sub-inputs can be transferred out of the processing engine.
Owner:NVIDIA CORP

Flexibly programmable processing unit

A digital data processing unit comprises: inputs for receiving input data; outputs for providing output data; one or more processing elements, and one or more switches. Each processing element have one or more inputs for receiving input data and one or more outputs for providing output data. Each processing element is configured to apply, during a time slot, a mathematical operation to its input data so as to generate output data. Each switch has outputs and at least one input for receiving a respective input value, wherein each switch is operated during a time slot based on control data to provide each received input value to one of its outputs selected based on the control data; wherein the one or more processing elements and the one or more switches are interconnected to form data paths between the inputs of the processing unit and the outputs of the processing unit, wherein at least one of the data paths is configured to provide, at a corresponding output of the processing unit, a final mathematical result from intermediate mathematical results generated by the one or more processing elements in the considered data path.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Depthwise-convolution implementation on a neural processing core

A core of neural processing units is configured to efficiently process a depthwise convolution by maximizing spatial feature-map locality using adder trees. Data paths of activations and weights are inverted, and 2-to-1 multiplexers are every 2 / 9 multipliers along a row of multipliers. During a depthwise convolution operation, the core is operated using a RSxHW dataflow to maximize the locality of feature maps. For a normal convolution operation, the data paths of activations and weights may be configured for a normal convolution configuration and in which multiplexers are idle.
Owner:SAMSUNG ELECTRONICS CO LTD

Computing devices with transpose operations

The present specification discloses a computing device designed for efficiently transposing an N×N matrix using a spatial architecture of processing elements (PEs). The device can include a plurality of PEs arranged, literally or representatively, in a two-dimensional grid, each PE configured to store a single element of the matrix. A controller initializes the matrix across the PEs and sequentially increases the sub-matrix size, performing boundary calculations and data swaps within the PEs. The controller utilizes hardware primitives to facilitate parallel processing and lower compute cycles. The device adapts to various matrix shapes by padding them to the nearest N×N configuration, optimizing data distribution across the PEs.
Owner:AT-MEMORY COMPUTING LP

MEMORY CIRCUITS AND METHOD FOR ENCODER / DECODER DUAL MODE FOR COMPUTE-IN-MEMORY

An integrated circuit can have multiple compute-in-memory (CIM) circuits physically implemented on a substrate. Each of the multiple CIM circuits can have: an input circuit configured to receive multiple first data elements; a memory array coupled to the input circuit and configured to store the multiple first data elements; a data multiplexer configured to output the multiple first data elements via a first data path or a second data path; and multiple arithmetic cells coupled to the data multiplexer and configured to perform multiplication-accumulation (MAC) operations on the multiple first data elements and multiple second data elements.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Calculation device and data moving method

A calculation device multiple processors. The multiple processors are represented by a coordinate system, which includes two dimensions indicating X direction and Y direction and two or more different dimensions indicating different directions. Each processor is configured to perform data input or data output with the processor adjacent in the X direction or the Y direction, and is further configured to perform data input or data output with the processor adjacent in the different dimension.
Owner:DENSO CORP

Calculation device and data transfer method

An accelerator (10) includes multiple PEs (12). The multiple PEs (12) are represented by a coordinate system, which includes two dimensions of X direction and Y direction and two or more different dimensions indicating different directions. Each PE (12) is capable of performing data input or data output with another PE (12) adjacent in the X direction or the Y direction, and is also capable of performing data input or data output with another PE (12) adjacent in different dimension.
Owner:DENSO CORP

Multiple contexts for a memory unit in a reconfigurable data processor

A system includes a coarse-grained reconfigurable (CGR) processor and a compiler configured to generate one or more configuration files for an application for execution on the CGR processor including an array of pattern compute units (PCUs) and pattern memory units (PMUs). A PCU is configured to perform an operation. A PMU comprises a plurality of data structures including a plurality of portions of operation-specific data related to the operation. The PMU is coupled to the PCU via a multi-segment datapath pipeline. The CGR processor is coupled to configure a segment of the datapath pipeline using a set of configurations bits corresponding to a portion of the operation-specific data related to the operation to activate to the segment, to further communicate the operation-specific data to the PCU via the activated segment. The CGR processor is coupled to switch among multiple PMU contexts in various segments sequentially to concurrently.
Owner:SAMBANOVA SYSTEMS INC

Quantum block cipher resisting coprocessor for computer monitoring system and operation method

The invention discloses an anti-quantum block cipher coprocessor operation method for a computer monitoring system, and the method comprises the following steps: S1, loading data to generate a random number and a secret key, and splicing the random number and the secret key with a preset initialization vector to generate an input state; s2, receiving the associated parameters and the plaintext, and performing preprocessing to generate grouping mask shares; s3, performing parallel encryption on the grouped mask shares based on an input state to generate a plurality of mask ciphertexts and mask authentication tags; and S4, performing linear reconstruction on the mask ciphertext to finally run to obtain a complete ciphertext, decomposing plaintext data into a plurality of mask shares for parallel processing, effectively diluting data structure features, and fundamentally cutting off a way of acquiring key information by an attacker through power consumption analysis. Compared with the prior art, the parallel computing advantage of reconfigurable hardware is fully utilized, efficient execution of cryptographic operation is achieved, the whole data processing process is completed in the register, and the influence of memory access delay on system performance is remarkably reduced.
Owner:HANGZHOU HUADIAN BANSHAN POWER GENERATION

Processor system for performing neural network computations

A processor system (1) is disclosed herein for performing neural network computations. The processor system (1) comprises a plurality of synchronously operating processor clusters arranged in a two-dimensional array, wherein a first plurality of processor clusters in a first column of the array are logically coupled to a second plurality of processor clusters in a second, directly adjacent column by at least one shared external bus. Each processor cluster comprises at least one processor unit including a programmable processor and an instruction memory, and further comprises a data memory. The processor system is configured to allow a first programmable processor of one of the first plurality of processor clusters to write a value to the at least one shared external bus and to allow a plurality of other programmable processors comprised in the second plurality of processor clusters to concurrently read the value from the at least one external bus.
Owner:EUCLYD BV

Data processing method, apparatus and medium applied to a distributed system

The present disclosure provides a data processing method, an apparatus and a medium applied to a distributed system, relates to the field of computer technology, and particularly to the field of chip technology and distributed data processing technology. The implementation is: for each of a plurality of data blocks included in data to be processed, dividing the data block into a plurality of first sub-blocks corresponding to a plurality of computing unit groups respectively; for each first sub-block of the plurality of first sub-blocks, dividing the first sub-block, based on a target computing unit group corresponding to the first sub-block, into a plurality of second sub-blocks corresponding to the plurality of computing units in the target computing unit group, respectively; determining, by processing each second sub-block utilizing the corresponding computing unit of the second sub-block , a plurality of first processing results output by the plurality of computing units respectively; and determining a processing result of the first sub-block by performing a data reduction operation on the plurality of first processing results utilizing the target computing unit group.
Owner:KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD

Method, accelerator, and electronic device with tensor processing

A processor-implemented tensor processing method includes: receiving a request to process a neural network including a normalization layer by an accelerator; and generating an instruction executable by the accelerator in response to the request, wherein, by executing the instruction, the accelerator is configured to determine an intermediate tensor corresponding to a result of performing a portion of operations included in the normalization layer, by performing, in a channel axis direction, a convolution based on: a target tensor on which the portion of operations is to be performed; and a kernel having a number of input channels and a number of output channels determined based on the target tensor and including elements of scaling values determined based on the target tensor.
Owner:SAMSUNG ELECTRONICS CO LTD

Interconnect device, operation method of interconnect device, and artificial intelligence (AI) accelerator system including interconnect device

An interconnect device may include one or more hardware-implemented modules configured to: receive a command from a processing core; perform, based on the received command, an operation including either one or both of an accumulation operation on sets of data stored in a memory and an aggregation operation on results processed by the processing core; and provide a result of the performing of the operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Apparatus, method, and computer program for collecting diagnostic information

The device has a processing circuit for performing Single Instruction Multiple Data (SIMD) processing on an array having multiple data items, and the processing circuit supports SIMD processing for multiple array sizes. The processing circuit can select the array size for which to perform SIMD processing based on SIMD processing configuration information. A diagnostic information acquisition circuit is provided to collect diagnostic information about the software running on the processing circuit, and the diagnostic information acquisition circuit filters the collection of diagnostic information based on the SIMD processing configuration information.
Owner:ARM LTD

Rearranging data among processing elements of computational memory

An array of interconnected processing elements is modelled as a graph of nodes. Each layer of the graph represents a possible arrangement of data elements within the array of interconnected processing elements. An edge between nodes of adjacent layers of the graph represents a movement of a data element between the nodes. Constraints are set for a starting arrangement of the data elements stored in the array, an ending arrangement of the data elements stored in the array, and a limit for each node of the graph to have one input edge from a previous layer and one output edge to a subsequent layer. The model and constraints are processed with an integer programming solver to obtain a program of movements of data elements among the interconnected processing elements. The program implements a rearrangement of the data elements from the starting arrangement to the ending arrangement.
Owner:AT-MEMORY COMPUTING LP

Gather accelerated address space

In a system including a processing unit and a set of one or more stacked memory chips, the processing unit can request data. When the data is distributed such that there is at least one non-contiguous memory sector in the smallest unit of memory segments usable by the system, then a gather operation can be utilized to instruct the set of one or more stacked memory chips to gather the requested data into a virtual address space, e.g., a gather accelerated address space. The requested data can be aligned to the byte chunk size used by the processing unit and at least some of the unneeded memory segments can be skipped, e.g., not copied into the virtual address space. The requested data in the virtual address space can be communicated to the processing unit using less bandwidth resources than when not using the gather operation.
Owner:NVIDIA CORP

Vertical and horizontal broadcast of shared operands

The array processor includes an array of processor elements distributed in rows and columns. The processor element array performs operations on parameter values. The array processor also includes a memory interface that broadcasts a set of parameter values ​​to mutually exclusive subsets of the rows and columns of the processor element array. In some cases, the array processor includes a single instruction, multiple data (SIMD) unit that includes a subset of the processor element array of a corresponding row, a workgroup processor (WGP) that includes a subset of the SIMD units, and a memory fabric configured to interconnect with an external memory that stores the parameter values. The memory interface broadcasts the parameter values ​​to the SIMD unit that includes the row of processor element array associated with the memory interface and the column of processor element array implemented across the SIMD units in the WGP. The memory interface accesses the parameter values ​​from the external memory via the memory fabric.
Owner:ADVANCED MICRO DEVICES INC

Calculation device and data transfer method

An accelerator (10) includes multiple PEs (12). The multiple PEs (12) are represented by a coordinate system, which includes two dimensions of X direction and Y direction and two or more different dimensions indicating different directions. Each PE (12) is capable of performing data input or data output with another PE (12) adjacent in the X direction or the Y direction, and is also capable of performing data input or data output with another PE (12) adjacent in different dimension.
Owner:DENSO CORP