Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

246 results about "Register file" patented technology

A register file is an array of processor registers in a central processing unit (CPU). Modern integrated circuit-based register files are usually implemented by way of fast static RAMs with multiple ports. Such RAMs are distinguished by having dedicated read and write ports, whereas ordinary multiported SRAMs will usually read and write through the same ports.

Information processing method and device, electronic equipment and storage medium

The invention discloses an information processing method and device, electronic equipment and a storage medium, and is applied to the field of information processing. The information processing method comprises the steps that a first utilization rate is determined in combination with a data flow between a memory and a shared memory when a processor executes a general matrix multiplication operator, and the first utilization rate is used for indicating performance evaluation of the general matrix multiplication operator at a processor level; determining a second utilization rate in combination with a data stream between a shared memory and a register file in a single computing unit when the processor executes the general matrix multiplication operator, the second utilization rate being used for indicating performance evaluation of the general matrix multiplication operator at the computing unit level; and based on the smaller one of the first utilization rate and the second utilization rate, determining the theoretical performance of the processor when executing the universal matrix multiplication operator. At present, only chip-level performance evaluation is considered in performance evaluation of a GEMM operator, and double-level-dimension theoretical performance evaluation is provided, so that the evaluation result is more accurate.
Owner:SHANGHAI BIREN TECH CO LTD

DCU-based high-performance sparse stiffness matrix vector multiplication method

The invention provides a DCU-based high-performance sparse stiffness matrix vector multiplication method, which comprises the following steps of: according to a sparse stiffness matrix, dividing a non-zero element into a plurality of calculation unit blocks by rows, pre-loading non-zero element data to an L1 shared memory or a register file through an on-chip shared memory controller of the DCU, a high-bandwidth crossbar switch of the DCU is used for realizing data copying and transmission; starting multi-row fusion execution for a short row of which the row non-zero element is lower than a DCU single-instruction multi-data width threshold value; constructing a wavefront scheduler based on a DCU asynchronous computing engine: binding an independent instruction cache region for each wavefront, and loading a multiply-add operation instruction set in advance through a prefetch instruction queue; a calculation unit state register is established, and when a wavefront scheduler ready signal is triggered, a scalar unit of the DCU is activated to execute calculation; and realizing cross-thread block reduction by adopting a DCU atomic operation accelerator. According to the method, the calculation throughput and the memory bandwidth utilization rate of large-scale structural mechanics stiffness matrix vector multiplication are effectively improved.
Owner:HENAN POLYTECHNIC

Matrix multiplication and accumulation operation unit and operation method, hardware accelerator and electronic equipment

The embodiment of the invention provides a matrix multiplication and accumulation operation unit and method, a hardware accelerator and electronic equipment, and the matrix multiplication and accumulation operation unit comprises a data loading storage engine, a tensor register file and a matrix multiplication engine. The data loading and storage engine is used for loading data of a plurality of matrixes to be subjected to matrix multiplication and accumulation calculation; the tensor register file is used for storing data of a plurality of matrixes acquired from the data loading storage engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprises a plurality of tensor registers, and different tensor register groups are used for storing data of different matrixes in the plurality of matrixes; and the matrix multiplication engine is used for carrying out matrix multiplication accumulation calculation based on the data of the plurality of matrixes stored in the tensor register file. According to the embodiment of the invention, more efficient MMA calculation is realized under the conditions of low cost, low power consumption and less occupied space.
Owner:ALIBABA (CHINA) CO LTD

Low latency scratch memory path

An apparatus and method for efficiently processing vector memory accesses on an integrated circuit. In various implementations, a computing system includes a processing circuit with multiple compute circuits for executing wavefronts of a parallel data application. Each compute circuit includes a local memory subsystem for accessing data not found in vector register files of the compute circuit. The local memory subsystem includes a first execution pipeline and a second execution pipeline. The second execution pipeline processes vector stack access instructions that access temporary data such as stack data of a function call used by each wavefront that is generated based on the function call. The first execution pipeline processes other types of vector memory access instructions and includes multiple complex pipeline stages not found in the second execution pipeline. Thus, the second execution pipeline has a latency less than the latency of the first execution pipeline.
Owner:ADVANCED MICRO DEVICES INC

Matrix multiply accumulation operation unit and operation method, and hardware accelerator and electronic device

Provided in the embodiments of the present disclosure are a matrix multiply accumulation (MMA) operation unit and operation method, and a hardware accelerator and an electronic device. The MMA operation unit comprises a data load / store engine, a tensor register file and a matrix multiplication engine, wherein the data load / store engine is used for loading data of a plurality of matrixes to be subjected to MMA computation; the tensor register file is used for storing the data of the plurality of matrixes that is acquired from the data load / store engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprising a plurality of tensor registers, and different tensor register groups being used for storing data of different matrixes among the plurality of matrixes; and the matrix multiplication engine is used for performing MMA computation on the basis of the data of the plurality of matrixes that is stored in the tensor register file. By means of the embodiments of the present disclosure, more efficient MMA computation is realized with low cost, low power consumption and a smaller footprint.
Owner:ALIBABA (CHINA) CO LTD

Vector mask buffers in a vector instruction execution pipeline

Systems and methods related to vector mask buffers in a vector instruction execution pipeline are disclosed herein. The vector instruction execution pipeline may include several lanes. Each lane may include a vector register file, a vector mask buffer, and a functional processing unit. The vector register file may store operand data and the vector mask buffer may store a vector mask associated with the operand data. In a lane, the operand data may be read from the register file into a functional processing unit, and the vector mask may be read from the vector mask buffer to the functional processing unit. The functional processing unit may process the operand data based on the vector mask. The lane-specific vector mask buffers improve the efficiency of the vector instruction execution pipeline by storing the vector masks proximate to where the vector masks will be used.
Owner:TENSTORRENT USA INC

Matrix multiplication optimization method and system based on RISC-V architecture

The invention discloses a matrix multiplication optimization method and system based on an RISC-V architecture, belongs to the technical field of machine learning, and aims to solve the technical problem of how to effectively improve the performance of sparse matrix multiplication. Comprising the following steps: constructing a sparse matrix adaptive to a machine learning model; the sparse matrix is converted into a structured sparse matrix, non-zero elements in the structured sparse matrix are stored in a vector register file, and column indexes corresponding to the non-zero elements are stored in a scalar register file; determining expansion factors of the internal circulation and the external circulation, and adjusting the expansion factors of the internal circulation; executing instructions in different loop iterations in a staggered manner in a staggered manner; designing and realizing a user-defined vector index multiply-add instruction; a custom vector index multiply-add instruction is executed.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Method for calculating matrix multiplication, artificial intelligence chip, calculation device, medium and program product

The invention relates to a method for calculating matrix multiplication, an artificial intelligence chip, a calculation device, a medium and a program product. The method comprises the following steps of: accumulating a calculation result of a current cycle calculation and a calculation result of a previous cycle calculation of matrix multiplication executed in a thread bundle group granularity by utilizing a buffer which is configured in a calculation core and is used for accumulation operation; determining whether the last cycle calculation of the matrix multiplication performed at the thread bundle group granularity is completed; and in response to determining that the last cycle computation of the matrix multiplication performed at the thread bundle group granularity is completed, writing the computation results accumulated via the buffer to the register file. According to the invention, the write bandwidth of the register and the occupation of the register space can be obviously reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Efficient matrix engine architecture based on RISC-V matrix extension and calculation method

The invention provides a high-efficiency matrix engine (RVME) architecture based on RISC-V matrix extension and a calculation method, and the architecture comprises an instruction buffering and decoding module, a matrix loading / storage module, a matrix register file, a parallel outer product array and an element-by-element operation module; the matrix register file comprises a Tile register and an Acculator register; the storage modules are respectively used for storing an input matrix and an accumulation result and supporting efficient data access and parallel computing; the matrix loading / storage module significantly improves the data loading efficiency through cache line alignment and matrix transposition optimization; the instruction buffering and decoding module cooperates with a main processor through a reordering buffer area and an instruction buffer area to ensure efficient scheduling and execution of instructions. The parallel outer product array is adopted to replace a traditional systolic array, the idle period in the calculation process is eliminated through multicast data flow scheduling and a ping-pong buffer read-write mechanism, and matrix multiplication and addition operation with high calculation utilization rate and low delay is achieved.
Owner:SHANGHAI JIAOTONG UNIV

STORING FLOATING-POINT VALUES ACCORDING TO AN EXTENDED QFLOAT FLOATING-POINT (xqFP) FORMAT IN PROCESSOR DEVICES

Storing floating-point values according to an extended QFloat floating-point (xqFP) format in processor devices is disclosed herein. In some aspects, a processor device comprises a register file comprising a plurality of registers, and comprises a floating-point unit (FPU) circuit that is configured to store a first floating-point value in a register of the plurality of registers. The first floating-point value is formatted according to the xqFP format that comprises an exponent field and a significand field. The significand field is formatted as a signed one's complement value, and comprises a sign bit, an explicit most-significant-bit (MSB), a fractional field, and a deferred increment bit that represents a value of one-half (½) unit of least precision (ULP).
Owner:QUALCOMM INC

In-memory computing accelerator, system and method based on adaptive utilization

The invention discloses an in-memory computing accelerator, system and method based on a self-adaptive utilization rate. The in-memory computing accelerator based on the self-adaptive utilization rate comprises an in-memory computing processor, a global buffer area, a vector processor, a control module and a configuration register file, the control module and the configuration register file are respectively connected with the in-memory computing processor, the global buffer area and the vector processor, the in-memory computing processor is connected with the global buffer area, and the global buffer area is connected with the vector processor; the in-memory computing processor comprises a plurality of in-memory computing cores, and the plurality of in-memory computing cores form a dynamic parallel structure. According to the method and the device, at least one neural network reasoning calculation process can be executed on the input data aiming at the input data of a certain neural network, and relatively high space utilization rate and time utilization rate can be obtained when facing calculation requirements of different operators, so that efficient acceleration of a diversified neural network model is realized.
Owner:SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Memory device and method

A memory device includes a plurality of memory banks, and a processing-in-memory (PIM) block accessible to the plurality of memory banks, wherein the PIM block comprises a control circuit configured to receive a plurality of operation instructions from a host and, in response to a predicated instruction indicating a predication operation among the plurality of operation instructions, instruct an arithmetic logic unit (ALU) to perform the predication operation, a predicate register file (PRF) configured to store therein a predicate value determined by the predication operation, and the ALU configured to perform an operation according to a command signal translated by the control circuit based on the predicate value from an operation instruction that depends on the predicate value among the plurality of operation instructions.
Owner:SAMSUNG ELECTRONICS CO LTD +1

Distributed read-write method and device for large model training

The invention relates to the technical field of computer data storage and distributed systems, in particular to a distributed read-write method for large model training, which comprises the following steps of: constructing a hierarchical structure of local cache, distributed shared cache and persistent storage, and maintaining a global cache table copy at each client to register file blocks and cache positions. And during writing, firstly writing in a local cache and copying to other client nodes, and submitting persistent storage after registration is completed. During reading, the file blocks are obtained from the local cache and the distributed shared cache in sequence, if the file blocks are not hit, the file blocks are obtained from the persistent storage, meanwhile, the effective file blocks are written back to the local, and table items are supplemented. Unified visibility is achieved through position mapping and version constraint, nearby reading and fault recovery are achieved, back-end pressure and access delay are reduced, and the method is suitable for large-scale training and batch processing scenes.
Owner:CHONGQING ZHONGKE YUNCONG TECH CO LTD +1

Database concurrency control and memory access optimization method, device and equipment

The invention provides a database concurrency control and memory access optimization method which can be applied to the technical field of computer system software performance optimization. According to the method, multi-level instruction-level reconstruction is carried out on an execution path of a database kernel through the characteristics of an underlying hardware instruction set of a collaborative application processor platform, and the method comprises the following collaborative implementation optimization dimensions: based on a register file and a cache hierarchical structure of the processor platform, an access mode of a database core data structure is optimized; the number of memory access instructions is reduced; the data locality is improved; on the basis of atomic instruction set extension of a processor platform, instruction-level reconstruction is carried out on key primitives in a database multi-thread synchronization mechanism so as to reduce the overhead and contention of synchronization operation; on the basis of single-instruction multi-data-stream extension of a processor platform and a vector atomic operation instruction of the single-instruction multi-data-stream extension, parallel acceleration is carried out on batch life cycle management operation of objects in a database.
Owner:AEROSPACE INFORMATION RES INST CAS

Vector mask buffers in a vector instruction execution pipeline

Systems and methods related to vector mask buffers in a vector instruction execution pipeline are disclosed herein. The vector instruction execution pipeline may include several lanes. Each lane may include a vector register file, a vector mask buffer, and a functional processing unit. The vector register file may store operand data and the vector mask buffer may store a vector mask associated with the operand data. In a lane, the operand data may be read from the register file into a functional processing unit, and the vector mask may be read from the vector mask buffer to the functional processing unit. The functional processing unit may process the operand data based on the vector mask. The lane-specific vector mask buffers improve the efficiency of the vector instruction execution pipeline by storing the vector masks proximate to where the vector masks will be used.
Owner:TENSTORRENT USA INC

Scalar processor, high-performance processor, and electronic device

The invention provides a scalar processor, a high-performance processor and electronic equipment, and the scalar processor comprises an instruction fetching unit, a register renaming unit, an operation reservation stack unit, a storage reservation stack unit, a scalar operation unit, a memory access unit, a program control unit, a synchronization unit, a pipeline control unit and a register file unit, a special vector register file unit; wherein the synchronization unit is used for synchronizing the scalar processor and the vector processor. According to the scalar processor provided by the invention, the synchronization of the scalar processor and the vector processor is realized through the synchronization unit, so that an instruction is efficiently executed.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Cross-domain large file transmission method based on distributed soft bus

The invention provides a cross-domain large file transmission method based on a distributed soft bus, and belongs to the technical field of communication transmission. The method comprises the following steps: finishing networking of different local area networks; on a first local area network side, a large file to be transmitted is fragmented according to a preset fixed size, so that a plurality of file fragments are generated; registering file fragment information on the soft bus, wherein the file fragment information comprises the serial number, the size and the check code of the file fragment; according to file fragment information registered by the soft bus, dynamically sending file fragments through a soft bus transmission process; on a second local area network side, receiving the file fragments which arrive successively through a soft bus transmission process; the serial number and the check code of the file fragment are checked through the soft bus; and after all the file fragments are received, carrying out file recombination according to the file fragment information, and recovering the original large file. According to the invention, the efficiency, stability and safety of cross-domain transmission can be improved.
Owner:WEAPON EQUIP RES INST OF CHINA NAT WEAPON EQUIP GRP

Hardware-based implementation of secure hash algorithms

A processor includes a register file and an execution unit. The execution unit includes a hash circuit including at least a state register, a state update circuit coupled to the state register, and a control circuit. Based on a hash instruction, the hash circuit receives from the register file and buffers within the state register a current state of a message being hashed. The state update circuit performs state update function on contents of the state register, where performing the state update function includes performing a plurality of iterative rounds of processing on contents of the state register and returning a result of each of the plurality of iterative rounds of processing to the state register. Following completion of all of the plurality of iterative rounds of processing, the execution unit stores contents of the state register to the register file as an updated state of the message.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Complete K-SAT solver based on incremental update technology

The invention provides a completeness K-SAT solver based on an incremental updating technology, and the solver comprises an incremental updater, and the incremental updater comprises an assignment mask unit which is used for adjusting the current assignment of a target character based on the current working mode of the assignment mask unit, and obtaining the adjusted assignment information; the register file is used for storing historical assignment information of characters contained in each clause in the Boolean formula; one end of the updating logic unit is connected with the assignment mask unit, the other end of the updating logic unit is connected with the register file, and the updating logic unit is used for determining character assignment information based on the historical assignment information and the adjusted assignment information; and the assignment mask unit is also used for outputting the character assignment information, so that the solver carries out complete K-SAT solving based on the character assignment information. According to the method, efficient storage and management of historical assignment information are realized on the hardware level, and hardware resource consumption and solution delay caused by frequent storage and backtracking in completeness analysis are effectively reduced.
Owner:PEKING UNIV

ALU operation fusion processing module and method suitable for neural network

The invention discloses an ALU operation fusion processing module and method suitable for a neural network, and the module comprises a control unit which is used for receiving and decoding a machine instruction, managing the execution processes of internal and external circulation and microinstruction circulation, and generating a control signal of each stage of a microinstruction assembly line; the microinstruction buffer area is used for storing a microinstruction sequence pre-generated by the neural network compiler; the register file is used for storing source operands and results of ALU operation; each entry of the register file is composed of a valid bit, a tag bit and a data bit; the ALU computing core adopts an SIMD (Single Instruction Multiple Data) architecture and comprises a plurality of paths of parallel arithmetic logic function units; the Load / Store unit is used for processing data exchange between a register file and a local buffer area; and the data selection interface is used for selecting a data source or a target buffer area according to the storage tag field of the microinstruction. According to the method, the high efficiency and the flexibility of the ALU in the neural network hardware accelerator can be effectively considered.
Owner:ZHEJIANG UNIV

Register file for systolic array

A processing apparatus includes a general-purpose parallel processing engine including a set of multiple processing elements including a single precision floating-point unit, a double precision floating point unit, and an integer unit; a matrix accelerator including one or more systolic arrays; a first register file coupled with a first read control circuit, wherein the first read control circuit couples with the set of multiple processing elements and the matrix accelerator to arbitrate read requests to the first register file from the set of multiple processing elements and the matrix accelerator; and a second register file coupled with a second read control circuit, wherein the second read control circuit couples with the matrix accelerator to arbitrate read requests to the second register file from the matrix accelerator and limit access to the second register file by the set of multiple processing elements.
Owner:INTEL CORP

Stall-driven multi-processing

In a microprocessor having an instruction execution unit and first and second sets of process execution resources, context information for a process next-to-be-executed by the instruction execution unit is loaded into a register file, translation lookaside buffer and first-level data cache of the first set of process execution resources during a first interval. During the first interval and concurrently with the loading of context information for the process next-to-be executed, the instruction execution unit executes a current process, including accessing context information for the current process within the register file, translation lookaside buffer and first-level data cache of the second set of process execution resources.
Owner:RAMBUS INC

Vector memory loads return to cache

An apparatus and method for efficiently processing vector memory accesses on an integrated circuit. In various implementations, a computing system includes a parallel data processing circuit with multiple, replicated compute circuits. Each compute circuit includes a cache that stores temporary data. A control circuit that identifies an instruction of the in-fight wavefronts as a long-latency vector memory access instruction with an indication specifying that a destination targets data storage in the cache, rather than a vector register file. The control circuit assigns a cache line for the vector memory access instruction while foregoing vector register assignment. The control circuit sends a data retrieval request to one or more higher levels of a cache memory subsystem. When the requested data has arrived, the control circuit stores the requested data in the assigned cache line and sends a notification to the corresponding SIMD circuit.
Owner:ADVANCED MICRO DEVICES INC

Processor context storage and recovery method based on shadow register

The invention discloses a processor context storage and recovery method based on a shadow register. The main purpose of the invention is to realize more bottom-layer flexible processor context switching under complex processor scenes such as a multi-privilege mode based on a shadow register structure. By introducing a shadow register structure, when the context of the processor is saved and recovered, logic registers at different positions in a saving and loading instruction are mapped into different physical register files, so that the damage of a context saving program to an original site is avoided. Compared with a classic context storage and recovery method, the method can flexibly support the context switching of the processor in different modes, simplifies the context storage of a switching program, and improves the switching efficiency while guaranteeing the context integrity of the processor.
Owner:BEIHANG UNIV

Fully homomorphic encryption (FHE) operations on a unified FHE accelerator

PendingUS20260189363A1Theoretical computer scienceModular arithmetic
Techniques for fully homomorphic encryption are described. In some examples, a fully homomorphic encryption includes a butterfly compute circuitry to support a polynomial integer multiplication in response to an instance of a single instruction of a first type, wherein the instance of the single instruction is at least to include one or more fields for a register file address for a first source operand of the integer multiplication, one or more fields for a register file address for a second source operand of the integer multiplication, and one more fields for a register file address for a result of the integer multiplication, wherein the butterfly compute circuitry is to additionally support modular arithmetic operations.
Owner:INTEL CORP

Scheduling optimization processor, instruction scheduling method, apparatus, medium and product

This disclosure presents a scheduling optimization processor, instruction scheduling method, apparatus, medium, and product. The scheduling optimization processor includes multiple arithmetic pipelines and a physical register file. Each arithmetic pipeline is connected to at least two write-back caches. The number of instruction types with different execution latencies supported by the arithmetic pipeline is n. The write-back cache contains several cache partitions, with the number of cache partitions being n-1. According to the enqueue order of the cache partitions, the arithmetic pipeline writes the execution results of instructions into the idle write-back cache. In each clock cycle, the write-back cache writes the valid execution results from the last-level cache partition into the physical register file. This processor can achieve ordered buffering and parallel write-back of multiple instruction execution results by configuring multiple write-back caches for a single arithmetic pipeline and adopting a cache partition structure that matches the execution latency. This eliminates write conflicts in the physical register file at the hardware level, significantly improving processor issue efficiency and operating performance.
Owner:HYGON YUNXIN INTEGRATED CIRCUIT DESIGN (SHANGHAI) CO LTD

Soft and hard heterogeneous data processing system and method for atmospheric sounding lidar

ActiveCN122111888BData streamVirtual space
This application provides a hardware-software heterogeneous data processing system and method for atmospheric sounding lidar, relating to the technical field of embedded architecture for atmospheric remote sensing and scientific instruments. The system includes: a PL (Plug-in) terminal, which performs preliminary processing on lidar signal data to generate target data; a PS (Power Controller) terminal, including an internal cache and a processing core, with a virtual spatial filter configured within the processing core; the target data is stored in the internal cache via an ACP (Automatic Component Processing Filter); the processing core schedules the PL terminal to perform preliminary processing on the digital signal, reads the target data from the internal cache, and runs the virtual spatial filter to filter the target data; the virtual spatial filter performs register-level single-loop fusion of each independent algorithm node, completing the calculation of each independent algorithm within the single loop fusion, and all intermediate data resides in the register file. Relying on the ACP and single-loop fusion architecture, the bus occupancy rate and system latency of data flow are reduced, thereby achieving high throughput and zero bottlenecks at the system level.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Memory device, operating method thereof, and in-memory processing device

The invention discloses a memory device, an operating method thereof, and an in-memory processing device. The memory device includes an in-memory processing (PIM) block configured to perform an operation between a weight value and an input value, the weight value being represented by a weight scaling factor and a weight element, the input value being represented by an input scaling factor and an input element, where the PIM block includes: a first scaling register file storing the input scaling factor; a second scaling register file storing a weight scaling factor; a scalar register file (SRF) storing input elements; a plurality of arithmetic logic units (ALU) configured to perform a first operation between the input scaling factor and the weight scaling factor and a second operation between the input element and the weight element in parallel in response to an operation command received from a host; and an accumulator configured to accumulate and store operation results of the first operation and the second operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Data processing apparatus having streaming engine with read and read / advance operand coding

A streaming engine employed in a digital signal processor specified a fixed data stream. Once started the data stream is read only and cannot be written. Once fetched, the data stream is stored in a first-in-first-out buffer for presentation to functional units in the fixed order. Data use by the functional unit is controlled using the input operand fields of the corresponding instruction. A read only operand coding supplies the data an input of the functional unit. A read / advance operand coding supplies the data and also advances the stream to the next sequential data elements. The read only operand coding permits reuse of data without requiring a register of the register file for temporary storage.
Owner:TEXAS INSTRUMENTS INC