Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

106 results about "Register file" patented technology

A register file is an array of processor registers in a central processing unit (CPU). Modern integrated circuit-based register files are usually implemented by way of fast static RAMs with multiple ports. Such RAMs are distinguished by having dedicated read and write ports, whereas ordinary multiported SRAMs will usually read and write through the same ports.

Memory device and method

A memory device includes a plurality of memory banks, and a processing-in-memory (PIM) block accessible to the plurality of memory banks, wherein the PIM block comprises a control circuit configured to receive a plurality of operation instructions from a host and, in response to a predicated instruction indicating a predication operation among the plurality of operation instructions, instruct an arithmetic logic unit (ALU) to perform the predication operation, a predicate register file (PRF) configured to store therein a predicate value determined by the predication operation, and the ALU configured to perform an operation according to a command signal translated by the control circuit based on the predicate value from an operation instruction that depends on the predicate value among the plurality of operation instructions.
Owner:SAMSUNG ELECTRONICS CO LTD +1

Database concurrency control and memory access optimization method, device and equipment

The invention provides a database concurrency control and memory access optimization method which can be applied to the technical field of computer system software performance optimization. According to the method, multi-level instruction-level reconstruction is carried out on an execution path of a database kernel through the characteristics of an underlying hardware instruction set of a collaborative application processor platform, and the method comprises the following collaborative implementation optimization dimensions: based on a register file and a cache hierarchical structure of the processor platform, an access mode of a database core data structure is optimized; the number of memory access instructions is reduced; the data locality is improved; on the basis of atomic instruction set extension of a processor platform, instruction-level reconstruction is carried out on key primitives in a database multi-thread synchronization mechanism so as to reduce the overhead and contention of synchronization operation; on the basis of single-instruction multi-data-stream extension of a processor platform and a vector atomic operation instruction of the single-instruction multi-data-stream extension, parallel acceleration is carried out on batch life cycle management operation of objects in a database.
Owner:AEROSPACE INFORMATION RES INST CAS

Fully homomorphic encryption (FHE) operations on a unified FHE accelerator

PendingUS20260189363A1Theoretical computer scienceModular arithmetic
Techniques for fully homomorphic encryption are described. In some examples, a fully homomorphic encryption includes a butterfly compute circuitry to support a polynomial integer multiplication in response to an instance of a single instruction of a first type, wherein the instance of the single instruction is at least to include one or more fields for a register file address for a first source operand of the integer multiplication, one or more fields for a register file address for a second source operand of the integer multiplication, and one more fields for a register file address for a result of the integer multiplication, wherein the butterfly compute circuitry is to additionally support modular arithmetic operations.
Owner:INTEL CORP

Scheduling optimization processor, instruction scheduling method, apparatus, medium and product

This disclosure presents a scheduling optimization processor, instruction scheduling method, apparatus, medium, and product. The scheduling optimization processor includes multiple arithmetic pipelines and a physical register file. Each arithmetic pipeline is connected to at least two write-back caches. The number of instruction types with different execution latencies supported by the arithmetic pipeline is n. The write-back cache contains several cache partitions, with the number of cache partitions being n-1. According to the enqueue order of the cache partitions, the arithmetic pipeline writes the execution results of instructions into the idle write-back cache. In each clock cycle, the write-back cache writes the valid execution results from the last-level cache partition into the physical register file. This processor can achieve ordered buffering and parallel write-back of multiple instruction execution results by configuring multiple write-back caches for a single arithmetic pipeline and adopting a cache partition structure that matches the execution latency. This eliminates write conflicts in the physical register file at the hardware level, significantly improving processor issue efficiency and operating performance.
Owner:HYGON YUNXIN INTEGRATED CIRCUIT DESIGN (SHANGHAI) CO LTD

Soft and hard heterogeneous data processing system and method for atmospheric sounding lidar

ActiveCN122111888BData streamVirtual space
This application provides a hardware-software heterogeneous data processing system and method for atmospheric sounding lidar, relating to the technical field of embedded architecture for atmospheric remote sensing and scientific instruments. The system includes: a PL (Plug-in) terminal, which performs preliminary processing on lidar signal data to generate target data; a PS (Power Controller) terminal, including an internal cache and a processing core, with a virtual spatial filter configured within the processing core; the target data is stored in the internal cache via an ACP (Automatic Component Processing Filter); the processing core schedules the PL terminal to perform preliminary processing on the digital signal, reads the target data from the internal cache, and runs the virtual spatial filter to filter the target data; the virtual spatial filter performs register-level single-loop fusion of each independent algorithm node, completing the calculation of each independent algorithm within the single loop fusion, and all intermediate data resides in the register file. Relying on the ACP and single-loop fusion architecture, the bus occupancy rate and system latency of data flow are reduced, thereby achieving high throughput and zero bottlenecks at the system level.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Memory device, operating method thereof, and in-memory processing device

The invention discloses a memory device, an operating method thereof, and an in-memory processing device. The memory device includes an in-memory processing (PIM) block configured to perform an operation between a weight value and an input value, the weight value being represented by a weight scaling factor and a weight element, the input value being represented by an input scaling factor and an input element, where the PIM block includes: a first scaling register file storing the input scaling factor; a second scaling register file storing a weight scaling factor; a scalar register file (SRF) storing input elements; a plurality of arithmetic logic units (ALU) configured to perform a first operation between the input scaling factor and the weight scaling factor and a second operation between the input element and the weight element in parallel in response to an operation command received from a host; and an accumulator configured to accumulate and store operation results of the first operation and the second operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Vector Scatter and Gather with Single Memory Access

PendingUS20260119176A1Register arrangementsMemory hierarchyComputer architecture
Disclosed embodiments provide techniques for improved performance in processing vector instructions. A processor core is accessed. The processor core is coupled to a memory hierarchy, and the processor core includes one or more vector execution units (VUs), and one or more load store units (LSUs). The processor core includes a vector register file (VRF). The VRF includes multiple vector registers, and each vector register includes multiple vector elements. Vector elements that have a source or destination in contiguous memory are identified. Load store units (LSUs) take advantage of the contiguous memory condition by executing a vector load or vector store operation as a single memory access, requiring a reduced number of clock cycles. The single memory access satisfies each memory operation for each vector element within the vector register file.
Owner:AKEANA INC

Power management techniques using location-mapped chiplet configuration

This application is directed to an electronic device having a configurable group of voltage regulator cells. The electronic device includes a group of voltage regulator cells operating based on parameter settings of individual voltage regulator cells, output a rail voltage, and provide the rail voltage to multiple power rails. The electronic device includes a memory component coupled to the group of voltage regulator cells. The memory component stores an encoding table including multiple register files, and a register file defines parameter settings for the individual voltage regulator cells of the group of voltage regulator cells. The electronic device includes a setting interface for receiving a parameter setting signal applied to select the register file among the multiple register files for defining the parameter settings for the group of voltage regulator cells. The electronic device includes a substrate where the group of voltage regulator cells and the setting interface are integrated.
Owner:POWERLATTICE TECHNOLOGIES

Double-edge multi-phase clock generator and control method thereof

The invention discloses a double-edge multi-phase clock generator and a control method thereof, and belongs to the technical field of integrated circuit design. According to the method, the waveform freedom degree is improved through a three-stage time division architecture of'time period-round-period 'and in combination with a configuration mechanism of a plurality of programmable event points in each time period; configuration parameters are pre-stored in a register file module, and mode switching only needs to directly call pre-stored configuration, so that the real-time response capability is improved; the time control resolution is improved through double-edge counting, so that the sub-clock cycle control precision is achieved, and the method is suitable for high-precision application sensitive to power consumption and complex scene application.
Owner:INCORE MICRO TECH (WUXI) COM LTD

An ALU, an instruction execution method, a processor, a device, a medium and a program

This disclosure provides an ALU, an instruction execution method, a processor, a device, a medium, and a program, relating to the field of computer technology, specifically information processing, deep learning, artificial intelligence, and chip technology. The ALU is integrated into the processor and includes a first data conversion module and a calculation module. The first data conversion module and the calculation module are connected, and the first data conversion module is also connected to a target register file. The first data conversion module receives first-precision data required for the calculation of a target instruction read from the target register file and converts the first-precision data into second-precision data; the precision of the first-precision data is lower than that of the second-precision data. The calculation module receives the second-precision data and executes the calculation operation of the target instruction based on the second-precision data. The embodiments of this disclosure can fully utilize hardware computing resources to improve instruction execution efficiency, thereby improving the overall execution performance of the arithmetic logic unit.
Owner:KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD

Stochastic sampling of memory operations at a processing unit

During execution of software a processing unit issues asynchronous operations such that there are multiple asynchronous operations, such as memory operations, in flight (that is, pending execution completion) from a single set of instructions, such as a wavefront or warp. In some cases, the processing unit executes other operations while the multiple asynchronous operations are pending. Performance monitor circuitry records information, such as asynchronous operation count information, register file scoreboard information, and the like, that allows a software engineer to identify which of a plurality of asynchronous operations caused a stall.
Owner:ADVANCED MICRO DEVICES INC

Hardware-based power management integrated circuit register file write protection

Disclosed are devices and methods for protecting the register file of a power management integrated circuit (PMIC). In one embodiment, a device is disclosed comprising: a register file comprising a plurality of a registers, at least one register in the register file containing a write register bit (WRB); and an interface configured to receive messages from a host application, the messages including a WRB enablement signal, wherein the device is configured to enable writing to the register file in response to receiving the WRB enablement signal over the interface, write data in response to write messages while writing to the register file is enabled, and disable writing to the register file in response to receiving a stop bit over the interface.
Owner:LODESTAR LICENSING GROUP LLC

Generating iteration transfer information for code execution with a compute slice microarchitecture

A processor core is accessed. The core is configured to execute instructions associated with an instruction set architecture (ISA). The core comprises a plurality of compute slices, a plurality of barrier register files, and a control unit. Each compute slice includes at least one arithmetic logic unit (ALU), a local register file, and is coupled to a successor compute slice and a predecessor compute slice by a barrier register file. Code associated with the ISA is evaluated, where the code includes a first loop. The evaluating includes generating iteration transfer information associated with the first loop. Each slice task within a plurality of slice tasks associated with the first loop is distributed to a compute slice. The processor core executes the plurality of slice tasks. Data forwarding between successive compute slices is based on the plurality of barrier register files and the iteration transfer information.
Owner:ASCENIUM INC

A hardware configuration designed for the execution of ascon cryptographic methods while defending against side-channel attacks

A low-area hardware architecture to execute the ASCON cipher suite and resist side-channel attacks that have a co-processor having controllers, register file, a permutation operably configured to receive input from a multiplexor structure and execute ASCON permutation, an ASCON state register that can be optionally removed and substituted by the register layer inside the permutation unit, and XOR logic gates that are operably configured to receive input from the register file and the permutation unit and provide input to the multiplexor structure.
Owner:PQSECURE TECHNOLOGIES LLC

Data transmission method, functional unit, cluster of functional units, processor and device

This application provides a data transmission method, functional unit, functional unit cluster, processor, and device. The method includes: acquiring data while parsing a target field in an instruction; wherein the target field consists of a target field module index and a register file number; determining a data path based on the target field module index; and sending the data to the register file corresponding to the register number through the data path. The method provided by this application acquires data while parsing a target field in an instruction; wherein the target field consists of a target field module index and a register file number; determining a data path based on the target field module index; and sending the data to the register file corresponding to the register number through the data path. Data transmission can be completed with a single instruction, reducing the execution time of one instruction compared to existing technologies and improving transmission efficiency.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

A verification method and system of a RISC-V out-of-order superscalar processor

PendingCN122450760AVerificationSimics
The application belongs to the technical field of processors, and provides a verification method and system of a RISC-V out-of-order superscalar processor, a test program is generated based on a random instruction generator of a RISC-V instruction set, and the test program is run in parallel by an out-of-order superscalar processor and an instruction set simulator; in the running process of the out-of-order superscalar processor, retired instruction information is sequentially extracted from a reorder buffer thereof, and corresponding instruction execution results are read from a physical register file according to the information of the retired instruction, a check information queue table item of sequential execution is generated, and is stored into a check information queue; through a preset interface, the instruction set simulator is controlled to run in single step, and sequential submission results corresponding to the check information queue table item are extracted and stored into a submission result queue; the table items in the check information queue and the submission result queue are compared one by one, if the comparison is consistent, the verification is continued, and if the comparison is inconsistent, an error instruction is located. The efficiency of verification is improved.
Owner:SHANDONG UNIV

Vector mask buffers in a vector instruction execution pipeline

Systems and methods related to vector mask buffers in a vector instruction execution pipeline are disclosed herein. The vector instruction execution pipeline may include several lanes. Each lane may include a vector register file, a vector mask buffer, and a functional processing unit. The vector register file may store operand data and the vector mask buffer may store a vector mask associated with the operand data. In a lane, the operand data may be read from the register file into a functional processing unit, and the vector mask may be read from the vector mask buffer to the functional processing unit. The functional processing unit may process the operand data based on the vector mask. The lane-specific vector mask buffers improve the efficiency of the vector instruction execution pipeline by storing the vector masks proximate to where the vector masks will be used.
Owner:TENSTORRENT USA INC

Graphics processor register file including a low energy portion and a high capacity portion

One embodiment provides circuitry coupled with cache memory and a memory interface, the circuitry to compress compute data at multiple cache line granularity, and a processing resource coupled with the memory interface and the cache memory. The processing resource is configured to perform a general-purpose compute operation on compute data associated with multiple cache lines of the cache memory. The circuitry is configured to compress the compute data before a write of the compute data via the memory interface to the memory bus, in association with a read of the compute data associated with the multiple cache lines via the memory interface, decompress the compute data, and provide the decompressed compute data to the processing resource.
Owner:INTEL CORP

Memory device and method with processing-in-memory block

A memory device includes a first scalar register file storing a first input fragment, a second scalar register file storing a second input fragment, an arithmetic logic unit (ALU), and a control circuit. The control circuit is configured to perform, using the ALU, a first operation between the first input fragment and a first weight fragment based on a first operation command received from a host, and to perform, using the ALU, a second operation between the second input fragment and a second weight fragment based on a second operation command received from the host.
Owner:SAMSUNG ELECTRONICS CO LTD

3D in-pipeline private scratchpad memory for SIMT compute cores

A processor core (1) of the SIMT type and a related method of operation are disclosed. The processor core comprises frontend circuitry (10), a plurality of lanes (20a-f), a multi-banked and arbitration-free scratchpad memory (40), and an interconnect (50) to couple active ones of the plurality of lanes to different ones of the scratchpad memory banks. There are at least as many banks as lanes in the processor core. Each lane comprises a thread-private register file (21a-f) and address generation logic (22a-f) to independently calculate an effective address of an operand within the scratchpad memory. An execution pipeline of the processor core, associated with the thread-parallel execution of vector memory instructions, comprises at least the address generation logic of the different lanes, the interconnect, and the different banks as pipeline components. The processor core is implemented as a stack of dies (2a, 2b) including the frontend circuitry and the plurality of lanes on a first die (2a) and the scratchpad memory on a second die (2b).
Owner:INTERUNIVERSITAIR MICRO ELECTRONICS CENT (IMEC VZW)

Hardware instruction-level scheduling in processors

A processing core for executing threads is provided. The processing core comprises: one or more functional units; a shared register file comprising a plurality of registers, each register of the shared register file being configured to store data for executing any thread; and a multi-thread elaboration unit, MTEU, the MTEU comprising: a control module configured to: receive an indication of a plurality of active threads for a given scheduling window; and receive instructions for each of the plurality of active threads; a context manager module configured to load, for each of the plurality of active threads, data required for executing a next one or more of the instructions into a respective one or more registers of the shared register file; and an instruction scheduler module configured to schedule the instructions of the active threads for execution based on availability of the one or more functional units.
Owner:POLITECNICO DI MILANO +1

Register allocation decision

An apparatus comprises execution circuitry configured to execute a given instruction to produce a given data value. Value analysis circuitry is configured to perform an analysis of the given data value produced by the execution circuitry to determine at least one property of the given data value, and register allocation circuitry is configured to make a register allocation decision regarding storage of the given data value in a physical register file in dependence on the analysis of the given data value.
Owner:ARM LTD

Atomic operation execution method

The invention discloses an atomic operation execution method, which relates to the technical field of computers, and comprises the following steps: determining a target atomic register corresponding to an access request, a target address in the target atomic register and an operation type in a register file unit through an address decoder; and orderly executing atomic operation based on target atom register, the target address and the operation type through a control logic and sequencer, when it is determined that register value transformation information corresponding to the target address meets a preset interrupt condition through the interrupt control logic module, interrupt trigger information is sent to an interrupt signal generator; and generating a corresponding interrupt based on the interrupt trigger information through an interrupt signal generator, and sending the interrupt to a second target CPU corresponding to the interrupt trigger information. All atomic operations are realized in the module in a hardware mode, an atomic operation register and interrupt control logic are closely integrated on the hardware level, and atomic synchronization of a cross-CPU cluster is realized.
Owner:KINGTIGER TESTING TECH (SZ) LTD

Special execution control device and control method for in-memory calculation

The invention provides a special execution control device and control method for in-memory calculation. The control device comprises an instruction caching assembly used for receiving and caching batch bit operation instructions, a microprogram storage assembly used for caching microprograms, a microoperation code storage assembly used for storing a currently executed microoperation code sequence, an addressing assembly used for generating a DRAM row physical address, and a data processing assembly used for processing the DRAM row physical address. The register file component is used for storing a general register and control information, the counting component is used for recording the processing number of data blocks, the control component is used for decoding a microoperation code and controlling an execution process, and the counting control component is used for pointing to the currently executed microoperation code. According to the control device provided by the invention, through cooperative work of all the components, the execution process of in-memory calculation is optimized, the throughput and energy efficiency of the system are remarkably improved, and meanwhile, the control device has good flexibility and integration.
Owner:TIANJIN JINHANG COMP TECH RES INST

Batch data processing method and device suitable for database, equipment and medium

The invention discloses a batch data processing method and device suitable for a database, equipment and a medium, relates to the technical field of data processing, and is applied to a processor. Monitoring whether a target instruction meeting a preset stride condition exists in a data processing instruction stream corresponding to the preset database in the target assembly line or not, if yes, entering an instruction pre-execution mode, vectorizing the target instruction and a subsequent instruction having a dependency relationship with a register of the target instruction, and outputting the vectorized instruction; performing instruction pre-execution based on the vectorized instruction and a first execution unit, and storing an instruction pre-execution result to a first physical register file and a result cache region; and when the data processing instruction stream in the target assembly line reaches the input end of the second execution unit, matching the instruction pre-execution result in the result cache region, and completing batch data processing operation based on an information matching result. According to the invention, the efficiency and accuracy of batch data processing are improved.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Multi-thread processor pipeline architecture system, scheduling method and equipment

The invention relates to the technical field of processor design, discloses a multi-thread processor pipeline architecture system, a scheduling method and equipment, and designs an instruction prefetching scheduling algorithm based on mixed priorities, and the algorithm realizes load balancing of thread instruction queues through a static level decision mechanism and a dynamic level decision mechanism. The static level preferentially processes an empty queue thread to avoid starvation, and the dynamic level dynamically adjusts instruction fetching priority according to each queue depth, thread validity and a branch risk state to ensure that an instruction stream is continuously and stably supplied to a subsequent stage; a coupled polling emission scheduling algorithm is provided, so that high complexity of full-permutation search is avoided, and structural conflicts and data conflicts are effectively reduced; a special hardware architecture is designed around the algorithm, a multi-program counter and an instruction queue are integrated at the front end, a double-transmitting channel, a multifunctional unit and a thread private register file are configured at the rear end, and the performance of the single-core CPU is improved by optimizing the hardware architecture.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Fused comparison add instructions

An apparatus, system, and method for efficiently processing pairs of operations repeatedly used in applications. In various implementations, a computing system includes a parallel data processing circuit with multiple compute circuits. Each of the compute circuits includes multiple lanes of execution, each with a corresponding arithmetic logic unit (ALU). The ALU supports executing a single fused conditional ternary instruction that replaces two separate instructions that provide two operations (comparison and add). When executing the fused conditional ternary instruction, the ALU does not retrieve the intermediate result from the scalar register file, the vector register file, or bypass circuitry located externally from the ALU. Rather, the ALU generates the intermediate result and uses the intermediate result without routing the intermediate result externally from ALU.
Owner:ADVANCED MICRO DEVICES INC

Memory access optimization device for convolutional neural network accelerator

The invention provides a memory access optimization device for a convolutional neural network accelerator, and belongs to the technical field of integrated circuit design, the memory access optimization device comprises a pulsation processing unit array which is used as a computing core, adopts a regular two-dimensional grid structure and comprises processing units, and each processing unit is provided with a basic multiplication and addition operation unit and a local register; the ifmaps second-level cache system adapts to locality of convolution operation through hierarchical scheduling, the reuse rate of an input feature map is improved, and redundant memory access is reduced; and a part of the sum register file is directly connected with the pulse processing unit array and is used for caching a part of the sum register file and an intermediate result generated in the calculation process. According to the method, a hierarchical data multiplexing system is established, so that the data multiplexing efficiency of the input feature map and the partial sum is remarkably improved on the premise of not excessively increasing the hardware complexity, and the storage access frequency and the system power consumption are effectively reduced.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

GPU-oriented compiler constant storage method, apparatus and device, and medium

The invention provides a GPU-oriented compiler constant storage method and device, equipment and a medium, and belongs to the field of data processing. According to the method, a control buffer with a dynamic size is divided in a scalar register file and is shared by a whole kernel, a Workload Manager is responsible for space distribution, recovery and initialization, and constants are stored in the control buffer divided by a scalar register during compiling, so that a GPU of an SIMT architecture does not need to have multiple copies in a scalar register file at the same moment, and the number of copies in the scalar register file at the same moment is reduced. And a constant does not need to be obtained from a memory, and the instruction code can accommodate a plurality of 64-bit immediate operands, so that the problem of insufficient resource utilization of a large number of scalar registers in the prior art is solved.
Owner:METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD

Computing devices, methods of operation, and machine readable storage media

A computing device, a method of operation thereof, and a machine-readable storage medium are provided. The computing device includes execution unit (EU) circuitry and instruction scheduler (IS) circuitry. The IS circuitry can issue instructions to the EU circuitry for execution. The IS circuitry includes a descriptor address register (DAR). The DAR is used to store address information of a resource descriptor. In a case where a current instruction being processed by the IS circuitry uses the resource descriptor, the IS circuitry accesses the address information stored in the DAR local to the IS circuitry without triggering the EU circuitry to read the address information of the resource descriptor from a scalar register file external to the IS circuitry to the IS circuitry. After obtaining the address information of the resource descriptor, the IS circuitry issues the current instruction using the resource descriptor to the EU circuitry for execution based on the address information.
Owner:SHANGHAI BIREN TECH CO LTD