Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

14 results about "Address generation unit" patented technology

The address generation unit (AGU), sometimes also called address computation unit (ACU), is an execution unit inside central processing units (CPUs) that calculates addresses used by the CPU to access main memory. By having address calculations handled by separate circuitry that operates in parallel with the rest of the CPU, the number of CPU cycles required for executing various machine instructions can be reduced, bringing performance improvements.

Access system, memory system and access method for three-dimensional data access

The embodiment of the invention relates to the technical field of integrated circuit design, in particular to an access system, a memory system and an access method for three-dimensional data access. The access system comprises a configuration register unit, an address generation unit, a data assembling and splitting unit and a double-pipeline controller, the address generation unit is connected with the configuration register unit and the double-pipeline controller, and is used for calculating a one-dimensional physical address based on a target index value input by a user and sending the one-dimensional physical address to the double-pipeline controller; the dual-pipeline controller is connected to the memory interface and is used for independently managing a read flow and a write flow based on a user read-write request so as to perform read-write in the memory based on a one-dimensional physical address; the data assembling and splitting unit is connected to the double-assembly-line controller and used for carrying out data assembling and splitting in the read-write process and then sending the data to the double-assembly-line controller. According to the scheme, three-dimensional data indexing is achieved in a hardware mode, and the data processing efficiency and the system flexibility can be improved.
Owner:SKYRELAY (BEIJING)TECH CO LTD

Data processing apparatus for layered decoding of LDPC codes

The application discloses a data processing device for LDPC code layered decoding, which comprises a configurable computing unit capable of switching between row update mode and column update mode, a storage system physically divided into multiple independent storage banks, an address generation unit and a control unit. The address generation unit is configured to map access addresses of parallel data requests to different independent storage banks, and the control unit is configured to schedule column update tasks to be performed first and then closely related row update tasks in a decoding layer, wherein the row update tasks directly read the latest variable node messages generated by the column update tasks. The application guarantees the immediate reuse requirement of the layered decoding algorithm for the latest messages by constructing a high-bandwidth parallel data access path, eliminates the memory access bottleneck, and thus can fully exert the performance advantage of the layered decoding algorithm to realize high-throughput decoding processing.
Owner:SHANGHAI JINGJI COMM TECH CO LTD

Joint scheduling load in data processing apparatus

Data processing apparatus and methods of operating such data processing apparatus are disclosed. An issue circuit buffers operations until operands required by the operations are available in a set of registers before execution. When a first load operation and a second load operation both depend on a common operand, and when the common operand is available in the set of registers, the first load operation and the second load operation are identified in the issue circuit. A load circuit has a first address generation unit to generate a first address for the first load operation and a second address generation unit to generate a second address for the second load operation. An address comparison unit compares the first address and the second address. The load circuit is arranged to cause a merged lookup to be performed in a local scratchpad storing copies of data values from a memory when the address comparison unit determines that the first address and the second address differ by less than a predetermined address range characteristic of the local scratchpad.
Owner:ARM LTD

Instruction processing apparatus, method and chip

This application discloses an instruction processing apparatus, method, and chip, belonging to the field of integrated circuit design technology. The apparatus includes: an instruction fetch unit for fetching instructions and breaking them down into multiple micro-operations, including vector load instructions or vector store instructions, where one of the micro-operations processes at least one vector element; an address generation unit for determining an address generation strategy matching the instruction type, and sequentially generating addresses corresponding to the multiple micro-operations according to the address generation strategy, wherein multiple instruction types have a one-to-one correspondence with multiple address generation strategies, and the execution of the address generation strategy is based on the address correlation between vector elements in the corresponding instruction type; and an execution unit for executing the memory access processing corresponding to the multiple micro-operations according to the addresses. This accelerates address generation efficiency.
Owner:BEIJING ESWIN COMPUTING TECH CO LTD

Hardware layer-by-layer program control in a reconfigurable dataflow processor

ActiveUS12717592B2Address generation unitComputer architecture
A reconfigurable data processor includes a bus system, an array of configurable units, and a configuration load controller connected to the bus system and coupled to a memory. The configuration load controller incorporates a first set of registers accessible from a host processor for storing addresses of a first configuration file, a second set of registers loaded by loading a configuration file for storing addresses of a second configuration file, and an address generation unit with working address registers. The processor is configured to load a first configuration file from the memory and initiate execution based on a request from runtime software. Additional configuration files are automatically loaded upon completion of a previous configuration file based on information stored in the previous configuration file.
Owner:SAMBANOVA SYSTEMS INC

A heterogeneous computing direct memory access method and system based on tensor semantic perception

The application provides a tensor semantic perception-based heterogeneous computing direct memory access method and system, including the following steps: constructing a tensor DMA descriptor in the memory of a transmission end, wherein the tensor DMA descriptor includes a sparse mask bitmap; a first DMA engine reads the sparse mask bitmap; a first tensor address generation unit in the first DMA engine queries a mask bit in the sparse mask bitmap for address generation and filtering; when the data corresponding to the mask bit is a nonzero value, the first DMA engine generates a read request; when the data corresponding to the mask bit is a zero value, the first DMA engine does not generate a read request; and the first DMA engine packs the read nonzero value data to generate a data packet.
Owner:SHANGHAI XINLIJI SEMICON CO LTD

I2S data DSP short-range storage and transmission system and method based on arbitration mechanism

ActiveCN121880242ASound input/outputData conversionAddress generation unitComputer architecture
The invention discloses an I2S data DSP short-range storage and transmission system and method based on an arbitration mechanism, relates to the technical field of embedded audio and real-time signal processing, and aims to solve the technical problems that in the prior art, a plurality of hardware autonomous deterministic direct storage channels for short-range storage from I2S to DSP cannot be constructed, transmission delay cannot be predicted, and software overhead is large. The system comprises an I2S interface, a special hardware bridging module, a DSP containing short-range storage, a system bus and a DSP core, and the special hardware bridging module integrates an FIFO, an arbiter, an address generation unit and other components and can autonomously complete data caching, arbitration, address management and direct writing operation. The method comprises five steps of system initialization, hardware autonomous direct storage, hardware accurate notification, core direct processing and hardware synchronous feedback. According to the method, hardware autonomous deterministic low-delay transmission of the I2S data is realized, the software overhead is reduced, the system adaptability and the hardware reusability are good, and the method is suitable for various embedded audio and real-time signal processing scenes.
Owner:JUDI (SHANGHAI) TECHNOLOGY CO LTD

Storage controller for self-reconfiguration and self-evolution AI chip and chip thereof

The invention relates to a storage controller for a self-reconfiguration and self-evolution AI chip and a chip thereof, in particular to the technical field of chips, the storage controller comprises a plurality of data cache modules, and each data cache module is configured to manage storage and access of a three-dimensional data block in an on-chip SRAM array; each data caching module comprises an instruction decoding unit which is configured to receive and analyze an input instruction, and if it is determined that the instruction is a configuration instruction, configuration parameters analyzed from the instruction are updated to a corresponding configuration register in the instruction decoding unit; if it is determined that the instruction is a working instruction, working parameters analyzed from the instruction are updated to a corresponding configuration register in the instruction, and after the parameters are updated, starting control signals are sent to a DDR address generation unit and an SRAM management unit respectively, so that efficient and configurable access to data blocks in the three-dimensional space can be achieved.
Owner:XIAN UNIV OF POSTS & TELECOMM

Resource memory access device and method, chip and computer equipment

The invention discloses a resource access device and method, a chip and computer equipment, and belongs to the technical field of resource management. The device comprises a shader core, an address generation unit and a memory management unit, the shader core is used for sending a resource access request to the address generation unit, and the resource access request comprises resource address information corresponding to a to-be-accessed resource; the address generation unit is used for determining a block mapping table address and an in-block offset corresponding to the to-be-accessed resource based on the resource address information under the condition that the to-be-accessed resource is a sparse resource, and determining a virtual address corresponding to the to-be-accessed resource based on the block mapping table address and the in-block offset, transmitting the virtual address to the memory management unit; and the memory management unit is used for converting the received virtual address into a physical address and executing a resource access operation based on the physical address. According to the scheme provided by the invention, the resource access process can be optimized, and the flexibility of block mapping management of sparse resources is improved.
Owner:MOORE THREADS TECH CO LTD

Data operation module applied to processor, processor, equipment and method

PendingCN121277560APhysical realisationMachine execution arrangementsAddress generation unitData pack
The invention discloses a data operation module applied to a processor, the processor, equipment and a method, and relates to the technical field of chips. The data operation module comprises a first address generation unit, a request preprocessing unit, a first buffer, a first input unit and an operation unit, the first address generating unit is used for generating first address information, and the first address information is used for acquiring m pieces of data from the first buffer; the request preprocessing unit is used for shifting n pieces of data stored in the request preprocessing unit to obtain n pieces of shifted data under the condition that the second address information is stored in the request preprocessing unit and the n pieces of data obtained from the first buffer based on the second address information comprises m pieces of data; the first input unit is used for loading m pieces of data into the operation unit based on the n pieces of shifted data; and the operation unit is used for executing operation according to the m data. According to the scheme, the access frequency of the first buffer is reduced.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

I2s data dsp short-range storage transmission system and method based on arbitration mechanism

ActiveCN121880242BSound input/outputData conversionAddress generation unitComplete data
The application discloses an I2S data DSP short-range storage transmission system and method based on an arbitration mechanism, relates to the technical field of embedded audio and real-time signal processing, and aims to solve the technical problems that the prior art cannot construct a multi-channel I2S to DSP short-range storage hardware autonomous deterministic direct storage path, transmission delay is unpredictable, and software overhead is large. The system comprises an I2S interface, a special hardware bridging module, a DSP containing short-range storage, a system bus and a DSP core. The special hardware bridging module is integrated with components such as FIFO, an arbitrator and an address generation unit, and can autonomously complete data caching, arbitration, address management and direct writing operation. The method comprises five steps of system initialization, hardware autonomous direct storage, hardware precise notification, core direct processing and hardware synchronous feedback. The application realizes hardware autonomous deterministic low-delay transmission of I2S data, reduces software overhead, and has good system adaptability and hardware reusability, and is suitable for various embedded audio and real-time signal processing scenes.
Owner:JUDI (SHANGHAI) TECHNOLOGY CO LTD

Reconfigurable cryptographic accelerator based on mod 30 two-dimensional sieve table and fpga implementation method thereof

PendingCN122457248AAddress generation unitComputer architecture
The application discloses a reconfigurable cryptographic accelerator based on a modulo-30 two-dimensional sieve table and an FPGA implementation method thereof, and belongs to the technical field of FPGA hardware acceleration and cryptographic chips. The application constructs an 8-column and multi-row two-dimensional sieve table structure by using a modulo-30 reduction residue system, and compresses a candidate prime number space by 73.3%; a modulo-30 address generation unit is adopted to generate a storage address by using a 3-bit state machine and a row counter, only shift and addition operations are needed, and division and modulus are completely replaced; a matrix deletion operation unit is used to realize parallel screening of multiple processing units, and address iteration and bit operation are used to complete composite marking; a reconfigurable multi-mode application module is used to dynamically distribute prime number streams to different functional modules such as prime number direct output, random number generation, key derivation, hash calculation and the like, and supports runtime mode switching. The application realizes high-speed, low-resource and deterministic prime number generation and multi-mode cryptographic acceleration on an FPGA, and can be applied to cryptographic chips, blockchain nodes, post-quantum cryptographic systems and secure terminal devices.
Owner:刘诗章 +1

Apparatus, system, and method of compiling code for a processor

PendingUS20260140726A1Code compilationAddress generation unitComputer architecture
For example, a compiler may be configured to identify a plurality of memory-access operations in a loop based on a source code to be compiled into a target code to be executed by a target processor, the plurality of memory-access operations may include at least a first memory-access operation and a second memory-access operation to a same memory pointer. For example, the first memory-access operation may have a first offset, and the second memory-access operation may have a second offset different from the first offset. For example, the compiler may configure Address Generation Unit (AGU) configuration code to configure the plurality of memory-access operations by a same AGU. For example, the compiler may generate the target code based on compilation of the source code. For example, the target code may be based on the AGU configuration code.
Owner:MOBILEYE VISION TECH LTD

Heterogeneous computing direct memory access method and system based on tensor semantic perception

The invention provides a heterogeneous computing direct memory access method and system based on tensor semantic perception, and the method comprises the following steps: constructing a tensor DMA descriptor in a memory of a transmission end, the tensor DMA descriptor comprising a sparse mask bitmap; a first DMA engine reads the sparse mask bitmap; a first tensor address generation unit in the first DMA engine inquires a mask bit in the sparse mask bitmap to perform address generation and filtering, when data corresponding to the mask bit is a non-zero value, the first DMA engine generates a reading request, and when data corresponding to the mask bit is a zero value, the first DMA engine generates a reading request; the first DMA engine does not generate a read request; and the first DMA engine packages the read non-zero value data to generate a data packet.
Owner:SHANGHAI XINLIJI SEMICON CO LTD