Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

110 results about "Register bank" patented technology

Register Banks. A register bank is used for the programmable registers used by assembly language programmers. It can be viewed as the hardware equivalent of a software array. It has ports for reading and writing data given an index.

Vector and matrix calculation-oriented memory access system

The invention provides a memory access system oriented to vector and matrix calculation, the system comprises a vector memory access unit, a matrix memory access unit, a vector register group and a matrix register group, the vector memory access unit is connected with a memory interface and the vector register group, reads elements of a one-dimensional data structure or a two-dimensional data structure from a memory, and stores the elements of the one-dimensional data structure or the two-dimensional data structure; vector data are generated through data reorganization operation, a matrix access unit is connected with a memory interface and a matrix register set, matrix block data of a two-dimensional data structure are read from a memory together with the matrix access unit, matrix data are generated after data reorganization, and the matrix data are broadcasted to one or more computing units according to rows or columns. And performing calculation on the matrix data and the vector data. In the memory access system, the vector memory access unit and the matrix memory access unit can load data in parallel, and the utilization rate is improved through a plurality of computing units, so that the problems of low memory access efficiency and low memory bandwidth utilization rate are solved.
Owner:NANJING UNIV

Data processing device, data processing method, chip and electronic equipment

The invention discloses a data processing device, a data processing method, a chip and electronic equipment, and relates to the technical field of data processing. The data processing device comprises a storage module which is configured to store x groups of original data vectors, the original data vectors comprise m data elements, and x and m are positive integers; the calculation module is configured to access the original data vectors stored in the storage module, and reconstruct the x groups of original data vectors into one or more intermediate matrixes according to the parallel processing capability of the single-instruction multi-data execution unit and a target value k of a Top-k operator; loading the intermediate matrix into a vector register block according to columns, carrying out odd-even merging sorting by utilizing a single-instruction multi-data execution unit, and realizing in-line sorting on the intermediate matrix; and loading the intermediate matrix subjected to in-line sorting into a vector register block according to columns, and performing one or more rounds of merging sorting by utilizing a single-instruction multi-data execution unit to obtain top-k data elements in x groups of original data vectors.
Owner:BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

AXI bus monitoring system and method

The invention provides an AXI bus monitoring system and method, the system comprises a host side monitor, a slave side monitor and an error summarization processing module, and the error summarization processing module is respectively connected with the host side monitor and the slave side monitor; a plurality of register groups are arranged in each monitor, and each register group corresponds to one timer and one state machine and is used for monitoring one read or write transaction; the timer is used for carrying out timeout counting on the read / write transaction, and when the count exceeds a preset value, a timeout error is triggered to be reported to the error summarization processing module; the state machine is used for tracking which stage the read or write transaction is performed in real time; the register block is used for latching an AXI bus signal so that a monitor constructs data to complete access; and the error summarization processing module is used for executing corresponding processing according to the error codes reported by the host side monitor and / or the slave side monitor and the information latched by the register block. The system can effectively avoid system deadlock.
Owner:ARKMICRO TECH

RISC-V tensor instruction set extension method and device based on intelligent processor

The invention provides an RISC-V tensor instruction set extension method based on an intelligent processor. The method comprises the steps that a special tensor configuration register set is introduced to store tensor element information and operation element information; encoding a tensor extension instruction by adopting a 64-bit fixed-length instruction encoding format; and setting high-capacity on-chip tight coupling storage with uniform addressing in a processor core, and taking the high-capacity on-chip tight coupling storage as a direct operation space of all tensor expansion instructions. The invention further provides an intelligent processor-based RISC-V tensor instruction set extension device, a storage medium and electronic equipment. According to the method, the RISC-V tensor instruction set extension which supports high-dimensional tensor operation, is dense in calculation and is friendly to hardware can be realized.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Large language model operator compiling method and system oriented to GPU (Graphics Processing Unit) platform

The invention provides a large language model operator compiling method oriented to a GPU platform, and belongs to the field of GPU operator compilation. A large language model operator is compiled through a Python interface to obtain a Python code; performing lexical analysis and grammatical analysis on the Python code to generate an abstract syntax tree; mapping the abstract syntax tree into an MMIR extension intermediate representation; tensor layout in the MLIR extension intermediate representation is uniformly converted into a binary linear mapping matrix, and optimization problems of global memory layout, shared memory layout and register block layout are analyzed based on the binary linear mapping matrix; automatically deducing a non-key tensor layout based on a program control flow diagram, generating a global optimal layout scheme, and obtaining an optimized MMIR extension intermediate representation; enabling the optimized MLIR extension intermediate representation to pass through a GGPU (Graphics Graphics Processing Unit) target code generator of the MLIR, and generating a PTX code adaptive to the NVIDIA GPU; the PTX code is loaded to a GPU, and an operator calculation result is output; the invention further provides a compiling system. The problem that the current large language model operator compiling efficiency is low is solved.
Owner:江淮前沿技术协同创新中心

In-system testing architecture for autonomous systems and applications

Embodiments of the present disclosure relate to applications, platforms, architecture, etc. for using a master test image that may be used for multiple different tests. For example, a testing system may include a register bank that may be loaded with test configurations corresponding to one or more tests. The test configurations may respectively correspond to sets of control packets included in the master test image that may be used or executed for corresponding tests. The test configurations may indicate execution orders of their respective sets of control packets in which the execution order of one or more of the control packets included in the sets of control packets may differ from a default execution order of such control packets as indicated in the master test image. Such a configuration may accordingly allow for the flexibility of performing many different tests using a single master test image.
Owner:NVIDIA CORP

AI chip reasoning acceleration method and system and medium

The invention relates to the related technical field of reasoning calculation, in particular to an AI chip reasoning acceleration method and system and a medium, and the method comprises the steps of collecting adaptive scheduling parameters, setting a hierarchical pipeline response gradient under a reasoning acceleration framework, loading the optimized scheduling parameters to a configuration register group, and setting reasoning self-driven circulation. The technical problems that load feature perception is one-sided, scheduling parameters are statically solidified, dynamic adjustment along with reasoning tasks cannot be achieved, adaptability to feature graph structure changes and activation mode updating in the reasoning process is poor, and reasoning stability is insufficient are solved, and a tracing strategy based on a three-dimensional load feature vector and attention guidance is achieved. The method has the technical effects that the matching degree of scheduling parameters and task states is improved, the scheduling parameters and model compression granularity are dynamically adjusted in combination with real-time energy consumption constraints, a federated on-chip learning architecture enables the parameter iteration updating efficiency of a layer sensitivity analysis unit to be improved, the reasoning stability and the continuous operation reliability are enhanced, and the method adapts to end side complex dynamic scene requirements.
Owner:TIANJIN RUILEX ENVIRONMENTAL PROTECTION ENG CO LTD

Message encryption method and apparatus

The application discloses a message encryption method and device, and relates to the technical field of data encryption. The message encryption method comprises the following steps of: in a previous clock cycle of processing adjacent first and second message words, a first hash module calculates a first e value and a first a value according to the first message word, a hash register value and a hash constant; a second hash module calculates a second e value and a second a value according to the second message word, the hash register value, the first e value, the first a value and the hash constant; and the second a value, the first a value, the second e value and the first e value are sequentially cached into a third hash register, a fourth hash register, a seventh hash register and an eighth hash register of a hash register group. Two hash modules are adopted to process two message words in the same clock cycle, the clock cycle used for calculating the message word digest value in the message encryption process is reduced, and therefore the user experience can be obviously improved.
Owner:BEIJING YONGDING INTELLIGENT TECH CO LTD

Multi-channel LED digital control system and method

The invention discloses a multichannel LED digital control system and method. The system comprises a main controller which generates an LED control instruction and sends a data frame through a differential bus; the plurality of LED driving modules are connected with the main controller through a differential bus, and each LED driving module comprises a differential receiver for receiving a differential signal of the differential bus; the UART communication module is connected with the differential receiver and converts the differential signal into serial data; the data processing state machine module is connected with the UART communication module and is used for analyzing the received data frame, and the data frame comprises a synchronous byte, a device address / command composite byte, a register address byte, a data byte and a CRC (Cyclic Redundancy Check) check byte; the digital logic counter is used for measuring the bit width of each data bit in the synchronous bytes and calibrating the sampling time sequence of the UART communication module based on a measurement result; the register group is used for storing LED control parameters; and the multi-path PWM generator is connected with the register block and generates multi-path PWM signals according to the LED control parameters so as to drive the LEDs. The power consumption is reduced.
Owner:PURESEMI CO LTD

Display system capable of automatically brushing pictures

The invention relates to the technical field of display, in particular to a display system capable of automatically brushing pictures, which comprises a TDDI chip integrating a main control function and a display driving function, and a storage module which is in communication connection with the TDDI chip and is used for storing background picture data and windowing picture data, a software application layer, a register block, a storage controller and an RGB time sequence generator are arranged in the TDDI chip, the software application layer configures automatic picture brushing enabling and related parameters to the register block, and the storage controller automatically reads picture data in a storage module according to configuration and achieves automatic switching between a background picture and a windowing picture. The RGB timing generator generates a display timing and drives the display panel. The system can also be selectively provided with an external MCU which is only used for image switching. Through the TDDI chip master control and display integrated design, an external master control does not need to participate in storage, reading and continuous transmission of image data, and the problems that a traditional display system is high in complexity, large in power consumption, high in master control requirement and large in resource occupation are solved.
Owner:SHENZHEN AIXIESHENG TECH CO LTD

Access address configuration method, processor, multiprocessor system, medium and product

The invention discloses an access address configuration method, a processor, a multiprocessor system, a medium and a product. The method comprises the following steps: configuring an independent register block for each thread bundle in the processor; wherein the register group is used for storing coordinate information of tensor data used in the execution process of a corresponding thread bundle program; reading the coordinate information of the tensor data to be processed from the register group, and sending the read coordinate information to a coprocessor to carry out access address calculation of the corresponding tensor data; according to the method, each thread bundle is provided with a group of entity registers for storing the coordinate information of the tensor data, multiple multiplexing of the coordinate information can be supported, and a large number of instructions for specifying coordinates are saved, so that the number of assembly instructions is reduced, the instruction overhead is effectively reduced, the execution time of a processor is saved, and the processing efficiency of the processor is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Processor system and interrupt response method

The invention provides a processor system and an interrupt response method. In the processor system, an interrupt controller is used for setting the state of a control register, receiving an interrupt request signal which is sent from the outside and is used for requesting to execute an interrupt service program, setting the state of the control register to be a first state value when the interrupt request signal is received, and executing the interrupt service program after the interrupt server program is executed. Setting the state of the control register as a second state value; the multiplexer is used for reading the state value of the control register, when the read state value of the control register is a first state value, the multiplexer is communicated with the shadow register group and the arithmetic unit, and when the read state value of the control register is a second state value, the multiplexer is communicated with the main register group and the arithmetic unit; the main register group is used for storing data when the main program is operated; and the shadow register group is used for storing data when the interrupt server program runs. The delay in the interrupt response process can be reduced, and the system performance is improved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Key vaults with active register banks

Systems and techniques are provided for secure computing. For instances, a process can include generating a master private key; generating a set of first dummy values; dividing the master private key into a first set of shares, wherein a sum of shares of the first set of shares equals a value of the master private key; initiating portions of an active memory bank with a sequence of integer modular additions, wherein the integer modular additions comprise: masking the first set of shares of the master private key; and adding the masked first set of shares of the master private key to portions of the active memory bank using a sequence of adds that adds and mixes the masked first set of shares of the master private key with dummy values.
Owner:QUALCOMM INC

Shared operation method, system and equipment of network acceleration unit and medium

The invention belongs to the technical field of hardware design, and provides a sharing operation method, system and equipment of a network acceleration unit and a medium, and the method comprises the step of configuring a shared general register, an exclusive general register and an exclusive special register group for multiple hosts in a single network acceleration unit. Each host accesses the register through the low-speed bus based on the unique identifier. And the acceleration unit performs data interaction with each host through a high-speed bus according to the configuration of the register. When the host operation is processed, the management message and the data transmission message are distinguished by judging whether the basic communication resource register group is accessed or not, and high-order modification and resource mapping are carried out on a load data address in the data transmission message. When an external data packet is received, a target host is determined through analysis and address matching, and resource identifier mapping and data access are carried out. According to the invention, each host is ensured to have independent communication resources and data channels, the complexity and the cost of hardware implementation are reduced, and the resource utilization rate and the system expansibility are improved.
Owner:XIAN AVIATION COMPUTING TECH RES INST OF AVIATION IND CORP OF CHINA

Quick transposition method and device based on DMA (Direct Memory Access) hardware

The invention relates to the technical field of computers, in particular to a rapid transposition method and device based on DMA (direct memory access) hardware, which uses a group of better configuration parameters, adopts source matrix parameters and target matrix parameters to describe three-dimensional data blocks and supports multi-batch matrix transposition. The block size matched with the transmission bit width of the preset communication protocol is adopted, data transmission is flexibly and efficiently completed, and the dynamic change problem in the data transposition process can be flexibly solved; part of source data is cached in the DMA hardware in advance for subsequent rapid extraction and use; matrix transposition of source data is achieved through the register set, data reading and writing are conducted in the row direction, the problem that a memory jumps when one column of the matrix is read is avoided under the condition that few hardware resources are used, the method is flexibly suitable for various transposition types, and transposition efficiency is improved.
Owner:BEIJING FENGHUA CHUANGZHI TECHNOLOGY CO LTD

Memory system, memory controller and memory device

A memory system includes a memory device, a control circuit, and a memory controller. The memory device includes a plurality of memory regions, a candidate register set storing a plurality of candidate parameter sets corresponding to temperature ranges, respectively, a mode register set storing a temperature range and a refresh management parameter set. The control circuit stores, based on the temperature range, a first candidate parameter set of the plurality of candidate parameter sets as the refresh management parameter set in the mode register set. The memory controller obtains the refresh management parameter set from the memory device, based on an updated temperature range, counts activation counts of the plurality of memory regions, and provides a refresh management command to the memory device based on the activation counts and the refresh management parameter set.
Owner:SAMSUNG ELECTRONICS CO LTD

Programmable logic device dynamic register configuration device and control method thereof

The invention provides a programmable logic device dynamic register configuration device and a control method thereof. The configuration device for the dynamic register of the programmable logic device comprises an external controller used for providing a configuration instruction and configuration data; the main controller is used for receiving and analyzing the configuration data, generating a gating signal and recombining a data packet; at least one multiplexing module, each multiplexing module being configured to include a register block; and one or more stages of sub-controllers are arranged between the main controller and the multiplexing module and are used for further analyzing the data packet and controlling the target register group to complete read-write operation. In addition, the invention further provides a control method corresponding to the configuration device. Some embodiments of the invention aim to solve the problems of high configuration switching time consumption and uncertain configuration switching time in the prior art on the premise of not increasing too much complexity and area overhead.
Owner:SHANGHAI XINLU TECH CO LTD

An attention mechanism calculation method, device, storage medium and product

The application discloses an attention mechanism calculation method and device, a storage medium and a product. The method comprises the following steps: performing matrix multiplication operation on a query matrix block and a key matrix block to obtain a first product matrix block of a first precision type, performing exponential operation on the first product matrix block written in a second register group to obtain a numerator matrix block; writing the quantized numerator matrix block into a third register group according to a second precision type; writing the numerator matrix block into a fourth register group according to a third precision type and performing layout conversion through a shared memory; performing softmax operation on the layout-converted numerator matrix block and a value matrix block to obtain an attention result matrix block of the first precision type and write the attention result matrix block into a fifth register group; and writing the attention result matrix block into the shared memory according to the third precision type. The embodiment of the application can reduce the cumulative error of an intermediate operation result.
Owner:SHANGHAI BIREN TECH CO LTD

Circuitry to detect cycle count for increased throughput reads and write operations for memory

Control circuitry for memory includes a state machine including a number of state elements corresponding to a maximum number of available columns of a blast operation for memory; a set of registers including a corresponding register for each state element; and a column selection control circuit that combines outputs of the state machine and the set of registers to trigger an appropriate column during an appropriate clock cycle, available at a start of a corresponding clock cycle. The state machine receives a clock and various inputs associated with a start of memory operations and provides intermediate state element outputs and a final state element output as the outputs. Each register of the set of registers is available to store, from an address enable signal, a value indicating that a column in memory to which that register corresponds is to be accessed.
Owner:ARM LTD

Handshaking system, method and related device for three-dimensional convolutional neural network

The application discloses a handshake system, method and device of a three-dimensional convolutional neural network, computer equipment and a medium. The method comprises the following steps: obtaining to-be-transported data, and writing the to-be-transported data into a cache as a cache data block in a preset size; writing the size of the cache data block into a register group according to a preset storage mode to obtain data block record information; when a data operation request is received, obtaining required data information from the data operation request, wherein the required data information comprises three-dimensional coordinates; determining whether the cache data block contains data corresponding to the required data information based on the three-dimensional coordinates and the data block record information; and if the cache data block contains the data corresponding to the required data information, generating a response signal indicating that there is valid data in the cache data block. The handshake efficiency is improved by using the application.
Owner:SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

Clock peak-shaving control device, control method and chip

This application provides a clock peak shifting control device, control method, and chip. The device includes a first selection module, a first delay module, a second selection module, and a first latch module. The first delay module generates a target delay clock signal. The first selection module selects between a test clock signal and the target delay clock signal. The second selection module switches between different operating modes of the chip to obtain a drive clock signal. The latch module controls the timing of different groups of registers. The drive clock signal provides a clock signal for each register in each register group, ensuring that registers in each group do not flip at the same time. This reasonable peak shifting control of register flipping time effectively reduces the peak power consumption during the shift process.
Owner:KTMICRO ELECTRONICS

Validation system based on shared register address

The present application relates to the technical field of chip verification, and more particularly to a verification system based on a shared register address, comprising: a design under test including a plurality of register groups; a register model for generating register access request information and sending it to an adapter; the adapter for decoding a target logical address, obtaining target PFVF information and a target register address, and generating target transaction information and sending it to a bus agent module; the bus agent module for driving the design under test based on the target transaction information; the design under test is also used to determine a target register group based on the target register address, determine a target register in the target register group based on the target PFVF information, and then execute target access information based on the target register. The present application can be applied to verification of a shared register address, simplifies the verification environment, and improves the efficiency and accuracy of verification.
Owner:沐曦集成电路(南京)有限公司

Real-time debug in low-power devices

ActiveUS20260063704A1Digital circuit testingComputer architectureCritical signal
According to an embodiment, a method for debugging a low-power domain in a multi-power domain device is disclosed. The method includes mapping a critical signal of the low-power domain to a latch circuit of a register bank; capturing, by the latch circuit, transitions of the critical signal in registers of the register bank; observing, by a core circuit in a switchable power domain of the multi-power domain device, registers of the register bank; and determining a fault corresponding to the critical signal based on a value of the registers of the register bank.
Owner:STMICROELECTRONICS INT NV

System with a Communication Interface and a Restart Method Thereof

An example method for a primary device to restart a secondary device through a communication interface. The secondary device includes a register table. The method includes performing, by the primary device and through the communication interface, a write transaction on the secondary device for a predetermined address in the register table. The method includes performing, by the primary device and through the communication interface, a read transaction on the secondary device for reading a predetermined register group in the register table to cause the secondary device to restart after completing the reading of the predetermined register group.
Owner:ARTILUX INC

Data processing device, data processing method, chip and electronic equipment

This application discloses a data processing apparatus, a data processing method, a chip, and an electronic device, relating to the field of data processing technology. The data processing apparatus includes: a storage module configured to store x sets of original data vectors, each original data vector containing m data elements, where x and m are positive integers; and a computation module including a single instruction multiple data (SIMD) execution unit and a vector register group, the SIMD execution unit being coupled to the vector register group, and the vector register group being coupled to the storage module. The computation module is configured to: determine whether x is an integer multiple of the SIMD processing capability, split x, and obtain the top-k data elements in the original data vectors based on different data strategies. The algorithm of this application is flexible, requires no dedicated hardware sorting network, avoids dependence on dedicated hardware, and improves programming flexibility and versatility, facilitating the expansion and migration of large models across different hardware platforms.
Owner:BEIJING TSINGMICRO INTELLIGENT TECH CO LTD

A signal processing combination performance checking system based on DSP communication test

PendingCN122332233APathPingData stream
The application discloses a signal processing combination performance verification system based on DSP communication test, relates to the real-time test technical field of a digital signal processor, and specifically comprises: a homomorphism copying module, a slot injection module and a performance judgment module; the dynamic configuration parameters of a current DSP operation core are captured in real time through hardware sniffing of a bus, and the dynamic configuration parameters are mirrored to shadow registers; a peek unit is used to identify the bubble period of a pipeline through prediction logic; at the moment when the main computing unit is in a waiting period, the operation of the shadow register group is enabled by hardware, a preset characteristic test vector is input into a computing engine, and a shadow operation is completed; the real-time data flow of a main path and the reference data flow of a shadow path are compared in time domain and frequency domain by a hardware comparison array, the second derivative of a response time delay is calculated, a comprehensive performance factor is output, online verification of the combination performance of a DSP signal processing is realized, fault types can be accurately distinguished, and the reliability and maintainability of the system are improved.
Owner:BEIJING SHIJICHEN DATA TECH CO LTD

A register access system based on a common address

This invention relates to the field of chip design technology, and more particularly to a register access system based on a shared address, comprising a host subsystem and a chip subsystem. The host subsystem includes one physical functional module and N virtual functional modules. The chip subsystem includes a protocol conversion module, a register on-chip network, and M IP modules. Each IP module corresponds to a register address space, each register address space contains one or more register addresses, each register address corresponds to a register group, and each register group includes N+1 registers. Each of the N+1 registers in each register group corresponds one-to-one with the physical functional module and the N virtual functional modules. The host subsystem sends a register access request to the protocol conversion module through the physical functional module or the virtual functional module to access the target register. This invention can achieve effective isolation and flexible control between physical and virtual functions while ensuring efficient use of address resources.
Owner:沐曦集成电路(南京)有限公司

Attention computation implementation method and device, medium, equipment and product

The application discloses an attention calculation implementation method and device, medium, equipment and product, and the method comprises the following steps: by multiplexing a first register group, sequentially loading N1 query subblocks of the ith query block from the shared memory to the first register group, and based on a matrix multiplication instruction, calculating the query subblock obtained each time and the corresponding key subblock in the jth key block to obtain the jth attention score block; sequentially obtaining M K attention score blocks, and performing mixed-precision attention fusion calculation on the corresponding value block to obtain the corresponding attention output block, and then performing target precision type conversion and writing back to the shared memory on the attention output block by batch multiplexing the specified register group storing the target precision type data in the fusion calculation process. The application can significantly reduce the register overhead, support large-scale attention calculation, reduce memory access delay and significantly improve the overall calculation performance.
Owner:SHANGHAI BIREN TECH CO LTD

Memory system, memory controller, and memory device

A memory system, a memory controller, and a memory device are provided. The memory system includes a memory device, a control circuit, and a memory controller. The memory device includes a plurality of memory regions, a candidate register group storing a plurality of candidate parameter sets corresponding to temperature ranges, respectively, and a mode register group storing the temperature ranges and refresh management parameter sets. The control circuit stores a first candidate parameter set of the plurality of candidate parameter sets as a refresh management parameter set in a mode register set based on a temperature range. The memory controller obtains a refresh management parameter set from the memory device based on the updated temperature range, counts an activation count of the plurality of memory regions, and provides a refresh management command to the memory device based on the activation count and the refresh management parameter set.
Owner:SAMSUNG ELECTRONICS CO LTD

Chip optimization design method, chip, electronic equipment and storage medium

The application provides a chip optimization design method, a chip, an electronic device and a computer readable storage medium. The chip optimization design method comprises: determining a register combination with similarity according to register array information of a chip; wherein the register combination with similarity comprises a plurality of registers, and the plurality of registers have the same fan-in information and / or fan-out information; and optimizing design of a register array of the chip according to the register combination with similarity. In this way, the problem that a large amount of manpower is consumed and a long time is consumed when the chip is currently optimized and designed can be improved.
Owner:HYGON INFORMATION TECH CO LTD