Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

700results about "Register arrangements" patented technology

Universal measurement and control system data flow control bus architecture method and system

The invention provides a universal measurement and control system data flow control bus architecture method and system, and relates to the technical field of measurement and control, and the method comprises the steps: distributing independent address spaces for a plurality of driving modules through a main control module, building a parameter mapping table and a buffer region, and enabling a plurality of test channels to share an interrupt signal and store the interrupt signal in an interrupt vector register; generating an interrupt signal based on a relationship between the buffer data volume and a threshold; sequencing according to the channel priority reference value, and dynamically adjusting the priority based on the load rate; and the main control module determines an interrupt source according to the interrupt vector and dynamically adjusts a prefetching strategy according to the measurement and control data flow. The system data processing efficiency and the resource utilization rate are improved.
Owner:BEIJING TIANCHEN HECHUANG TECH CO LTD

Matrix multiplication and accumulation operation unit and operation method, hardware accelerator and electronic equipment

The embodiment of the invention provides a matrix multiplication and accumulation operation unit and method, a hardware accelerator and electronic equipment, and the matrix multiplication and accumulation operation unit comprises a data loading storage engine, a tensor register file and a matrix multiplication engine. The data loading and storage engine is used for loading data of a plurality of matrixes to be subjected to matrix multiplication and accumulation calculation; the tensor register file is used for storing data of a plurality of matrixes acquired from the data loading storage engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprises a plurality of tensor registers, and different tensor register groups are used for storing data of different matrixes in the plurality of matrixes; and the matrix multiplication engine is used for carrying out matrix multiplication accumulation calculation based on the data of the plurality of matrixes stored in the tensor register file. According to the embodiment of the invention, more efficient MMA calculation is realized under the conditions of low cost, low power consumption and less occupied space.
Owner:ALIBABA (CHINA) CO LTD

Instruction execution method, processor, electronic equipment and storage medium

The invention provides an instruction execution method, a processor, electronic equipment and a storage medium, the method is applied to the processor, the processor comprises a first execution unit, a second execution unit and a pipeline register, and a first execution instruction of an atomic instruction group is obtained through the first execution unit; writing a first execution result of the first execution instruction into the pipeline register in response to the first execution instruction; a second execution instruction of the atomic instruction group is obtained through the second execution unit, and the first execution result is read from the pipeline register as the source operand of the second execution instruction in response to the second execution instruction, so that the instruction execution method, the processor, the electronic equipment and the storage medium can reduce the access frequency of the general register and improve the access efficiency of the general register. Processor power consumption and register port conflicts can be reduced.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Register resource management method and device, equipment and storage medium

The invention relates to the technical field of artificial intelligence chips, and provides a register resource management method and device, equipment and a storage medium, and the method comprises the steps: obtaining register resource information needed by an operator; based on the register resource information, target register addresses are searched from register resource pools of all the thread bundles, and the target register addresses are register addresses which are the same in address and are not used in the register resource pools of all the thread bundles; and allocating the target register address to the operator, so that each thread bundle shares the target register address when executing the operator. According to the method, through an intelligent register resource searching and distributing mechanism, independent application in the processing logic of each thread bundle is not needed, and only one-time application outside the logic of the thread bundle is needed, so that the requirements of different logic among multiple thread bundles and application of same address register resources can be met, and the expense of register resource application is reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Optimization method for sharing one group of physical registers by multiple groups of logic registers

The invention provides an optimization method for sharing one group of physical registers by multiple groups of logic registers, which comprises the following steps: S1, defining a shared physical register file, the step comprises a physical register structure link and a quantity constraint link, the physical register structure link comprises N vector physical registers with VLEN bit width, and the quantity constraint link comprises N vector physical registers with VLEN bit width; each vector physical register can be divided into VLEN / FLEN floating point physical registers; in order to solve the problem of hardware redundancy caused by independence of a floating point register and a vector register in an RISC-V architecture, instruction pipeline sharing of floating point and vector expansion is realized through a scheme that multiple groups of logic registers share a physical register, and the hardware redundancy is avoided. And hardware overhead is reduced.
Owner:BEIJING YIHUA CLOUD NETWORK TECH CO LTD

Processor program address buffering method and device

The invention provides a processor program address buffering method and device, and belongs to the technical field of processors, and the processor program address buffering device comprises the following components: an instruction fetching unit, a program address comparison unit, a program address buffer area, a program address buffer area tail value register, an instruction transmitting and executing unit and a reordering buffer area; the instruction fetching unit is in control connection with the program address comparison unit, and the program address comparison unit is in control connection with the program address buffer area, the program address buffer area tail value register, the instruction transmitting and executing unit and the reordering buffer area. In order to solve the problem that the disadvantage of the method is obvious in design iteration of instruction fetching and decoding width increasing and assembly line stage increasing of the superscalar processor, the invention provides a judgment basis for whether a program address is written into a program address buffer area or not through a program address comparison unit; the program addresses of all instructions are not written into the program address buffer area, so that the number of implementation table entries of the program address buffer area can be smaller.
Owner:SHANGHAI YIHUA TECHNOLOGY CO LTD

Data processing device and method, processor and chip

The invention relates to a data processing device and method, a processor and a chip, and relates to the technical field of computers, the data processing device comprises a processing module and a driving module; the processing module obtains address information of each thread in a first instruction when judging that the instruction to be processed is the first instruction, the first instruction is a multi-thread parallel loading instruction, and the address information of different threads is mutually independent; the driving module determines a parallel loading result according to address information of each thread in the first instruction, the address information is used for determining an off-chip address and an on-chip address of each thread, and the thread is used for loading data at a position indicated by the off-chip address to a position indicated by the on-chip address. According to the embodiment of the invention, multi-thread parallel execution of different data accesses can be realized, and the calculation efficiency is remarkably improved.
Owner:MOORE THREADS TECH CO LTD

Heterogeneous processor-oriented reciprocal calculation instruction sequence generation method

The invention discloses a reciprocal calculation instruction sequence generation method oriented to a heterogeneous processor, and belongs to the field of compilation optimization and code generation. Aiming at the problems of instruction redundancy, weak precision control, poor hardware adaptation and high manual dependence of an existing method in a heterogeneous environment, characteristics of a reciprocal instruction and an operand are accurately identified by linearly scanning heterogeneous object codes (including vectorization, scalar and complex instruction sequences); in combination with hardware characteristics of RISC / SIMD / VLIW / DSP and the like, a multi-round iteration precision improvement and temporary register optimization allocation strategy is adopted, differential generation logic is formulated, and a high-precision low-redundancy instruction sequence is generated. The method comprises linear code scanning classification, reciprocal instruction and operand identification, cross-architecture generation logic rule formulation, instruction sequence generation and legality verification. Full-process automation is achieved, manual intervention is reduced, the execution efficiency and precision of reciprocal calculation of the heterogeneous processor are improved, and the method is suitable for embedded systems, high-performance calculation and other scenes.
Owner:HUNAN UNIV OF SCI & TECH

Asynchronous multi-granularity memory access control method in microprocessor, asynchronous circuit and memory access module

The invention discloses an asynchronous multi-granularity memory access control method in a microprocessor, an asynchronous circuit and a memory access module, the memory access control method adopts asynchronous clock-free control, introduces a multi-granularity memory access control mechanism, and divides the memory access granularity into two classes of multi-word memory access and non-multi-word memory access for processing. According to the method, the access path and the cache strategy can be flexibly adjusted according to different access granularities, redundant operation during access is reduced, and the access efficiency is greatly improved. The asynchronous circuit is a delay-limited asynchronous circuit and comprises an asynchronous controller and an asynchronous micro-pipeline structure based on a'sending-relay-receiving 'structure. The memory access module based on the asynchronous circuit comprises a memory access state management module, a data access updating module, an instruction / data cache and peripheral interaction interface, and a data interaction interface between the memory access module and the write-back module. According to the invention, lower dynamic power consumption and higher energy efficiency ratio are realized, and high efficiency and low delay are realized under large data volume operation.
Owner:LANZHOU UNIV

Matrix multiply accumulation operation unit and operation method, and hardware accelerator and electronic device

Provided in the embodiments of the present disclosure are a matrix multiply accumulation (MMA) operation unit and operation method, and a hardware accelerator and an electronic device. The MMA operation unit comprises a data load / store engine, a tensor register file and a matrix multiplication engine, wherein the data load / store engine is used for loading data of a plurality of matrixes to be subjected to MMA computation; the tensor register file is used for storing the data of the plurality of matrixes that is acquired from the data load / store engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprising a plurality of tensor registers, and different tensor register groups being used for storing data of different matrixes among the plurality of matrixes; and the matrix multiplication engine is used for performing MMA computation on the basis of the data of the plurality of matrixes that is stored in the tensor register file. By means of the embodiments of the present disclosure, more efficient MMA computation is realized with low cost, low power consumption and a smaller footprint.
Owner:ALIBABA (CHINA) CO LTD

Padding and suppressing rows and columns of data

A method is described herein. The method generally includes receiving stream parameters that defines an array, wherein the stream parameters include a first null element count and a second null element count. The method generally includes forming a stream of vectors for the multidimensional array responsive to the stream parameters. The stream of vectors generally includes a vector of null elements at a beginning of the stream of vectors based on the first null element count. The stream of vectors generally includes a null element at a beginning of each vector of the stream of vectors based on the second null element count. The stream of vectors generally includes a set of data distributed across a subset of the stream of vectors. The method generally includes providing the stream of vectors.
Owner:TEXAS INSTRUMENTS INC

Processor core, processor, and method for processor

The embodiment of the invention provides a processor core, a processor and a method for the processor. The processor core includes a processing pipeline configured to rename and execute a first instruction including at least one of a first architectural register and a second architectural register, and includes: a first type of physical register configured to be mapped by the first architectural register and configured to store a first type of data; the first type of physical register is configured to be mapped by the first architecture register and configured to store data of a first type, the second type of physical register is configured to be mapped by the second architecture register and configured to store data of a second type, the first architecture register comprises a vector architecture register, the data of the first type stored by the first type of physical register comprises vector data, and the second architecture register comprises a mask architecture register; the second type of data stored in the second type of physical register comprises mask data. The processor core may mitigate processing pipeline stagnation and increase less area.
Owner:HYGON INFORMATION TECH CO LTD

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Register overflow optimization method and device and storage medium

The embodiment of the invention provides a register overflow optimization method and device and a storage medium, and is applied to the technical field of chips. In the method, for each virtual register in a target program, based on a physical register type supported by an instruction operand where the virtual register is located, the virtual register is optimized; selecting a corresponding target register class from N candidate register classes, wherein N is greater than 1; allocating a first physical register in the target register class to the virtual register; when register overflow occurs, a target register is selected from the allocated first physical registers, the instruction operand stored in the target register overflows to the second physical registers in the other N-1 candidate register classes, and compared with the mode that the instruction operand overflows to the memory to generate read-write operation on the memory, the instruction operand stored in the target register overflows to the second physical registers in the other N-1 candidate register classes; according to the method, different types of physical registers are overflowed, read-write operation aiming at the physical registers is generated, the pressure of the registers is relieved, the performance overhead of the overflowed memories is reduced, and the register distribution efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Automatic illegal character cleaning system

The invention relates to the technical field of computer data processing and network security, in particular to an illegal character automatic cleaning system, which comprises a rule compiling module for monitoring rule change, performing semantic fusion and topological mapping on a rule set, constructing a deterministic finite automaton and mapping the deterministic finite automaton into a state transition table; the state switching module is used for constructing a double-buffer context and operating lock-free switching through an atomic pointer to realize hot updating; the speculation execution module is used for carrying out vectorization pre-scanning based on a state transition table by utilizing single-instruction multi-data stream parallel loading, identifying a walk path, falling into a safe state for releasing and falling into a trap state for triggering external verification; the self-adaptive feedback module is used for counting trap state triggering frequency, generating a rule allergy report and dynamically adjusting the size of a read fragment; according to the method, the contradiction between rule flexibility and execution efficiency is solved, and high-performance cleaning based on speculative execution is realized.
Owner:北京啄木鸟云健康科技有限公司

Enabling high-performance scalable matrix extension (SME) instruction issue in processor devices

Enabling high-performance Scalable Matrix Extension (SME) instruction issue in processor devices is disclosed herein. In some aspects, a processor device comprises a reservation station circuit configured to perform, during a first phase, a reduced-precision vector accumulator (ZA) tracking operation on micro-ops for which corresponding vector (Z) registers and corresponding predicate (P) registers are ready. Based on the reduced-precision ZA tracking operation, the reservation station circuit selects a first micro-op and a second micro-op having no Read-After-Write (RAW) hazard with respect to the ZA registers. During a subsequent second phase, the reservation station circuit performs a full-precision ZA tracking operation on the first micro-op and the second micro-op, and selects one as a micro-op for issue for which the full-precision ZA tracking operation indicates no RAW hazard exists with respect to the ZA registers. The reservation station circuit then issues the selected micro-op for execution.
Owner:QUALCOMM INC

Special instruction set processor for polar code coding and decoding algorithm and implementation method

The invention relates to the technical field of communication and computer instruction execution and processing, in particular to a polar code encoding and decoding algorithm-oriented special instruction set processor and an implementation method, and a vector processor has efficient parallel computing capability. A coding and decoding algorithm of a polar code relates to parallel computing of a large number of log-likelihood ratios, a vector processor of a special instruction set can process a plurality of LLR values at the same time through a parallel processing unit, and the throughput rate is remarkably improved; special instructions such as polarization shuffling and vector G operation instructions are designed for polarization code encoding and decoding and are used for accelerating algorithm operation. The method can achieve the balance among the performance, the energy efficiency and the flexibility, can adapt to the continuous iteration upgrading of the algorithm due to the programmable and reusable characteristics, and has a good application prospect.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Optimization method and device for protocol calculation, computer equipment, readable storage medium and program product

The invention relates to a protocol calculation optimization method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps that address mapping processing is conducted on a target register, memory mapping addresses corresponding to all threads in the target register are determined, a target initialization value is determined according to the operation type of protocol calculation, then data stored in the memory mapping addresses corresponding to all the threads are initialized into the target initialization value, protocol calculation is executed, and the target initialization value is stored in the memory mapping addresses corresponding to all the threads. And obtaining a protocol calculation result. By adopting the method, the calculation efficiency of protocol calculation can be improved.
Owner:SHANGHAI BIREN TECH CO LTD

RISC-V data processing system and data processing method

The embodiment of the invention discloses a data processing system and a data processing method based on RISC-V. The data processing system comprises an instruction coding space and a reconfigurable execution unit, the instruction coding space comprises an operation code area, a reserved area and an expansion area; the instruction coding space is used for defining a basic operation to be executed by an instruction in the operation code area, defining an extended operation to be executed by the instruction in the reserved area and defining the length of the instruction in the extended area so as to configure different types of instructions; the reconfigurable execution unit is used for constructing a data path matrix of the instruction based on the configuration information of the instruction; transmitting data required when the instruction is executed based on the data path matrix; and executing an operation corresponding to the instruction based on the data of the instruction.
Owner:CHENGDU KAIYUAN COMPUTING ECOLOGICAL TECHNOLOGY CO LTD

Accelerator architecture for near IO pipeline computing, and ai acceleration system

The present application relates to the field of accelerators, and discloses an accelerator architecture for near IO pipeline computing, and an AI acceleration system. A multi-channel direct access module comprises N DRAM controllers, and the N DRAM controllers are connected to DRAMs in a one-to-one correspondence; each DRAM controller is at least connected to k DMA controllers, and is connected to a corresponding core group cluster by means of the DMA controllers; a pipeline synchronization ring is connected to N core group clusters and comprises M cascaded forward transmission blocks and M cascaded backward transmission blocks, and the head-end forward transmission block and the tail-end forward transmission block are respectively connected to a data receiving module and a data transmitting module; the output of the tail-end forward transmission block is cascaded to the first backward transmission block; the ith forward transmission block and the (M-i)th backward transmission block correspond to each other in a front-rear direction, and are jointly connected to at least one computing core group. The accelerator of the architecture replaces a traditional multi-level cache structure, reduces the time delay caused multi-layer search, and accelerates the computing rate between the interior of the accelerator and the accelerator.
Owner:STORAGEX TECHNOLOGY INC

Scan chain optimization method, computer equipment, storage medium and program product

The invention provides a scan chain optimization method, computer equipment, a storage medium and a program product. The method comprises the following steps: determining a plurality of initial scan chains to be optimized; mixing the scan registers of the plurality of initial scan chains to obtain a scan register set; sorting the scan registers in the scan register set based on a scan chain wiring length to obtain a long-chain scan chain; cutting the long-chain scanning chain to obtain a plurality of scanning chain segments in one-to-one correspondence with the plurality of initial scanning chains; sorting the plurality of scanning chain segments according to the cutting sequence, and determining whether the scanning chain bit width corresponding to each scanning chain segment meets the bit width requirement or not; in response to the situation that the scanning chain bit width corresponding to any scanning chain segment does not meet the bit width requirement, the scanning register is moved between the adjacent scanning chain segments until the scanning chain bit widths corresponding to all the scanning chain segments meet the bit width requirement, and multiple optimized target scanning chains are obtained.
Owner:X TIMES DESIGN AUTOMATION CO LTD

High-efficiency INT6 quantification method, device and equipment for large language model

The invention discloses an efficient INT6 quantification method, device and equipment for a large language model, and the method comprises the steps: carrying out the mixing precision quantification of the large language model, and obtaining a quantized large language model; performing bit-level data packaging on the weight and the activation value in the quantized large language model to obtain bit-level data; loading the bit-level data to a register of the GPU, and carrying out matrix product accumulation operation and weighted summation by utilizing BTC to obtain output data; and storing the output data back to a global memory of the GPU so as to complete the quantitative reasoning process of the large language model. According to the method, mixed precision quantification is carried out on a large language model by utilizing different precisions, the reasoning speed is improved through a scheduling strategy of the GPU while relatively high quantification precision is achieved, and the reasoning potential of the GPU is fully mined, so that all potential of 6-bit quantification is released.
Owner:XIDIAN UNIV

Matrix multiplication operation method, matrix transposition operation method and device of processor

The invention provides a matrix multiplication operation method and device and a matrix transposition operation method and device of a processor. The matrix multiplication operation method of the processor comprises the following steps: taking a first matrix of which the storage sequence is inconsistent with the loading sequence for carrying out element-by-element multiplication operation as a second matrix, loading the second matrix through a register set in the processor, and dividing the second matrix into a plurality of third matrixes, replacing elements in a plurality of nth register groups included in the register set; for each nth register group, reading elements in the nth register group according to an alternating sequence of the elements of the first register and the elements of the second register through the (n + 1) th register group; when iteration is carried out for n to N-1, the obtained sixth matrix is stored in the memory; and performing element-by-element multiplication and accumulation on the elements at each position of the first matrix and the sixth matrix. According to the invention, the time required by matrix multiplication operation can be reduced.
Owner:SHANGHAI ORIENTAL COMPUTER TECHNOLOGY CO LTD

SoC system control method and device, electronic equipment and readable storage medium

The invention provides an SoC system control method and device, electronic equipment and a readable storage medium, and the method comprises the steps: controlling a CPU large core to generate a control instruction, and writing the control instruction into a register; the control register analyzes the control instruction, determines at least one target function module associated with the control instruction, generates a corresponding enable signal and sends the enable signal to a control circuit corresponding to each target function module; and aiming at the control circuit of each target function module, controlling the control circuit to receive and analyze the corresponding enable signal, generating control information and sending the control information to the corresponding target function module, so that each target function module is configured according to the control information and executes the control instruction. Therefore, the configuration time of the CPU for configuring the functional modules one by one through the on-chip internet can be reduced, and the control efficiency of the SoC system is further improved.
Owner:CIX TECH (SHANGHAI) CO LTD

Vector extract and merge instruction

There is provide an apparatus, method and medium. The apparatus comprises decoder circuitry to generate control signals in response to a vector extract and merge instruction specifying a control parameter, a first vector register, a second vector register, and a destination vector register. The apparatus comprises processing circuitry responsive to the control signals, to perform plural beats of processing, each beat comprising processing corresponding to a portion of at least the first vector register and the destination vector register. The processing, for a Kth beat comprises: extracting bits, specified by the control parameter, from a Kth portion of the first vector register, concatenating the bits with further bits, and storing the result in the Kth portion of the destination register. The further bits are, for a first portion, extracted from a first portion of the second vector register and, otherwise, from a (K−1)th portion of the first vector register.
Owner:ARM LTD

Integrated circuit time sequence optimization method and device, electronic equipment and storage medium

The invention provides an integrated circuit time sequence optimization method and device, electronic equipment and a storage medium, and is applied to the technical field of computers.The integrated circuit time sequence optimization method comprises the steps that after registers of an integrated circuit are extracted, a target register set is determined based on the registers of the integrated circuit; the target register set comprises at least two target registers, and a time sequence path exists between any two target registers, counting target parameters related to time sequence performance on the time sequence path to which each target register belongs, and adjusting the time sequence path to which each target register belongs based on the target parameters, the time sequence path optimization is carried out on the target registers in the target register set, the overall workload of the optimization design can be reduced, the optimization process is carried out on the target registers where time sequence violation is prone to occurring, the optimization effect is more obvious, and therefore the time sequence optimization efficiency is improved on the whole, and the integrated circuit design period is shortened.
Owner:PHYTIUM TECH CO LTD

Computing device and method based on RISC-V extension instruction

The invention provides a computing device and method based on RISC-V extension instructions, the computing device supports approximate computation of mixed precision according to approximate computation instructions in an extended approximate computation instruction set, and the computing device comprises an out-of-order scheduling and register reading module used for executing instruction dependency analysis and operand preloading, scheduling the non-approximate calculation instruction and the approximate calculation instruction to different transmitting queues respectively; the first instruction transmitting queue is used for temporarily storing a to-be-transmitted non-approximate calculation instruction; the second instruction transmitting queue is used for temporarily storing approximate calculation instructions to be transmitted; the precise calculation module is used for completing precise calculation related tasks according to the instruction from the first instruction transmitting queue; and the approximate calculation module is used for completing approximate calculation related tasks according to the instructions from the second instruction transmitting queue, and supports approximate calculation of various precisions.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Matrix multiplication in dynamically spatially and dynamically temporally divisible architectures

A data processing apparatus includes a first vector register and a second vector register, both of which are dynamically spatially and dynamically temporally divisible. A decode circuit receives one or more matrix multiplication instructions indicating a set of first elements in the first vector register and a set of second elements in the second vector register, and generates a matrix multiplication operation in response to receiving the matrix multiplication instructions. The matrix multiplication operation causes one or more execution units to perform a matrix multiplication of the set of first elements and the set of second elements, and an average bit width of the first elements is different from an average bit width of the second elements.
Owner:ARM LTD