Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1066results about "Register arrangements" patented technology

Optimization method for improving parallelism and performance of RISC-V vector instruction executed by hardware

The invention provides an optimization method for improving parallelism and performance of hardware executing RISC-V vector instructions, which comprises the following steps: S1, an instruction fetching and decoding stage step: the step comprises an instruction fetching link and a decoding link, and the instruction fetching link reads instructions from a memory according to a program sequence and stores the instructions into an instruction queue; the decoding link comprises the steps of analyzing an instruction, identifying whether the instruction is a vector mask instruction or a vector length control instruction, and if the instruction needs an old value of a destination register, marking that the instruction needs to carry an old value dependency identifier, belonging to the technical field of optimization methods. According to the method, the instruction is transmitted without waiting for the old value of the target vector register to be ready, the instruction can be transmitted in advance after a certain condition is met, and the correct old value is obtained, so that the aim of improving the parallelism and the performance of hardware for executing the RISC-V vector instruction is fulfilled.
Owner:BEIJING YIHUA CLOUD NETWORK TECH CO LTD

Universal measurement and control system data flow control bus architecture method and system

The invention provides a universal measurement and control system data flow control bus architecture method and system, and relates to the technical field of measurement and control, and the method comprises the steps: distributing independent address spaces for a plurality of driving modules through a main control module, building a parameter mapping table and a buffer region, and enabling a plurality of test channels to share an interrupt signal and store the interrupt signal in an interrupt vector register; generating an interrupt signal based on a relationship between the buffer data volume and a threshold; sequencing according to the channel priority reference value, and dynamically adjusting the priority based on the load rate; and the main control module determines an interrupt source according to the interrupt vector and dynamically adjusts a prefetching strategy according to the measurement and control data flow. The system data processing efficiency and the resource utilization rate are improved.
Owner:BEIJING TIANCHEN HECHUANG TECH CO LTD

Processor, graphics card, equipment, resource allocation method and device

The invention discloses a processor, a graphics card, equipment and a resource allocation method and device, and belongs to the technical field of computer resource management. The processor comprises a thread group assembling unit, a thread group scheduling execution unit and at least two processing units, the thread group assembling unit is used for assembling the tasks into one or more thread groups based on the attribute information of the tasks in a task assembling stage; and allocating computing resources for one or more thread groups with the thread group granularity, wherein each thread group is allocated with a processing unit; the thread group scheduling execution unit is used for scheduling one or more thread groups to the processing unit corresponding to each thread group for execution according to a unified scheduling rule, and the unified scheduling rule is a scheduling rule shared by different types of tasks; and the at least two processing units are used for executing the scheduled thread groups. And the computing resources are allocated according to the thread group granularity, so that the execution efficiency of the processor is improved.
Owner:MOORE THREADS TECH CO LTD

Matrix multiplication and accumulation operation unit and operation method, hardware accelerator and electronic equipment

The embodiment of the invention provides a matrix multiplication and accumulation operation unit and method, a hardware accelerator and electronic equipment, and the matrix multiplication and accumulation operation unit comprises a data loading storage engine, a tensor register file and a matrix multiplication engine. The data loading and storage engine is used for loading data of a plurality of matrixes to be subjected to matrix multiplication and accumulation calculation; the tensor register file is used for storing data of a plurality of matrixes acquired from the data loading storage engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprises a plurality of tensor registers, and different tensor register groups are used for storing data of different matrixes in the plurality of matrixes; and the matrix multiplication engine is used for carrying out matrix multiplication accumulation calculation based on the data of the plurality of matrixes stored in the tensor register file. According to the embodiment of the invention, more efficient MMA calculation is realized under the conditions of low cost, low power consumption and less occupied space.
Owner:ALIBABA (CHINA) CO LTD

Instruction execution method, processor, electronic equipment and storage medium

The invention provides an instruction execution method, a processor, electronic equipment and a storage medium, the method is applied to the processor, the processor comprises a first execution unit, a second execution unit and a pipeline register, and a first execution instruction of an atomic instruction group is obtained through the first execution unit; writing a first execution result of the first execution instruction into the pipeline register in response to the first execution instruction; a second execution instruction of the atomic instruction group is obtained through the second execution unit, and the first execution result is read from the pipeline register as the source operand of the second execution instruction in response to the second execution instruction, so that the instruction execution method, the processor, the electronic equipment and the storage medium can reduce the access frequency of the general register and improve the access efficiency of the general register. Processor power consumption and register port conflicts can be reduced.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Register resource management method and device, equipment and storage medium

The invention relates to the technical field of artificial intelligence chips, and provides a register resource management method and device, equipment and a storage medium, and the method comprises the steps: obtaining register resource information needed by an operator; based on the register resource information, target register addresses are searched from register resource pools of all the thread bundles, and the target register addresses are register addresses which are the same in address and are not used in the register resource pools of all the thread bundles; and allocating the target register address to the operator, so that each thread bundle shares the target register address when executing the operator. According to the method, through an intelligent register resource searching and distributing mechanism, independent application in the processing logic of each thread bundle is not needed, and only one-time application outside the logic of the thread bundle is needed, so that the requirements of different logic among multiple thread bundles and application of same address register resources can be met, and the expense of register resource application is reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Optimization method for sharing one group of physical registers by multiple groups of logic registers

The invention provides an optimization method for sharing one group of physical registers by multiple groups of logic registers, which comprises the following steps: S1, defining a shared physical register file, the step comprises a physical register structure link and a quantity constraint link, the physical register structure link comprises N vector physical registers with VLEN bit width, and the quantity constraint link comprises N vector physical registers with VLEN bit width; each vector physical register can be divided into VLEN / FLEN floating point physical registers; in order to solve the problem of hardware redundancy caused by independence of a floating point register and a vector register in an RISC-V architecture, instruction pipeline sharing of floating point and vector expansion is realized through a scheme that multiple groups of logic registers share a physical register, and the hardware redundancy is avoided. And hardware overhead is reduced.
Owner:BEIJING YIHUA CLOUD NETWORK TECH CO LTD

Binary translation system and method applied to X86 program

The invention provides a binary translation system applied to an X86 program, which is used for translating a source program following X86 semantics into a target program following other semantics, and comprises a data acquisition module used for acquiring the source program; the disassembling module is used for dividing a source program into a plurality of basic blocks and analyzing subsequent basic blocks corresponding to each basic block; the backward data flow analysis module is used for sequentially analyzing each instruction in each basic block from back to front so as to obtain a target definition set and a target subsequent use set corresponding to each instruction in the source program; and the translation module is used for eliminating redundant instructions of high-order zero clearing or high-order retention of the general register generated in translation. According to the technical scheme, the register state of the general register corresponding to each instruction of the source program is analyzed through the backward data flow analysis module so as to eliminate redundant instructions generated in the translation process.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Processor program address buffering method and device

The invention provides a processor program address buffering method and device, and belongs to the technical field of processors, and the processor program address buffering device comprises the following components: an instruction fetching unit, a program address comparison unit, a program address buffer area, a program address buffer area tail value register, an instruction transmitting and executing unit and a reordering buffer area; the instruction fetching unit is in control connection with the program address comparison unit, and the program address comparison unit is in control connection with the program address buffer area, the program address buffer area tail value register, the instruction transmitting and executing unit and the reordering buffer area. In order to solve the problem that the disadvantage of the method is obvious in design iteration of instruction fetching and decoding width increasing and assembly line stage increasing of the superscalar processor, the invention provides a judgment basis for whether a program address is written into a program address buffer area or not through a program address comparison unit; the program addresses of all instructions are not written into the program address buffer area, so that the number of implementation table entries of the program address buffer area can be smaller.
Owner:SHANGHAI YIHUA TECHNOLOGY CO LTD

Data processing device and method, processor and chip

The invention relates to a data processing device and method, a processor and a chip, and relates to the technical field of computers, the data processing device comprises a processing module and a driving module; the processing module obtains address information of each thread in a first instruction when judging that the instruction to be processed is the first instruction, the first instruction is a multi-thread parallel loading instruction, and the address information of different threads is mutually independent; the driving module determines a parallel loading result according to address information of each thread in the first instruction, the address information is used for determining an off-chip address and an on-chip address of each thread, and the thread is used for loading data at a position indicated by the off-chip address to a position indicated by the on-chip address. According to the embodiment of the invention, multi-thread parallel execution of different data accesses can be realized, and the calculation efficiency is remarkably improved.
Owner:MOORE THREADS TECH CO LTD

Heterogeneous processor-oriented reciprocal calculation instruction sequence generation method

The invention discloses a reciprocal calculation instruction sequence generation method oriented to a heterogeneous processor, and belongs to the field of compilation optimization and code generation. Aiming at the problems of instruction redundancy, weak precision control, poor hardware adaptation and high manual dependence of an existing method in a heterogeneous environment, characteristics of a reciprocal instruction and an operand are accurately identified by linearly scanning heterogeneous object codes (including vectorization, scalar and complex instruction sequences); in combination with hardware characteristics of RISC / SIMD / VLIW / DSP and the like, a multi-round iteration precision improvement and temporary register optimization allocation strategy is adopted, differential generation logic is formulated, and a high-precision low-redundancy instruction sequence is generated. The method comprises linear code scanning classification, reciprocal instruction and operand identification, cross-architecture generation logic rule formulation, instruction sequence generation and legality verification. Full-process automation is achieved, manual intervention is reduced, the execution efficiency and precision of reciprocal calculation of the heterogeneous processor are improved, and the method is suitable for embedded systems, high-performance calculation and other scenes.
Owner:HUNAN UNIV OF SCI & TECH

Asynchronous multi-granularity memory access control method in microprocessor, asynchronous circuit and memory access module

The invention discloses an asynchronous multi-granularity memory access control method in a microprocessor, an asynchronous circuit and a memory access module, the memory access control method adopts asynchronous clock-free control, introduces a multi-granularity memory access control mechanism, and divides the memory access granularity into two classes of multi-word memory access and non-multi-word memory access for processing. According to the method, the access path and the cache strategy can be flexibly adjusted according to different access granularities, redundant operation during access is reduced, and the access efficiency is greatly improved. The asynchronous circuit is a delay-limited asynchronous circuit and comprises an asynchronous controller and an asynchronous micro-pipeline structure based on a'sending-relay-receiving 'structure. The memory access module based on the asynchronous circuit comprises a memory access state management module, a data access updating module, an instruction / data cache and peripheral interaction interface, and a data interaction interface between the memory access module and the write-back module. According to the invention, lower dynamic power consumption and higher energy efficiency ratio are realized, and high efficiency and low delay are realized under large data volume operation.
Owner:LANZHOU UNIV

Sparse polynomial multiplication accelerator applied to HQC algorithm

The invention discloses a sparse polynomial multiplication accelerator circuit applied to an HQC algorithm, and belongs to the field of post quantum cryptography algorithm hardware acceleration. Comprising a sparse polynomial non-zero coefficient index register set, a state register set, a multiplication and addition unit, a control state machine and an address generation module. The sparse polynomial non-zero coefficient index register set is used for storing indexes of sparse polynomial non-zero coefficients and providing the indexes for the control state machine, and the control state machine further inputs the obtained indexes to the address generation module; the state register group is used for configuring and representing the working state of the accelerator, storing information of registers and providing information for the address generation module and the control state machine; the address generation module is used for calculating address information and feeding back a result to the control state machine; and the control state machine reads the coefficient of the dense polynomial from the external RAM and inputs the coefficient into the multiplication and addition unit for calculation, and after the multiplication and addition unit completes calculation, the control state machine writes a calculation result into the external RAM.
Owner:ZHEJIANG UNIV

Matrix multiply accumulation operation unit and operation method, and hardware accelerator and electronic device

Provided in the embodiments of the present disclosure are a matrix multiply accumulation (MMA) operation unit and operation method, and a hardware accelerator and an electronic device. The MMA operation unit comprises a data load / store engine, a tensor register file and a matrix multiplication engine, wherein the data load / store engine is used for loading data of a plurality of matrixes to be subjected to MMA computation; the tensor register file is used for storing the data of the plurality of matrixes that is acquired from the data load / store engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprising a plurality of tensor registers, and different tensor register groups being used for storing data of different matrixes among the plurality of matrixes; and the matrix multiplication engine is used for performing MMA computation on the basis of the data of the plurality of matrixes that is stored in the tensor register file. By means of the embodiments of the present disclosure, more efficient MMA computation is realized with low cost, low power consumption and a smaller footprint.
Owner:ALIBABA (CHINA) CO LTD

Data processing method and device, storage medium and electronic equipment

The embodiment of the invention provides a data processing method and device, a storage medium and electronic equipment, and relates to the technical field of computers, and the device comprises a data filling module which is used for carrying out data filling on received first data to obtain second data; splitting the second data to obtain at least one block message; the message word extension module is used for performing message word extension on each block message to obtain a message word for compression calculation; the compression function module is used for performing compression calculation on the message word to obtain a calculation result; the compression function module adopts an assembly line calculation framework formed based on a compressor and a working variable register, calculation variables required by compression calculation are stored in the compression function module, and the calculation variables change along with the change of the round number. In this way, the problem that in the prior art, a cryptographic chip is low in large-scale data safety calculation efficiency is solved through a pipeline calculation architecture, and the data calculation efficiency is improved in combination with the pre-stored calculation variables.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Register control method, processor, system on chip and computing equipment

The invention discloses a register control method, a processor, a system on chip and computing equipment. The method comprises the steps that in the process that a first process accesses a target register, in response to a processor, the first process is switched into a second process, the first state of the target register is obtained, and the target register is allowed to be accessed by a plurality of processes; under the condition that the first state meets a first preset state, storing the value of the target register into a target storage space; in response to the processor, switching the second process back to the first process, and obtaining a second state of the target register; and under the condition that the second state meets a second preset state, recovering the target register based on the data stored in the target storage space. According to the method and the device, the technical problems that the processing overhead of the processor is relatively high when the corresponding register needs to be updated during process switching of the processor in the related technology, and the process which needs to be realized by a processor hardware manufacturer in advance is relatively complex, time-consuming and relatively high in cost are solved.
Owner:DAMO ACAD (SHANGHAI) TECH CO LTD

Padding and suppressing rows and columns of data

A method is described herein. The method generally includes receiving stream parameters that defines an array, wherein the stream parameters include a first null element count and a second null element count. The method generally includes forming a stream of vectors for the multidimensional array responsive to the stream parameters. The stream of vectors generally includes a vector of null elements at a beginning of the stream of vectors based on the first null element count. The stream of vectors generally includes a null element at a beginning of each vector of the stream of vectors based on the second null element count. The stream of vectors generally includes a set of data distributed across a subset of the stream of vectors. The method generally includes providing the stream of vectors.
Owner:TEXAS INSTRUMENTS INC

Padding and suppressing rows and columns of data

A method is described herein. The method generally includes receiving stream parameters that defines an array, wherein the stream parameters include a first null element count and a second null element count. The method generally includes forming a stream of vectors for the multidimensional array responsive to the stream parameters. The stream of vectors generally includes a vector of null elements at a beginning of the stream of vectors based on the first null element count. The stream of vectors generally includes a null element at a beginning of each vector of the stream of vectors based on the second null element count. The stream of vectors generally includes a set of data distributed across a subset of the stream of vectors. The method generally includes providing the stream of vectors.
Owner:TEXAS INSTRUMENTS INC

Configurable time sequence control system

The invention discloses a configurable time sequence control system, which relates to the technical field of electronic components, can replace the requirements of most scenes on a time sequence controller, can be directly reused in a development process, can greatly reduce the development workload and can improve the development efficiency. And the time sequence can be flexibly adjusted by adjusting configuration parameters, so that the working difficulty of debugging and maintenance is reduced. The system is constructed by single-channel basic modules, and the maturity and reliability of the system are easy to guarantee. Pure Veri log / VHDL design is adopted, an IP core can be independently packaged and added into an IP library of a programmable logic device integrated development environment, the IP core can be directly called during design and use, and due to the fact that pure logic codes are adopted for implementation, limitation of device models and development platforms is avoided.
Owner:CHINA ORDNANCE EQUIP GRP AUTOMATION RES INST CO LTD

Instruction processing method and device, electronic equipment and readable storage medium

The embodiment of the invention provides an instruction processing method and device, electronic equipment and a readable storage medium, and the method comprises the steps: based on multiple pieces of source data and branch conditions corresponding to a to-be-executed branch jump instruction in a source program, selecting at least two to-be-selected branches contained in the branch jump instruction, determining jump branches respectively corresponding to the plurality of pieces of source data; taking at least two pieces of source data with the same jump branches as at least two pieces of target data, and writing the at least two pieces of target data into a first vector register; the operation types of the jump branches of the at least two pieces of target data serve as target types, a target vector instruction is generated based on the target types and the first vector register, and the target vector instruction is used for executing the operation of the target types on the at least two pieces of target data in the first vector register; and executing the target vector instruction. And the instruction execution efficiency is improved.
Owner:LOONGSON TECH CORP

Processor core, processor, and method for processor

The embodiment of the invention provides a processor core, a processor and a method for the processor. The processor core includes a processing pipeline configured to rename and execute a first instruction including at least one of a first architectural register and a second architectural register, and includes: a first type of physical register configured to be mapped by the first architectural register and configured to store a first type of data; the first type of physical register is configured to be mapped by the first architecture register and configured to store data of a first type, the second type of physical register is configured to be mapped by the second architecture register and configured to store data of a second type, the first architecture register comprises a vector architecture register, the data of the first type stored by the first type of physical register comprises vector data, and the second architecture register comprises a mask architecture register; the second type of data stored in the second type of physical register comprises mask data. The processor core may mitigate processing pipeline stagnation and increase less area.
Owner:HYGON INFORMATION TECH CO LTD

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Register overflow optimization method and device and storage medium

The embodiment of the invention provides a register overflow optimization method and device and a storage medium, and is applied to the technical field of chips. In the method, for each virtual register in a target program, based on a physical register type supported by an instruction operand where the virtual register is located, the virtual register is optimized; selecting a corresponding target register class from N candidate register classes, wherein N is greater than 1; allocating a first physical register in the target register class to the virtual register; when register overflow occurs, a target register is selected from the allocated first physical registers, the instruction operand stored in the target register overflows to the second physical registers in the other N-1 candidate register classes, and compared with the mode that the instruction operand overflows to the memory to generate read-write operation on the memory, the instruction operand stored in the target register overflows to the second physical registers in the other N-1 candidate register classes; according to the method, different types of physical registers are overflowed, read-write operation aiming at the physical registers is generated, the pressure of the registers is relieved, the performance overhead of the overflowed memories is reduced, and the register distribution efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Automatic illegal character cleaning system

The invention relates to the technical field of computer data processing and network security, in particular to an illegal character automatic cleaning system, which comprises a rule compiling module for monitoring rule change, performing semantic fusion and topological mapping on a rule set, constructing a deterministic finite automaton and mapping the deterministic finite automaton into a state transition table; the state switching module is used for constructing a double-buffer context and operating lock-free switching through an atomic pointer to realize hot updating; the speculation execution module is used for carrying out vectorization pre-scanning based on a state transition table by utilizing single-instruction multi-data stream parallel loading, identifying a walk path, falling into a safe state for releasing and falling into a trap state for triggering external verification; the self-adaptive feedback module is used for counting trap state triggering frequency, generating a rule allergy report and dynamically adjusting the size of a read fragment; according to the method, the contradiction between rule flexibility and execution efficiency is solved, and high-performance cleaning based on speculative execution is realized.
Owner:北京啄木鸟云健康科技有限公司

In-memory computing accelerator, system and method based on adaptive utilization

The invention discloses an in-memory computing accelerator, system and method based on a self-adaptive utilization rate. The in-memory computing accelerator based on the self-adaptive utilization rate comprises an in-memory computing processor, a global buffer area, a vector processor, a control module and a configuration register file, the control module and the configuration register file are respectively connected with the in-memory computing processor, the global buffer area and the vector processor, the in-memory computing processor is connected with the global buffer area, and the global buffer area is connected with the vector processor; the in-memory computing processor comprises a plurality of in-memory computing cores, and the plurality of in-memory computing cores form a dynamic parallel structure. According to the method and the device, at least one neural network reasoning calculation process can be executed on the input data aiming at the input data of a certain neural network, and relatively high space utilization rate and time utilization rate can be obtained when facing calculation requirements of different operators, so that efficient acceleration of a diversified neural network model is realized.
Owner:SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Enabling high-performance scalable matrix extension (SME) instruction issue in processor devices

Enabling high-performance Scalable Matrix Extension (SME) instruction issue in processor devices is disclosed herein. In some aspects, a processor device comprises a reservation station circuit configured to perform, during a first phase, a reduced-precision vector accumulator (ZA) tracking operation on micro-ops for which corresponding vector (Z) registers and corresponding predicate (P) registers are ready. Based on the reduced-precision ZA tracking operation, the reservation station circuit selects a first micro-op and a second micro-op having no Read-After-Write (RAW) hazard with respect to the ZA registers. During a subsequent second phase, the reservation station circuit performs a full-precision ZA tracking operation on the first micro-op and the second micro-op, and selects one as a micro-op for issue for which the full-precision ZA tracking operation indicates no RAW hazard exists with respect to the ZA registers. The reservation station circuit then issues the selected micro-op for execution.
Owner:QUALCOMM INC

Special instruction set processor for polar code coding and decoding algorithm and implementation method

The invention relates to the technical field of communication and computer instruction execution and processing, in particular to a polar code encoding and decoding algorithm-oriented special instruction set processor and an implementation method, and a vector processor has efficient parallel computing capability. A coding and decoding algorithm of a polar code relates to parallel computing of a large number of log-likelihood ratios, a vector processor of a special instruction set can process a plurality of LLR values at the same time through a parallel processing unit, and the throughput rate is remarkably improved; special instructions such as polarization shuffling and vector G operation instructions are designed for polarization code encoding and decoding and are used for accelerating algorithm operation. The method can achieve the balance among the performance, the energy efficiency and the flexibility, can adapt to the continuous iteration upgrading of the algorithm due to the programmable and reusable characteristics, and has a good application prospect.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Operator execution method and device, computer equipment and storage medium

ActiveCN120276769ARegister arrangementsFusion operatorProcessing
The invention discloses an operator execution method and device, computer equipment and a storage medium, and belongs to the technical field of artificial intelligence, in the method, after an output result of first precision of a first operation in a fusion operator is obtained, the output result is stored in a TLR array according to a preset data arrangement mode, and the output result of the first precision of the first operation in the fusion operator is obtained; and converting the data in the TLR array from the first precision to a second precision based on a preset data arrangement mode, the second precision being an input precision corresponding to a second operation in the fusion operator, and after determining that the converted data meets a data arrangement requirement corresponding to the second operation, processing the data through each thread to obtain a calculation result of the fusion operator, the processing is determined according to the category of the second operation. Thus, the data arrangement mode of the output result of the previous operation in the fusion operator during output is changed, the data arrangement requirement corresponding to the next operation can be met after precision conversion, data exchange between threads is not needed, hardware overhead is small, and therefore the performance of the fusion operator can be improved.
Owner:SHANGHAI BIREN TECH CO LTD