Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

506 results about "Operand" patented technology

In mathematics an operand is the object of a mathematical operation, i.e., it is the object or quantity that is operated on.

Integer parallel computing method and device based on distributed storage and computer equipment

The invention belongs to the field of high-performance computing, and relates to an integer parallel computing method and device based on distributed storage and computer equipment, and the method comprises the steps of collecting real-time resource indexes, dynamically identifying fault nodes, triggering task migration, and performing data verification and hard disk fault detection. The weight value of each node is calculated, the nodes are arranged according to the descending order of the weight values, and the nodes with high load capacity are selected to distribute tasks; dynamically distributing a data generation task to a computing node, executing parallel computing, and performing distributed storage on a result; obtaining an operand, converting the operand into a first-order tensor form of a basic operand, serializing tensor data, and sending the serialized tensor data to a parallel computing layer; distributing a search task to a computing node, retrieving storage data in parallel, reading effective data from a storage layer, and combining search results into a partial sum; and summarizing and then outputting. The system has dynamic resource management and fault-tolerant capabilities, and can realize efficient task allocation and load balancing.
Owner:SHENZHEN Y& D ELECTRONICS CO LTD

Instruction execution method, processor, electronic equipment and storage medium

The invention provides an instruction execution method, a processor, electronic equipment and a storage medium, the method is applied to the processor, the processor comprises a first execution unit, a second execution unit and a pipeline register, and a first execution instruction of an atomic instruction group is obtained through the first execution unit; writing a first execution result of the first execution instruction into the pipeline register in response to the first execution instruction; a second execution instruction of the atomic instruction group is obtained through the second execution unit, and the first execution result is read from the pipeline register as the source operand of the second execution instruction in response to the second execution instruction, so that the instruction execution method, the processor, the electronic equipment and the storage medium can reduce the access frequency of the general register and improve the access efficiency of the general register. Processor power consumption and register port conflicts can be reduced.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Multiplication hardware block with adaptive fidelity control system

Methods and systems relating to computational hardware are disclosed herein. One disclosed method for executing a multiplication computation using a computational hardware block includes storing a first operand and a second operand for the multiplication computation. The first operand includes a first set of bit strings. The second operand includes a second set of bit strings. The method also includes multiplying the first set of bit strings and the second set of bit strings in a set of temporal phases using the computational hardware block. Each temporal phase uses a different group of bit strings from the first set of bit strings and the second set of bit strings. A cardinality of the set of temporal phases is determined by a fidelity control value. The fidelity control value adaptively sets a fidelity of execution of the multiplication computation.
Owner:TENSTORRENT AI ULC

Random calculation processing unit and method based on adaptive compensation mechanism

ActiveCN120653304AMachine execution arrangementsStochastic computingOperand
The invention relates to a random calculation processing unit and method based on a self-adaptive compensation mechanism, and the method employs a self-adaptive compensation core formed by solidifying a lightweight neural network to replace a conventional solidification compensation rule, and the self-adaptive compensation core can carry out the random calculation according to two operands participating in multiplication. And an optimal compensation parameter is dynamically predicted in real time for compensation. According to the scheme, the calculation error of random calculation is accurately, continuously and individually compensated in a data driving mode, and hardware implementation is performed by adopting a constant coefficient multiplier technology, so that the self-adaptive compensation core has extremely low area and power consumption overhead in hardware. Compared with the prior art, the method has the advantages that the fidelity of random calculation can be greatly improved on the premise that the hardware cost is not remarkably increased, and a new effective way is provided for constructing a high-performance and high-energy-efficiency random calculation neural network accelerator.
Owner:NAT UNIV OF DEFENSE TECH

Heterogeneous processor-oriented reciprocal calculation instruction sequence generation method

The invention discloses a reciprocal calculation instruction sequence generation method oriented to a heterogeneous processor, and belongs to the field of compilation optimization and code generation. Aiming at the problems of instruction redundancy, weak precision control, poor hardware adaptation and high manual dependence of an existing method in a heterogeneous environment, characteristics of a reciprocal instruction and an operand are accurately identified by linearly scanning heterogeneous object codes (including vectorization, scalar and complex instruction sequences); in combination with hardware characteristics of RISC / SIMD / VLIW / DSP and the like, a multi-round iteration precision improvement and temporary register optimization allocation strategy is adopted, differential generation logic is formulated, and a high-precision low-redundancy instruction sequence is generated. The method comprises linear code scanning classification, reciprocal instruction and operand identification, cross-architecture generation logic rule formulation, instruction sequence generation and legality verification. Full-process automation is achieved, manual intervention is reduced, the execution efficiency and precision of reciprocal calculation of the heterogeneous processor are improved, and the method is suitable for embedded systems, high-performance calculation and other scenes.
Owner:HUNAN UNIV OF SCI & TECH

SRT operational circuit

The SRT operational circuit comprises an input module, a floating point conversion module, a calculation module and an output module, the input module is configured to output an initial operand, a first end of the floating point conversion module is connected with the input module, a first end of the calculation module is connected with the floating point conversion module and the input module, and a second end of the calculation module is connected with the output module. The output module is connected with the second end of the calculation module; when the initial operand is an integer, the floating point conversion module is configured to perform floating point conversion on the initial operand to output a first operand; if the calculation module is configured to perform division operation or root extraction operation by adopting the first operand based on the SRT algorithm, the output module is configured to perform mantissa rounding on a division result or a root extraction result after the calculation of the calculation module is completed and output a floating point result, so that the operation accuracy of an integer in the circuit is improved, and the universality of the circuit is improved.
Owner:GUANGDONG LEAPFIVE TECH CO LTD

Methods and circuits for streaming data to processing elements in stacked processor-plus-memory architecture

A stacked processor-plus-memory device includes a processing die with an array of processing elements of an artificial neural network. Each processing element multiplies a first operand—e.g. a weight—by a second operand to produce a partial result to a subsequent processing element. To prepare for these computations, a sequencer loads the weights into the processing elements as a sequence of operands that step through the processing elements, each operand stored in the corresponding processing element. The operands can be sequenced directly from memory to the processing elements or can be stored first in cache. The processing elements include streaming logic that disregards interruptions in the stream of operands.
Owner:RAMBUS INC

Vector instruction processing method, electronic equipment, readable medium and program product

The invention provides a vector instruction processing method, electronic equipment, a readable medium and a program product. The vector instruction processing method comprises the steps of obtaining a vector length of a vector instruction; when the vector length of the vector instruction is the first value and the effective indication of the second source operand is the first value, cancelling the transmission of the vector instruction, and setting the effective indication of the second source operand as a second value; a first value of the second source operation valid indication indicates that the old value of the destination register is not required as a valid indication of the second source operand, and a second value of the second source operation valid indication indicates that the old value of the destination register is required as a valid indication of the second source operand; under the condition that the destination register is ready, retransmitting the vector instruction; a first source operand and a second source operand are obtained based on the vector instruction, and the vector instruction is executed based on the first source operand and the second source operand, the second source operand being an old value of the destination register.
Owner:SANECHIPS TECH CO LTD

Vector mask buffers in a vector instruction execution pipeline

Systems and methods related to vector mask buffers in a vector instruction execution pipeline are disclosed herein. The vector instruction execution pipeline may include several lanes. Each lane may include a vector register file, a vector mask buffer, and a functional processing unit. The vector register file may store operand data and the vector mask buffer may store a vector mask associated with the operand data. In a lane, the operand data may be read from the register file into a functional processing unit, and the vector mask may be read from the vector mask buffer to the functional processing unit. The functional processing unit may process the operand data based on the vector mask. The lane-specific vector mask buffers improve the efficiency of the vector instruction execution pipeline by storing the vector masks proximate to where the vector masks will be used.
Owner:TENSTORRENT USA INC

Floating point arithmetic device and method of operating the same

A floating point arithmetic device with two floating point operands and its operation method are disclosed. The floating point arithmetic device includes an exponent subtraction circuit, an exponent calculation circuit, a mantissa calculation circuit, and a conversion circuit. The exponent subtraction circuit calculates the difference between the exponents of the two operands and generates a sign bit and an exponent difference. The exponent calculation circuit generates the post-operation exponent bits according to the larger one of the exponents of the two operands. The mantissa calculation circuit aligns the mantissa bits of the two operands and performs one of addition and subtraction on the aligned mantissa bits. To improve the calculation efficiency and reduce the power consumption, the floating point arithmetic device can complete the floating point addition or subtraction operation in one step (one clock cycle) without moving the intermediate floating point data between the registers and the functional circuit units as in the multi-step operation.
Owner:XINLIJIA INTEGRATED CIRCUIT (SHANGHAI) CO LTD

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Variable-precision parallelization index floating-point number storage and calculation integrated method based on FPGA

The invention belongs to the field of computers, and discloses a variable-precision parallelization index floating point number storage and calculation integrated method based on an FPGA (Field Programmable Gate Array), which specifically comprises the following steps of: inputting decimal floating point data; an index floating-point number binary sequence format is defined, decimal floating-point data values are converted into the index floating-point number binary sequence format, parameters of the index floating-point number binary sequence format comprise sign bits, index bits, mantissa bits, specified index offset and mantissa bit width, and the mantissa bit width is dynamically adjusted according to calculation requirements; based on decimal floating point data numerical values, extracting parameters of an index floating point number binary sequence, and storing the parameters in a register of the FPGA as operands; performing displacement sorting on the two operands stored in the register, and calculating an absolute value of a difference between exponent bits and a difference between mantissa bits of the two operands; the problems that in the prior art, a floating-point number storage scheme is low in calculation efficiency and insufficient in precision in an application scene where the dynamic range changes drastically are solved.
Owner:SOUTHEAST UNIV

Register overflow optimization method and device and storage medium

The embodiment of the invention provides a register overflow optimization method and device and a storage medium, and is applied to the technical field of chips. In the method, for each virtual register in a target program, based on a physical register type supported by an instruction operand where the virtual register is located, the virtual register is optimized; selecting a corresponding target register class from N candidate register classes, wherein N is greater than 1; allocating a first physical register in the target register class to the virtual register; when register overflow occurs, a target register is selected from the allocated first physical registers, the instruction operand stored in the target register overflows to the second physical registers in the other N-1 candidate register classes, and compared with the mode that the instruction operand overflows to the memory to generate read-write operation on the memory, the instruction operand stored in the target register overflows to the second physical registers in the other N-1 candidate register classes; according to the method, different types of physical registers are overflowed, read-write operation aiming at the physical registers is generated, the pressure of the registers is relieved, the performance overhead of the overflowed memories is reduced, and the register distribution efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Processor supporting data multiplexing and instruction multiplexing and data multiplexing method

The invention provides a processor supporting data multiplexing and instruction multiplexing, which is used for supporting multiplexing of operands and instructions of a previous calculation task in a later-executed calculation task, and realizes operand multiplexing and instruction multiplexing based on configuration information stored by a host control module connected with the processor. The processor comprises a configuration information caching module used for storing a storage offset and a storage offset use mark which are acquired from configuration information and correspond to each instruction in each calculation task; the base address dynamic generation module is used for storing a data cache index address of each calculation task and generating a real base address of each instruction in each calculation task; and the calculation module is used for acquiring an operand to execute each calculation task based on the real base address of each instruction in each calculation task generated by the base address dynamic generation module.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Loading storage circuit and graphics processor

The invention provides a loading storage circuit and a graphics processor, and relates to the technical field of graphics processing. The loading storage circuit comprises an instruction scheduling module, an address generation module and a data service module, and is provided with a buffer area module which comprises an operand buffer area and an effective data buffer area and is used for caching information required by address calculation and data access; the instruction scheduling module is used for collecting a data access instruction and outputting instruction information; the address generation module calculates a target access address according to the instruction information and the operand; and the data service module executes corresponding data loading or storage operation on the target cache unit. According to the scheme, the operands and the valid data are pre-cached, so that overflow and stagnation caused by inconsistent processing rhythms among modules can be avoided, and the parallel processing capability of an assembly line is improved; through the independent address calculation and data access process, the stability of the access time sequence and the data processing efficiency can be improved.
Owner:MOORE THREADS TECH CO LTD

Multi-precision floating point fusion multiply-add structure and microprocessor architecture

The invention provides a multi-precision floating point fusion multiply-add structure and a microprocessor architecture, and the structure is characterized in that in a first-stage assembly line, an input preprocessing module decomposes an input operand to obtain a sign bit, an index bit and a mantissa bit; the Booth encoder generates partial products through a radix-4 Booth encoding algorithm, and the tree array multiplier compresses the partial products into three groups; in the second-stage assembly line, the CSA4-2 compression adder compresses the mantissa bits after the three groups of partial products and shift alignment into two groups; the summator module sums the two groups of partial products to obtain a mantissa summation result; in the third-stage assembly line package, a leading zero detection module performs leading zero detection on the mantissa summation result, and performs index adjustment and mantissa adjustment to obtain a normalized result; and in the fourth-stage assembly line, rounding of floating point data is carried out. According to the application, the operation performance, the energy efficiency ratio and the adaptability of the floating point fusion multiply-add structure are improved, so that the high performance and the high energy efficiency of the microprocessor are balanced.
Owner:EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD

Systems and methods for energy-efficient, bit-parallel, multiply-accumulate for artificial intelligence and deep neural networks

A system and method for providing a tunable floating-point multiply-accumulate (MAC) unit are disclosed. The unit maintains full arithmetic precision while enabling dynamic elimination of ineffectual computation through operand decomposition and selective activation of partial product generation logic. The disclosed MAC unit is suitable for drop-in replacement in existing deep-learning accelerators and improves energy efficiency without requiring architectural changes.
Owner:KAXIRAS STEFANOS +3

Vector mask buffers in a vector instruction execution pipeline

Systems and methods related to vector mask buffers in a vector instruction execution pipeline are disclosed herein. The vector instruction execution pipeline may include several lanes. Each lane may include a vector register file, a vector mask buffer, and a functional processing unit. The vector register file may store operand data and the vector mask buffer may store a vector mask associated with the operand data. In a lane, the operand data may be read from the register file into a functional processing unit, and the vector mask may be read from the vector mask buffer to the functional processing unit. The functional processing unit may process the operand data based on the vector mask. The lane-specific vector mask buffers improve the efficiency of the vector instruction execution pipeline by storing the vector masks proximate to where the vector masks will be used.
Owner:TENSTORRENT USA INC

Data layout optimization method and device applied to NPU code compiling and medium

The invention discloses a data layout optimization method and device applied to NPU code compilation and a medium. The method comprises the steps that an intermediate representation IR input by an upper layer is split into an operation type OP and an operand, and the operand input by the upper layer is divided into logic data and a logic mask; converting the logic data of the upper layer into data vector representation of the abstraction layer, and converting the logic mask of the upper layer into mask vector representation of the abstraction layer; obtaining a corresponding bottom hardware instruction capability according to an operation type OP input by an upper layer, and performing legalization operation on the operation type OP; and performing legalization processing on the converted abstract data vector type representation and mask vector representation according to the target hardware capability, performing instruction mapping on the processed abstract instruction according to the underlying hardware pair operation type OP, and generating a target LLVM instruction compatible with the underlying hardware. Register layout abstraction and conversion can be carried out on the NPU code under mixed precision input, and the utilization efficiency of the NPU bottom layer register is improved.
Owner:SOUTH CHINA UNIV OF TECH

Method for building instruction set acceleration component verification platform based on UVM

The invention provides a UVM-based instruction set acceleration component verification platform construction method, which comprises the following steps: firstly, generating a required instruction signal according to an input signal of an acceleration component by a verification platform, then generating a floating-point number meeting calculation requirements according to requirements, and meanwhile, ensuring the configurability of various parameters which are random and excited in an effective range of operational data; a C-based soft library is called as a reference model for comparison with a hardware output result, so that the verification accuracy is ensured; then collecting and merging coverage rate files obtained by excitation operation through scripts to realize convergence of the coverage rate and ensure the completeness of verification; and finally, when hardware design finds an error and is modified, a verification platform can be called to carry out a regression test, so that the method has relatively high reusability, and the verification efficiency is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Processor, method, device and storage medium for data processing

According to an embodiment of the present disclosure, a processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand and a target operand. The target opcode indicates the vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading data to be processed. The target operand specifies a target storage location in the memory for writing a processing result. The processor also includes an arithmetic logic unit coupled to the instruction decoder and the memory. The arithmetic logic unit is configured to: read data to be processed from a source storage location in the memory; perform an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed; and write the processing result to a target storage location in the memory. In this way, the efficiency of vector calculations can be improved.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Efficient execution of group-sparsified neural networks

Creating neural network (NN) code may include for each row in a kernel matrix, finding the first non-zero element; and creating a loop including multiply and add instructions. On each iteration of the loop, the multiply and add instructions may be executed, and the position of the kernel matrix operand operated on by each multiply and add may be correlated to the loop iteration number. Instructions may be issued or created to be executed in the loop. A method may execute a NN by executing a loop including a series of multiply and add instructions to multiply a kernel matrix A by an input, such that on each iteration of the loop the series of multiply and add instructions are executed; and the position of the matrix A operand operated on by each multiply and add instruction in the series is correlated to the iteration number of the loop.
Owner:RED HAT INC

Optimization system and method for data delivery based on speculative wakeup and related equipment

The invention is suitable for the technical field of processors, and particularly relates to a speculative wake-up-based data forwarding optimization system and method and related equipment. According to the method and the device, the data forwarding flag bit and the data forwarding pipeline number are set for the to-be-executed instruction, and whether the to-be-executed instruction needs to be subjected to data forwarding or not is judged according to the data forwarding flag bit; if yes, the instruction to be executed needs to obtain the corresponding source operand from the front instruction, data forwarding needs to be carried out, and the source operand is read from the corresponding pipeline of the assembly line according to the serial number of the data forwarding pipeline; and if the data forward flag bit is the second data forward flag bit, the instruction to be executed does not need to acquire the source operand from the front instruction and is directly read from the physical register, so that resource conflicts are reduced. Compared with the prior art, the method has the advantages that the logic level of the data forwarding behavior can be effectively reduced, and the running efficiency and performance of the processor are improved.
Owner:BLUECORE COMPUTING POWER (SHENZHEN) TECHNOLOGY CO LTD

Processor, method, device and storage medium for data processing

A processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading to-be-processed data. The target operand specifies a target storage location in the memory for writing a processed result. The processor further includes an arithmetic logic unit configured to: read the to-be-processed data from the source storage location of the memory; perform, on the to-be-processed data, an arithmetic logic operation associated with the vector operation specified by the target instruction; and write the processed result to the target storage location of the memory.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Speculative wakeup instruction dependency chain cancellation systems, methods, and related devices

The invention discloses a speculative wake-up instruction dependency chain cancelling system and method and related equipment, and the system comprises a memory access miss judgment module which is used for judging whether a front instruction of a to-be-executed instruction can hit memory access or not; the register number comparison module is used for comparing whether the source register number of the instruction to be executed is the same as the destination register number of the operand; and the cancel processing module is used for informing the emission queue module to clear the emission flag bits of the to-be-executed instruction and the instruction having the dependency relationship with the to-be-executed instruction when the to-be-executed instruction is supposed to fail to be awakened. Compared with the prior art, the cancel processing module notifies the transmission queue module to initiate the cancel behavior of clearing the transmission flag bit in different time periods, so that a long instruction dependency chain is decoupled into a plurality of short dependencies, and on the premise that the performance is not influenced, the performance of the system is greatly improved. Two register number comparison circuits are reduced from a time sequence critical path of an instruction when speculative wake-up is cancelled.
Owner:BLUECORE COMPUTING POWER (SHENZHEN) TECHNOLOGY CO LTD

ALU operation fusion processing module and method suitable for neural network

The invention discloses an ALU operation fusion processing module and method suitable for a neural network, and the module comprises a control unit which is used for receiving and decoding a machine instruction, managing the execution processes of internal and external circulation and microinstruction circulation, and generating a control signal of each stage of a microinstruction assembly line; the microinstruction buffer area is used for storing a microinstruction sequence pre-generated by the neural network compiler; the register file is used for storing source operands and results of ALU operation; each entry of the register file is composed of a valid bit, a tag bit and a data bit; the ALU computing core adopts an SIMD (Single Instruction Multiple Data) architecture and comprises a plurality of paths of parallel arithmetic logic function units; the Load / Store unit is used for processing data exchange between a register file and a local buffer area; and the data selection interface is used for selecting a data source or a target buffer area according to the storage tag field of the microinstruction. According to the method, the high efficiency and the flexibility of the ALU in the neural network hardware accelerator can be effectively considered.
Owner:ZHEJIANG UNIV

Large-scale matrix restructuring and matrix-scalar operations

Embodiments of apparatuses and methods for copying and operating on matrix elements are described. In embodiments, an apparatus includes a hardware instruction decoder to decode a single instruction and execution circuitry, coupled to hardware instruction decoder, to perform one or more operations corresponding to the single instruction. The single instruction has a first operand to reference a base address of a first representation of a source matrix and a second operand to reference a base address of second representation of a destination matrix. The one or more operations include copying elements of the source matrix to corresponding locations in the destination matrix and filling empty elements of the destination matrix with a single value.
Owner:INTEL CORP

Apparatus and method for vector packed signed / unsigned shift, round, and saturate

Apparatus and method for signed and unsigned shift, round and saturate using different data element values. For example, one embodiment of an apparatus comprises a decoder to decode an instruction having fields for a first packed data source operand to provide a first source data element and a second source data element, a second packed data source operand or immediate to provide a first shift value and a second shift value corresponding to the first source data element and second source data element, respectively, and a packed data destination operand to indicate a first result value and a second result value corresponding to the first source data element and second source data element, and execution circuitry to execute the decoded instruction to: shift the first source data element by an amount based on the first shift value to generate a first shifted data element; shift the second source data element by an amount based on the second shift value to generate a second shifted data element; update a saturation indicator responsive to detecting a saturation condition resulting from the shift of the first and / or second source data elements; round and / or saturate the first and second shifted data elements in accordance with a specified rounding mode and the saturation indicator, respectively, to generate the first and second result data elements; and store the first result value and the second result value in a first data element location and a second data element location in a destination register.
Owner:INTEL CORP

Computing device and method based on RISC-V extension instruction

The invention provides a computing device and method based on RISC-V extension instructions, the computing device supports approximate computation of mixed precision according to approximate computation instructions in an extended approximate computation instruction set, and the computing device comprises an out-of-order scheduling and register reading module used for executing instruction dependency analysis and operand preloading, scheduling the non-approximate calculation instruction and the approximate calculation instruction to different transmitting queues respectively; the first instruction transmitting queue is used for temporarily storing a to-be-transmitted non-approximate calculation instruction; the second instruction transmitting queue is used for temporarily storing approximate calculation instructions to be transmitted; the precise calculation module is used for completing precise calculation related tasks according to the instruction from the first instruction transmitting queue; and the approximate calculation module is used for completing approximate calculation related tasks according to the instructions from the second instruction transmitting queue, and supports approximate calculation of various precisions.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI