Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

162 results about "Arithmetic logic unit" patented technology

An arithmetic logic unit (ALU) is a combinational digital electronic circuit that performs arithmetic and bitwise operations on integer binary numbers. This is in contrast to a floating-point unit (FPU), which operates on floating point numbers. An ALU is a fundamental building block of many types of computing circuits, including the central processing unit (CPU) of computers, FPUs, and graphics processing units (GPUs). A single CPU, FPU or GPU may contain multiple ALUs.

Storage and calculation integrated server optimization method based on NPU

The invention discloses a storage and calculation integrated server optimization method based on an NPU, relates to the technical field of server architecture, and discloses the storage and calculation integrated server optimization method based on the NPU by dynamically adjusting the configuration of an arithmetic logic unit, a multiplication and accumulation calculation unit and a cache module to generate a reconstruction calculation unit. By combining heterogeneous hardware performance evaluation and a task migration mechanism, dynamic optimal configuration and load balancing of computing resources are realized, the task execution efficiency can be improved, and the heterogeneous hardware collaborative energy efficiency ratio can be optimized.
Owner:四川华鲲振宇智能科技有限责任公司

Arithmetic logic unit of processor and floating-point number calculation method

The invention provides an arithmetic logic unit of a processor and a floating-point number calculation method, and belongs to the technical field of floating-point numbers. The arithmetic logic unit comprises a format conversion module and a mixed precision multiplier-adder, the format conversion module performs format conversion on a high-precision floating-point number to obtain a low-precision floating-point number, and retains an index lost in the conversion process as index auxiliary information, and when the mixed precision multiplier-adder performs multiply-add operation on the low-precision floating-point number, the mixed precision multiplier-adder performs multiply-add operation on the low-precision floating-point number. And the index auxiliary information is used for compensating the lost index, so that the precision loss of the floating-point number can be reduced when the mixed precision multiplier-adder is used.
Owner:HUAWEI TECH CO LTD

Arithmetic logic unit parallel processing system and method and electronic equipment

The invention provides an arithmetic logic unit parallel processing system and method and electronic equipment, and relates to the field of chip architecture processing, and the arithmetic logic unit parallel processing system comprises an instruction transmitting unit, a write-back unit and a plurality of ALU groups containing a plurality of arithmetic logic units; wherein the instruction transmitting unit is respectively connected with the input end of each ALU group; the output end of each ALU group is connected with the write-back unit; the input end of each ALU group is respectively connected with the input ends of all the arithmetic logic units in the current ALU group; the output end of each ALU group is respectively connected with the output ends of all the arithmetic logic units in the current ALU group; according to the scheme, the reduction instruction and the vector arithmetic logic operation can be realized by using the same arithmetic logic operation component, the number and the area of the vector arithmetic logic operation component are reduced, and the utilization rate of the vector arithmetic logic operation component is improved.
Owner:BEIJING YIHUA CLOUD NETWORK TECH CO LTD

Memory device and method

A memory device includes a plurality of memory banks, and a processing-in-memory (PIM) block accessible to the plurality of memory banks, wherein the PIM block comprises a control circuit configured to receive a plurality of operation instructions from a host and, in response to a predicated instruction indicating a predication operation among the plurality of operation instructions, instruct an arithmetic logic unit (ALU) to perform the predication operation, a predicate register file (PRF) configured to store therein a predicate value determined by the predication operation, and the ALU configured to perform an operation according to a command signal translated by the control circuit based on the predicate value from an operation instruction that depends on the predicate value among the plurality of operation instructions.
Owner:SAMSUNG ELECTRONICS CO LTD +1

Processor, method, device and storage medium for data processing

According to an embodiment of the present disclosure, a processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand and a target operand. The target opcode indicates the vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading data to be processed. The target operand specifies a target storage location in the memory for writing a processing result. The processor also includes an arithmetic logic unit coupled to the instruction decoder and the memory. The arithmetic logic unit is configured to: read data to be processed from a source storage location in the memory; perform an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed; and write the processing result to a target storage location in the memory. In this way, the efficiency of vector calculations can be improved.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Constant data loading method, graphics processor and medium

The invention discloses a constant data loading method, a graphics processor and a medium. The constant data loading method comprises the following steps: storing data of at least part of data units in a constant cache into a target storage space; the target storage space comprises at least one vector register; acquiring a data access instruction, wherein the data access instruction is used for accessing target data in the constant cache; if it is detected that the target data is stored in a target storage space, compiling the data access instruction into an arithmetic logic unit instruction; the arithmetic logic unit instruction indicates the position of the target data in the target storage space. By adopting the scheme, the operation efficiency of the shader can be improved.
Owner:RICUN TECH (SHANGHAI) CO LTD

Processor, method, device and storage medium for data processing

A processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading to-be-processed data. The target operand specifies a target storage location in the memory for writing a processed result. The processor further includes an arithmetic logic unit configured to: read the to-be-processed data from the source storage location of the memory; perform, on the to-be-processed data, an arithmetic logic operation associated with the vector operation specified by the target instruction; and write the processed result to the target storage location of the memory.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Precision target optimization method and system adaptive to variable precision arithmetic logic unit, medium, terminal and program product

The invention provides a precision target optimization method and system adaptive to a variable precision arithmetic logic unit, a medium, a terminal and a program product. The method comprises the following steps: acquiring an output feature set of each group of an upper layer; the precision generation network layer generates a corresponding precision target according to the output feature set, and the ALU calculation layer generates a prediction result according to the generated precision target; the teacher model generates a reference target and a real label according to the output feature set; constructing a total loss function according to the calculated task loss, precision generation loss and adversarial loss; performing back propagation optimization on the student model based on the constructed total loss function; repeatedly and iteratively training the student model until convergence to obtain a final student model; and deploying the final student model to generate an optimal precision target corresponding to each group. According to the method provided by the invention, the fine precision adjustment of the bit granularity can be realized, the adaptive ability of the model is enhanced, and the matching degree of the precision and the task demand is improved.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Vector engine for efficient beam search

To provide native hardware support for complex operations such as a beam search operation, an integrated circuit device can be implemented with multiple computational circuit blocks coupled in series to form a pipeline. Each computational circuit block includes an arithmetic logic unit (ALU) circuit having a first numeric input, a second numeric input, a primary result output, and a secondary output. The ALU circuit is programmable to perform an arithmetic operation on the first numeric input and the second numeric input to generate the primary result output, and provide one of the first numeric input or the second numeric input to the secondary output, which is fed back to the ALU circuit.
Owner:AMAZON TECH INC

Greedy algorithm for in network computation trees

Techniques and architecture are described for a method that includes an in network compute (INC) manager receiving from switches of a fat tree configured network, arithmetic logic unit (ALU) capacity of the switches. Based at least in part on the ALU capacity of the switches and bandwidth, the INC manager determines one or more switches within each tier that are capable of supporting the processing units and based at least in part on the determining, the INC manager selects a first switch as a root, wherein the first switch is included within a tier of switches having intermediate tiers of switches located between the tier and the plurality of processing units within the fat tree configured network. The INC manager creates one or more paths of switches within each of the intermediate tiers from the root to the plurality of processing units to provide a constrained disjoint spanning tree of switches.
Owner:CISCO TECHNOLOGY INC

Sparse tensor processing in a machine learning accelerator

In an example, a processor for machine learning calculations is described. An adapter input circuit is operable to receive an input tensor. The adapter input circuit includes channels. A first channel of the channels is operable to process samples of the input tensor to generate pre-processed samples and to obtain locations of the samples. A location processor, coupled to the first channel, is operable to determine output locations in response to the locations. An arithmetic logic unit (ALU), coupled to the channels, is operable to calculate output samples from the pre-processed samples. An adapter output circuit, coupled to the location processor and the ALU, operable to process the output locations and the output samples to generate an output tensor.
Owner:AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD

Clustering of machine learning (ML) functional components

A graphics processing unit (GPU) for clustering of machine learning (ML) functional components, including: a plurality of compute units; a plurality of ML clusters, wherein each of the ML clusters comprises at least one arithmetic logic unit (ALU), and wherein each of the ML clusters is associated with a respective subset of the compute units; and a plurality of memory modules each positioned on the GPU adjacent to a respective ML cluster of the plurality of ML clusters, wherein each ML cluster is configured to directly access one or more adjacent memory modules.
Owner:ADVANCED MICRO DEVICES INC

Angle-based solution system based on CORDIC instructions

This invention belongs to the field of computer technology and relates to an angle calculation system based on CORDIC instructions. The system includes an IFU (Instruction Fetch Unit) and an EXU (Execution Unit). During the instruction fetch phase, the IFU reads CORDIC instructions from the instruction memory according to the address of the PC. The EXU includes a decoding module, an arithmetic logic unit, a load-memory unit, and a write-back unit. This invention extends the CORDIC angle calculation instruction set based on the RISC-V architecture and designs a complete pipelined hardware structure, supporting trigonometric function operations. It has a complete CORDIC operation path hardware circuit structure, reusing the same operation path for different operation processes, improving hardware resource utilization. It overcomes the limitation of convergence of the calculation result range when performing angle calculation based on the CORDIC algorithm by pre-storing special angle results to reduce calculation time and optimize the circuit operating frequency.
Owner:NORTH CHINA ELECTRIC POWER UNIV

Computer processing system and method configured to perform side-channel countermeasures

This invention presents a computer processing system and method designed to execute cryptographic operations while providing selective protection against side-channel attacks. It comprises a configuration of unprotected and protected hardware modules, the latter of which is equipped with data isolators, and a protected arithmetic logic unit (ALU) for secure data processing. The system enhances cryptographic security by selectively transmitting and computing input shares to generate side-channel protected output shares, ensuring robust protection during cryptographic operations.
Owner:PQSECURE TECHNOLOGIES LLC

Digital signal processing block

The present disclosure relates to digital signal processing blocks. A digital signal processor (DSP) chip is disclosed. The DSP chip includes an input stage for receiving a plurality of input signals; a pre-adder coupled to the input stage and configured to perform one or more operations on one or more of the plurality of input signals; and a multiplier coupled to the input stage and the pre-adder and configured to perform one or more multiplication operations on one or more of the plurality of input signals or an output of the pre-adder. The DSP chip also includes an arithmetic logic unit (ALU) coupled to the input stage, the pre-adder, and the multiplier. The ALU is configured to perform one or more mathematical or logical operations on one or more of the plurality of input signals, an output of the pre-adder, or an output of the multiplier. The DSP chip also includes an output stage coupled to the ALU, the output stage configured to generate one or more output signals based at least in part on one or more outputs of the ALU or at least one of the plurality of input signals.
Owner:XILINX INC

Multi-thread dynamic task scheduling circuit, computing chip and computing system

The invention relates to a multi-thread dynamic task scheduling circuit, a computing chip and a computing system, which are integrated in a parallel computing chip, and comprise a plurality of thread management units which are respectively used for storing context information of each task thread, reading an instruction stream from a memory through an internal instruction reading and transmitting unit, and decoding and transmitting the instruction stream; the arithmetic logic unit is used for executing arithmetic and logic operation instructions of the thread management unit so as to realize dynamic parameter derivation during operation; the task distribution unit is used for arbitrating the task configuration information transmitted by the plurality of thread management units and distributing the task configuration information to each processing unit in the chip; and the synchronization unit is used for synchronizing primitives through hardware so as to realize synchronization among the plurality of thread management units and synchronization between the thread management units and the processing unit. Each thread is programmable and adapts to complex scenes such as dynamic image sizes and dynamic calculation processes.
Owner:SHANGHAI JIAOTONG UNIV

Memory device, operating method thereof, and in-memory processing device

The invention discloses a memory device, an operating method thereof, and an in-memory processing device. The memory device includes an in-memory processing (PIM) block configured to perform an operation between a weight value and an input value, the weight value being represented by a weight scaling factor and a weight element, the input value being represented by an input scaling factor and an input element, where the PIM block includes: a first scaling register file storing the input scaling factor; a second scaling register file storing a weight scaling factor; a scalar register file (SRF) storing input elements; a plurality of arithmetic logic units (ALU) configured to perform a first operation between the input scaling factor and the weight scaling factor and a second operation between the input element and the weight element in parallel in response to an operation command received from a host; and an accumulator configured to accumulate and store operation results of the first operation and the second operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Arithmetic logic unit and imaging system

The present disclosure relates to image signal and phase detection autofocus signal extraction and storage in an arithmetic logic unit. An arithmetic logic unit (ALU) includes a front end latch stage coupled to a signal latch stage coupled to a Gray code (GC) to binary stage. A first input of an adder stage is coupled to receive an output of the GC to binary stage. An adder input latch stage includes first and second adder input latches including first and second inputs coupled to receive the output of the GC to binary stage. An adder input multiplexer stage includes an output coupled to a second input of the adder stage and first and second inputs coupled to outputs of the first and second adder input latches, respectively.
Owner:OMNIVISION TECHNOLOGIES INC

Method and apparatus for error detection in integer data processing

A system and method of preventing error propagation in variables and memory, including receiving, at an arithmetic logic unit, at least one input formatted in binary according to ones complement encoding, detecting a computation error or an input error associated with the at least one input, and outputting, from the arithmetic logic unit, a binary result with ones filling all positions as a NiN value in response to detecting the computation error or the input error.
Owner:ANDERSON MARK P

Single event upset irradiation test method based on modularized fifth-generation reduced instruction set soft core

The invention discloses a single event upset irradiation test method based on a modularized fifth-generation reduced instruction set soft core. The method is implemented on a ZYNQ chip. The method comprises the following steps: firstly, dividing a fifth-generation reduced instruction set soft core into a plurality of independent circuit modules according to functions, wherein the independent circuit modules comprise a decoder module, a jump processing module, an access control module, an arithmetic logic unit module, a data path module, a static memory module and a write-back circuit module; in the irradiation process, a program burnt by the ZYNQ chip outputs a running result data stream of each module to the upper computer in real time; the upper computer compares the received data with a pre-stored module expected correct value in real time; once data mismatch is found, an error event and a timestamp corresponding to the error event are immediately recorded, and circuit resources occupied by an error event occurrence bit are reversely positioned, so that module-level accurate positioning and vulnerability evaluation of errors caused by single event upset in the soft core are realized. The problem that the error source recognition rate is low when the whole soft core is tested is solved, and the fault resolution and the testing efficiency are remarkably improved.
Owner:XINJIANG TECH INST OF PHYSICS & CHEM CHINESE ACAD OF SCI

Compute optimizations for neural networks

One embodiment provides for a compute apparatus comprising a decode unit to decode a single instruction into a decoded instruction that specifies multiple operands including a multi-bit input value and a one-bit weight associated with a neural network, as well as an arithmetic logic unit including a multiplier, an adder, and an accumulator register. To execute the decoded instruction, the multiplier is to perform a fused operation including an exclusive not OR (XNOR) operation and a population count operation. The adder is configured to add the intermediate product to a value stored in the accumulator register and update the value stored in the accumulator register.
Owner:INTEL CORP

Processing unit with mixed precision arithmetic

A graphics processing unit (GPU) [100] implements an operation [105] having an associated operation code to perform a mixed-precision mathematical operation. The GPU includes an arithmetic logic unit (ALU) [104] having different execution paths [106, 107], where each execution path performs a different mixed-precision operation. By implementing the mixed-precision operation at the ALU in response to an operation code that specifies a description of the operation, the GPU efficiently improves the precision of the specified mathematical operation while reducing execution overhead.
Owner:ADVANCED MICRO DEVICES INC

An ALU, an instruction execution method, a processor, a device, a medium and a program

This disclosure provides an ALU, an instruction execution method, a processor, a device, a medium, and a program, relating to the field of computer technology, specifically information processing, deep learning, artificial intelligence, and chip technology. The ALU is integrated into the processor and includes a first data conversion module and a calculation module. The first data conversion module and the calculation module are connected, and the first data conversion module is also connected to a target register file. The first data conversion module receives first-precision data required for the calculation of a target instruction read from the target register file and converts the first-precision data into second-precision data; the precision of the first-precision data is lower than that of the second-precision data. The calculation module receives the second-precision data and executes the calculation operation of the target instruction based on the second-precision data. The embodiments of this disclosure can fully utilize hardware computing resources to improve instruction execution efficiency, thereby improving the overall execution performance of the arithmetic logic unit.
Owner:KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD

Convolutional neural networks with content-adaptive filters

Systems and methods to train a convolutional neural network having two or more filter layers having different filtering parameters corresponding to respective different portions of a digital representation of an image. A processor comprising one or more arithmetic logic units (ALUs) to be configured to identify one or more features within an image based, at least in part, on a convolutional neural network having two or more filter layers having different filtering parameters corresponding to respective different portions of a digital representation of the image.
Owner:NVIDIA CORP

Calculation instruction execution method and device, electronic equipment, medium and product

The invention provides a calculation instruction execution method and device, electronic equipment, a medium and a product, and relates to the field of artificial intelligence, in particular to the field of chips. According to the specific implementation scheme, a to-be-executed target calculation instruction is obtained; when the target calculation instruction meets a read-after-write condition for the target single-bit register, continuously constructing two target transmission instructions according to a current bit value in the target single-bit register; and respectively transmitting the two target transmitting instructions to an arithmetic logic unit to carry out calculation to obtain two calculation results, so that a target register in the target calculation instruction selects a correct calculation result from the two calculation results to complete data writing operation. When the read-after-write condition is met, two target transmitting instructions are constructed according to the current value of the target single-bit register and transmitted to the arithmetic logic unit for calculation, and a correct result is selected and written by the target register, so that the degree of dependence on the single-bit register is reduced, and the parallelism and efficiency of instruction execution are improved.
Owner:KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD

Drive state detection device

A drive state detection device includes a current detection unit configured to detect a drive current during driving of an inductive load; a voltage detection unit configured to detect an interterminal voltage during the driving of the inductive load; a variable filter configured to allow passage of a component of a predetermined frequency pass band, the component being of the drive current detected by the current detection unit; a signal generation unit configured to generate a pulse signal from a waveform of the drive current after the passage through the variable filter; and an arithmetic logic unit configured to detect a state of the inductive load in accordance with the drive current detected by the current detection unit, the interterminal voltage detected by the voltage detection unit, and the pulse signal generated by the signal generation unit.
Owner:ALPS ALPINE CO LTD

Generating iteration transfer information for code execution with a compute slice microarchitecture

A processor core is accessed. The core is configured to execute instructions associated with an instruction set architecture (ISA). The core comprises a plurality of compute slices, a plurality of barrier register files, and a control unit. Each compute slice includes at least one arithmetic logic unit (ALU), a local register file, and is coupled to a successor compute slice and a predecessor compute slice by a barrier register file. Code associated with the ISA is evaluated, where the code includes a first loop. The evaluating includes generating iteration transfer information associated with the first loop. Each slice task within a plurality of slice tasks associated with the first loop is distributed to a compute slice. The processor core executes the plurality of slice tasks. Data forwarding between successive compute slices is based on the plurality of barrier register files and the iteration transfer information.
Owner:ASCENIUM INC

Computer processing system and method configured to perform side-channel countermeasures

This invention presents a computer processing system and method designed to execute cryptographic operations while providing selective protection against side-channel attacks. It comprises a configuration of unprotected and protected hardware modules, the latter of which is equipped with data isolators, and a protected arithmetic logic unit (ALU) for secure data processing. The system enhances cryptographic security by selectively transmitting and computing input shares to generate side-channel protected output shares, ensuring robust protection during cryptographic operations.
Owner:PQSECURE TECHNOLOGIES LLC

Processing core with metadata actuated conditional graph execution

A processing core and associated methods for the efficient execution of a directed graph are disclosed. A disclosed processing core includes a memory and a first data tile stored in the memory. The first data tile includes a first set of data elements and metadata stored in association with the first set of data elements. The processing core also includes a second data tile stored in the memory. The second data tile includes a second set of data elements. The processing core also includes an arithmetic logic unit configured to conduct an arithmetic logic operation using data from the first set of data elements and the second set of data elements. The processing core also includes a control unit configured to evaluate the metadata and control the arithmetic logic unit to conditionally execute the arithmetic logic operation based on the evaluation of the metadata.
Owner:TENSTORRENT AI ULC

Enabling logical operations in low-power, double data rate (LPDDR) memory

A memory apparatus includes two independent subchannels of double data rate (DDR) memory. Each subchannel includes a number of memory banks, and an input / output (I / O) block having a number of meta data registers. The memory apparatus also includes a processing unit. The processing unit includes an arithmetic logic unit (ALU), an accumulator coupled to an output of the ALU, a controller coupled to the ALU, two multiplexers, each multiplexer coupling one subchannel to an ALU input, and coupling the accumulator to the ALU input, and two multiplexers / demultiplexers. Each multiplexer / demultiplexer is coupled between a memory bank, a corresponding I / O block, one of the two multiplexers, and the accumulator.
Owner:QUALCOMM INC