Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

104 results about "Arithmetic logic unit" patented technology

An arithmetic logic unit (ALU) is a combinational digital electronic circuit that performs arithmetic and bitwise operations on integer binary numbers. This is in contrast to a floating-point unit (FPU), which operates on floating point numbers. An ALU is a fundamental building block of many types of computing circuits, including the central processing unit (CPU) of computers, FPUs, and graphics processing units (GPUs). A single CPU, FPU or GPU may contain multiple ALUs.

Memory device and method

A memory device includes a plurality of memory banks, and a processing-in-memory (PIM) block accessible to the plurality of memory banks, wherein the PIM block comprises a control circuit configured to receive a plurality of operation instructions from a host and, in response to a predicated instruction indicating a predication operation among the plurality of operation instructions, instruct an arithmetic logic unit (ALU) to perform the predication operation, a predicate register file (PRF) configured to store therein a predicate value determined by the predication operation, and the ALU configured to perform an operation according to a command signal translated by the control circuit based on the predicate value from an operation instruction that depends on the predicate value among the plurality of operation instructions.
Owner:SAMSUNG ELECTRONICS CO LTD +1

Precision target optimization method and system adaptive to variable precision arithmetic logic unit, medium, terminal and program product

The invention provides a precision target optimization method and system adaptive to a variable precision arithmetic logic unit, a medium, a terminal and a program product. The method comprises the following steps: acquiring an output feature set of each group of an upper layer; the precision generation network layer generates a corresponding precision target according to the output feature set, and the ALU calculation layer generates a prediction result according to the generated precision target; the teacher model generates a reference target and a real label according to the output feature set; constructing a total loss function according to the calculated task loss, precision generation loss and adversarial loss; performing back propagation optimization on the student model based on the constructed total loss function; repeatedly and iteratively training the student model until convergence to obtain a final student model; and deploying the final student model to generate an optimal precision target corresponding to each group. According to the method provided by the invention, the fine precision adjustment of the bit granularity can be realized, the adaptive ability of the model is enhanced, and the matching degree of the precision and the task demand is improved.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Vector engine for efficient beam search

To provide native hardware support for complex operations such as a beam search operation, an integrated circuit device can be implemented with multiple computational circuit blocks coupled in series to form a pipeline. Each computational circuit block includes an arithmetic logic unit (ALU) circuit having a first numeric input, a second numeric input, a primary result output, and a secondary output. The ALU circuit is programmable to perform an arithmetic operation on the first numeric input and the second numeric input to generate the primary result output, and provide one of the first numeric input or the second numeric input to the secondary output, which is fed back to the ALU circuit.
Owner:AMAZON TECH INC

Greedy algorithm for in network computation trees

Techniques and architecture are described for a method that includes an in network compute (INC) manager receiving from switches of a fat tree configured network, arithmetic logic unit (ALU) capacity of the switches. Based at least in part on the ALU capacity of the switches and bandwidth, the INC manager determines one or more switches within each tier that are capable of supporting the processing units and based at least in part on the determining, the INC manager selects a first switch as a root, wherein the first switch is included within a tier of switches having intermediate tiers of switches located between the tier and the plurality of processing units within the fat tree configured network. The INC manager creates one or more paths of switches within each of the intermediate tiers from the root to the plurality of processing units to provide a constrained disjoint spanning tree of switches.
Owner:CISCO TECHNOLOGY INC

Multi-thread dynamic task scheduling circuit, computing chip and computing system

The invention relates to a multi-thread dynamic task scheduling circuit, a computing chip and a computing system, which are integrated in a parallel computing chip, and comprise a plurality of thread management units which are respectively used for storing context information of each task thread, reading an instruction stream from a memory through an internal instruction reading and transmitting unit, and decoding and transmitting the instruction stream; the arithmetic logic unit is used for executing arithmetic and logic operation instructions of the thread management unit so as to realize dynamic parameter derivation during operation; the task distribution unit is used for arbitrating the task configuration information transmitted by the plurality of thread management units and distributing the task configuration information to each processing unit in the chip; and the synchronization unit is used for synchronizing primitives through hardware so as to realize synchronization among the plurality of thread management units and synchronization between the thread management units and the processing unit. Each thread is programmable and adapts to complex scenes such as dynamic image sizes and dynamic calculation processes.
Owner:SHANGHAI JIAOTONG UNIV

Memory device, operating method thereof, and in-memory processing device

The invention discloses a memory device, an operating method thereof, and an in-memory processing device. The memory device includes an in-memory processing (PIM) block configured to perform an operation between a weight value and an input value, the weight value being represented by a weight scaling factor and a weight element, the input value being represented by an input scaling factor and an input element, where the PIM block includes: a first scaling register file storing the input scaling factor; a second scaling register file storing a weight scaling factor; a scalar register file (SRF) storing input elements; a plurality of arithmetic logic units (ALU) configured to perform a first operation between the input scaling factor and the weight scaling factor and a second operation between the input element and the weight element in parallel in response to an operation command received from a host; and an accumulator configured to accumulate and store operation results of the first operation and the second operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Arithmetic logic unit and imaging system

The present disclosure relates to image signal and phase detection autofocus signal extraction and storage in an arithmetic logic unit. An arithmetic logic unit (ALU) includes a front end latch stage coupled to a signal latch stage coupled to a Gray code (GC) to binary stage. A first input of an adder stage is coupled to receive an output of the GC to binary stage. An adder input latch stage includes first and second adder input latches including first and second inputs coupled to receive the output of the GC to binary stage. An adder input multiplexer stage includes an output coupled to a second input of the adder stage and first and second inputs coupled to outputs of the first and second adder input latches, respectively.
Owner:OMNIVISION TECHNOLOGIES INC

Method and apparatus for error detection in integer data processing

A system and method of preventing error propagation in variables and memory, including receiving, at an arithmetic logic unit, at least one input formatted in binary according to ones complement encoding, detecting a computation error or an input error associated with the at least one input, and outputting, from the arithmetic logic unit, a binary result with ones filling all positions as a NiN value in response to detecting the computation error or the input error.
Owner:ANDERSON MARK P

Single event upset irradiation test method based on modularized fifth-generation reduced instruction set soft core

The invention discloses a single event upset irradiation test method based on a modularized fifth-generation reduced instruction set soft core. The method is implemented on a ZYNQ chip. The method comprises the following steps: firstly, dividing a fifth-generation reduced instruction set soft core into a plurality of independent circuit modules according to functions, wherein the independent circuit modules comprise a decoder module, a jump processing module, an access control module, an arithmetic logic unit module, a data path module, a static memory module and a write-back circuit module; in the irradiation process, a program burnt by the ZYNQ chip outputs a running result data stream of each module to the upper computer in real time; the upper computer compares the received data with a pre-stored module expected correct value in real time; once data mismatch is found, an error event and a timestamp corresponding to the error event are immediately recorded, and circuit resources occupied by an error event occurrence bit are reversely positioned, so that module-level accurate positioning and vulnerability evaluation of errors caused by single event upset in the soft core are realized. The problem that the error source recognition rate is low when the whole soft core is tested is solved, and the fault resolution and the testing efficiency are remarkably improved.
Owner:XINJIANG TECH INST OF PHYSICS & CHEM CHINESE ACAD OF SCI

An ALU, an instruction execution method, a processor, a device, a medium and a program

This disclosure provides an ALU, an instruction execution method, a processor, a device, a medium, and a program, relating to the field of computer technology, specifically information processing, deep learning, artificial intelligence, and chip technology. The ALU is integrated into the processor and includes a first data conversion module and a calculation module. The first data conversion module and the calculation module are connected, and the first data conversion module is also connected to a target register file. The first data conversion module receives first-precision data required for the calculation of a target instruction read from the target register file and converts the first-precision data into second-precision data; the precision of the first-precision data is lower than that of the second-precision data. The calculation module receives the second-precision data and executes the calculation operation of the target instruction based on the second-precision data. The embodiments of this disclosure can fully utilize hardware computing resources to improve instruction execution efficiency, thereby improving the overall execution performance of the arithmetic logic unit.
Owner:KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD

Convolutional neural networks with content-adaptive filters

Systems and methods to train a convolutional neural network having two or more filter layers having different filtering parameters corresponding to respective different portions of a digital representation of an image. A processor comprising one or more arithmetic logic units (ALUs) to be configured to identify one or more features within an image based, at least in part, on a convolutional neural network having two or more filter layers having different filtering parameters corresponding to respective different portions of a digital representation of the image.
Owner:NVIDIA CORP

Generating iteration transfer information for code execution with a compute slice microarchitecture

A processor core is accessed. The core is configured to execute instructions associated with an instruction set architecture (ISA). The core comprises a plurality of compute slices, a plurality of barrier register files, and a control unit. Each compute slice includes at least one arithmetic logic unit (ALU), a local register file, and is coupled to a successor compute slice and a predecessor compute slice by a barrier register file. Code associated with the ISA is evaluated, where the code includes a first loop. The evaluating includes generating iteration transfer information associated with the first loop. Each slice task within a plurality of slice tasks associated with the first loop is distributed to a compute slice. The processor core executes the plurality of slice tasks. Data forwarding between successive compute slices is based on the plurality of barrier register files and the iteration transfer information.
Owner:ASCENIUM INC

Processing core with metadata actuated conditional graph execution

A processing core and associated methods for the efficient execution of a directed graph are disclosed. A disclosed processing core includes a memory and a first data tile stored in the memory. The first data tile includes a first set of data elements and metadata stored in association with the first set of data elements. The processing core also includes a second data tile stored in the memory. The second data tile includes a second set of data elements. The processing core also includes an arithmetic logic unit configured to conduct an arithmetic logic operation using data from the first set of data elements and the second set of data elements. The processing core also includes a control unit configured to evaluate the metadata and control the arithmetic logic unit to conditionally execute the arithmetic logic operation based on the evaluation of the metadata.
Owner:TENSTORRENT AI ULC

Enabling logical operations in low-power, double data rate (LPDDR) memory

A memory apparatus includes two independent subchannels of double data rate (DDR) memory. Each subchannel includes a number of memory banks, and an input / output (I / O) block having a number of meta data registers. The memory apparatus also includes a processing unit. The processing unit includes an arithmetic logic unit (ALU), an accumulator coupled to an output of the ALU, a controller coupled to the ALU, two multiplexers, each multiplexer coupling one subchannel to an ALU input, and coupling the accumulator to the ALU input, and two multiplexers / demultiplexers. Each multiplexer / demultiplexer is coupled between a memory bank, a corresponding I / O block, one of the two multiplexers, and the accumulator.
Owner:QUALCOMM INC

Methods and apparatuses for an arithmetic logic unit of a computational processor

A circuit is configured for transforming input data in a neural network to output data. The circuit includes inter-connected arithmetic logic circuits and a control circuit sending a sequence of configuration settings to configure the inter-connected arithmetic logic circuits over one or more cycles. The inter-connected arithmetic logic circuits jointly perform at least one of a softmax or a sigmoidal linear unit (SiLU) operation on the input data. The inter-connected arithmetic logic circuits include an exponential circuit generating a first exponential of a fractional part of an input and a second exponential of an integer part of the input, a multiplier circuit multiplying the first exponential and the second exponential into a first intermediate variable, an accumulator circuit summing multiple intermediate variables relating to exponentials of multiple inputs, and a reciprocal circuit generating a reciprocal of a second intermediate variable that is input to the reciprocal circuit.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Sparse 4C2+ phase detection auto focus and correlated multiple sampling

An arithmetic logic unit includes a GC to binary stage, an adder stage, an adder output stage, an adder input latch stage coupled to latch outputs of the GC to binary stage, a feedback multiplexer stage coupled to receive the outputs of the GC to binary stage, a latch output multiplexer coupled to receive outputs of the first adder input latches, where the latch output multiplexer is configured to multiply the outputs of the first adder input latches by either −1 or −2, and an adder input multiplexer stage, where first inputs of the adder input multiplexer stage are coupled to receive outputs of the latch output multiplexer and second inputs of the adder input multiplexer stage are coupled to receive outputs of the second adder input latches. The arithmetic logic performs adaptive correlated multiple sampling for image sensing pixels and phase detection auto focus for other pixels.
Owner:OMNIVISION TECHNOLOGIES INC

Memory device and method with processing-in-memory block

A memory device includes a first scalar register file storing a first input fragment, a second scalar register file storing a second input fragment, an arithmetic logic unit (ALU), and a control circuit. The control circuit is configured to perform, using the ALU, a first operation between the first input fragment and a first weight fragment based on a first operation command received from a host, and to perform, using the ALU, a second operation between the second input fragment and a second weight fragment based on a second operation command received from the host.
Owner:SAMSUNG ELECTRONICS CO LTD

Phase-locked loop, control method and frequency source

PendingCN122293080AConvertersControl signal
This application provides a phase-locked loop (PLL), a control method, and a frequency source. It includes a controller, a time-to-digital converter (TD-to-digital converter), an arithmetic logic unit (ALU), a digitally controlled oscillator (DCO), and a frequency divider. The DCO outputs a linearly swept frequency signal based on a frequency control signal from the ALU. The frequency divider generates a divided frequency signal based on the output frequency of the DCO. The TD-to-digital converter outputs the phase difference between the divided frequency signal and a reference signal. The ALU outputs a frequency control signal based on the phase difference and a control code to adjust the output frequency of the DCO. The controller acquires the output signal of the TD-to-digital converter after a step point in the division ratio, and adjusts the control code until the output signal of the TD-to-digital converter after the step point is zero when the output signal is not zero. This method improves the linearity of the sawtooth waveform generated by the PLL at the step point.
Owner:SHANGHAI INTEGRATED CIRCUIT RESEARCH & DEVELOPMENT CENTER CO LTD

Sparse 4c2+ phase detection auto focus and correlated multiple sampling

PendingUS20260095683A1Arithmetic logic unitMultiplexing
An arithmetic logic unit includes a GC to binary stage, an adder stage, an adder output stage, an adder input latch stage coupled to latch outputs of the GC to binary stage, a feedback multiplexer stage coupled to receive the outputs of the GC to binary stage, a latch output multiplexer coupled to receive outputs of the first adder input latches, where the latch output multiplexer is configured to multiply the outputs of the first adder input latches by either −1 or −2, and an adder input multiplexer stage, where first inputs of the adder input multiplexer stage are coupled to receive outputs of the latch output multiplexer and second inputs of the adder input multiplexer stage are coupled to receive outputs of the second adder input latches. The arithmetic logic performs adaptive correlated multiple sampling for image sensing pixels and phase detection auto focus for other pixels.
Owner:OMNIVISION TECHNOLOGIES INC

Method for solving multiple parallel read-write access conflicts of memory and artificial intelligence chip

The invention provides a memory multi-parallel read-write access conflict solution method, a chip storage circuit and a computing device, the memory multi-parallel read-write access conflict solution comprises a plurality of arithmetic logic units, a storage management module and a memory, sending a parallel read-write access request to the memory through a storage management module, and obtaining data required by matrix operation; the memory receives parallel read-write access requests from the plurality of arithmetic logic units and provides data required by matrix operation for the plurality of arithmetic logic units, and the memory is configured to comprise a specific number of storage sub-modules according to the degree of parallelism of data access of the plurality of arithmetic logic units, the specific number is a prime number. According to the technical scheme, in the data parallel read-write access process, access conflicts can be reduced, and the parallel efficiency of data read-write is improved.
Owner:SHENZHEN CORERAIN TECH CO LTD

Enabling logical operations in low-power, double data rate (LPDDR) memory

A memory apparatus includes two independent subchannels of double data rate (DDR) memory. Each subchannel includes a number of memory banks, and an input / output (I / O) block having a number of meta data registers. The memory apparatus also includes a processing unit. The processing unit includes an arithmetic logic unit (ALU), an accumulator coupled to an output of the ALU, a controller coupled to the ALU, two multiplexers, each multiplexer coupling one subchannel to an ALU input, and coupling the accumulator to the ALU input, and two multiplexers / demultiplexers. Each multiplexer / demultiplexer is coupled between a memory bank, a corresponding I / O block, one of the two multiplexers, and the accumulator.
Owner:QUALCOMM INC

Fused comparison add instructions

An apparatus, system, and method for efficiently processing pairs of operations repeatedly used in applications. In various implementations, a computing system includes a parallel data processing circuit with multiple compute circuits. Each of the compute circuits includes multiple lanes of execution, each with a corresponding arithmetic logic unit (ALU). The ALU supports executing a single fused conditional ternary instruction that replaces two separate instructions that provide two operations (comparison and add). When executing the fused conditional ternary instruction, the ALU does not retrieve the intermediate result from the scalar register file, the vector register file, or bypass circuitry located externally from the ALU. Rather, the ALU generates the intermediate result and uses the intermediate result without routing the intermediate result externally from ALU.
Owner:ADVANCED MICRO DEVICES INC

Non-private and private inference scheduling method and device based on reconfigurable chip

The application provides a non-privacy and privacy reasoning scheduling method based on a reconfigurable chip, which comprises the following steps: determining whether the privacy requirement of a neural network reasoning task is privacy reasoning; if yes, performing a first scheduling step; otherwise, performing a second scheduling step; in the first scheduling step, analyzing an arithmetic logic unit required by the privacy reasoning, combining a reconfigurable multiplier and an adder by a data layout converter to form the arithmetic logic unit, and reconfiguring an interconnection network of a processing unit array into a butterfly network or a SIMD data path to obtain a reconfigurable chip and perform a reasoning step; in the second scheduling step, setting a slice as a tensor mode by the data layout converter, independently and parallelly working all the slices, performing a multiply-accumulate operation by each slice, transmitting data between processing units PEs through horizontal and vertical links to form a pulsating data stream, forming a pulsating array by the processing unit array, obtaining the reconfigurable chip, and performing the reasoning step; and in the reasoning step, performing the neural network reasoning task by the reconfigurable chip to obtain a reasoning result.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Techniques for a physical side-channel-resilient arthmetic logic unit

Examples include techniques for a physical side-channel resilient arithmetic logic unit (ALU) The examples include use of circuitry to input a shared representation of a value having a plurality of shares to the ALU. Then, at the ALU, recombine the shared representation of the value to a single representation of the value and use the single representation to generate a result value or operate on the shared representation of the value to generate a shared representation of the result value. If the result value is based on operating on the single representation of the value, the result value is split to a shared representation of the result value. The shared representation of the result value can be output from the ALU and stored to at least one register from among a plurality of registers.
Owner:INTEL CORP

Dataflow-based field programmable gate array inference acceleration method and related apparatus

The application belongs to an edge intelligent computing optimization acceleration method, through multiple traversal of a computation graph, difference storage planning of fixed data and dynamic data is completed, and non-core operators are packaged and combined into an arithmetic logic unit module processing block to reduce data transfer and scheduling overhead. The division and cooperation of the arithmetic logic unit module and the general matrix multiplication module in the FPGA accelerator are used to process non-core operators and core operators respectively, and the smoothness of the calculation process is ensured through the interaction of the working states between the modules. The application optimizes the whole link from data storage, operator reconstruction to hardware execution, solves the problems of frequent data transfer and low hardware resource utilization in traditional schemes, and fully develops the reconfigurable characteristics and parallel computing advantages of FPGA through special module division.
Owner:CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2

Touch control device, touch control method, and robot

This application discloses a touch control device, comprising: a touch detection circuit and an arithmetic logic unit, wherein: the touch detection circuit is connected to the arithmetic logic unit; the touch detection circuit is configured to generate a first level for each touch electrode; the arithmetic logic unit is configured to receive a plurality of first levels output by the touch detection circuit, perform logical operations on the plurality of first levels to obtain a second level, and output the second level to the touch detection circuit; the touch detection circuit is further configured to determine the target touch electrode touched by the user from the plurality of touch electrodes based on the first level and the received second level. This application also discloses a touch control method and a robot.
Owner:JINGDONG KUNPENG (JIANGSU) TECH CO LTD

Method and memory module for performing in-memory computation

Methods and memory modules for performing in-memory computation are provided. The memory module includes a memory die including a plurality of dynamic random access memory (DRAM) banks, each DRAM bank including an array of DRAM cells arranged in pages, a row buffer storing a value of one of the pages, an input / output (IO) module, and an in-memory computation (IMC) module including an arithmetic logic unit (ALU) receiving operands from the row buffer or the IO module and computing an output based on the operands and one of a plurality of ALU operations, and a result register storing the output of the ALU, and a controller receiving operands and instructions from a host processor, determining a data layout based on the instructions, supplying the operands to the DRAM banks according to the data layout, and controlling the IMC module to perform one of the ALU operations on the operands according to the instructions.
Owner:SAMSUNG ELECTRONICS CO LTD

System and methods for data compression in low power double data rate-processing in memory on mobile system on chip

A device, system, and method are disclosed for processing-in-memory (PIM) compression. In an embodiment, a method includes obtaining, by an input / output sense amplifier (IOSA) from an associated RAM bank, compressed data; sending, by the IOSA, the compressed data divided into a plurality of portions; receiving, by a respective decompressor of a plurality of decompressors of a PIM block associated with the RAM bank, a respective portion of the compressed data; decompressing, by the respective decompressor, the respective portion of the compressed data to obtain a respective portion of decompressed data; and sending, by the respective decompressor, the respective portion of the decompressed data to a respective Arithmetic Logic Unit (ALU) of a plurality of ALUs of the PIM block for processing.
Owner:SAMSUNG ELECTRONICS CO LTD

Accelerated mathematical engine

Various embodiments of the disclosure relate to an accelerated mathematical engine. In certain embodiments, the accelerated mathematical engine is applied to image processing such that convolution of an image is accelerated by using a two-dimensional matrix processor comprising sub-circuits that include an ALU, output register and shadow register. This architecture supports a clocked, two-dimensional architecture in which image data and weights are multiplied in a synchronized manner to allow a large number of mathematical operations to be performed in parallel.
Owner:TESLA INC