Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

319 results about "Operand" patented technology

In mathematics an operand is the object of a mathematical operation, i.e., it is the object or quantity that is operated on.

Heterogeneous processor-oriented reciprocal calculation instruction sequence generation method

The invention discloses a reciprocal calculation instruction sequence generation method oriented to a heterogeneous processor, and belongs to the field of compilation optimization and code generation. Aiming at the problems of instruction redundancy, weak precision control, poor hardware adaptation and high manual dependence of an existing method in a heterogeneous environment, characteristics of a reciprocal instruction and an operand are accurately identified by linearly scanning heterogeneous object codes (including vectorization, scalar and complex instruction sequences); in combination with hardware characteristics of RISC / SIMD / VLIW / DSP and the like, a multi-round iteration precision improvement and temporary register optimization allocation strategy is adopted, differential generation logic is formulated, and a high-precision low-redundancy instruction sequence is generated. The method comprises linear code scanning classification, reciprocal instruction and operand identification, cross-architecture generation logic rule formulation, instruction sequence generation and legality verification. Full-process automation is achieved, manual intervention is reduced, the execution efficiency and precision of reciprocal calculation of the heterogeneous processor are improved, and the method is suitable for embedded systems, high-performance calculation and other scenes.
Owner:HUNAN UNIV OF SCI & TECH

SRT operational circuit

The SRT operational circuit comprises an input module, a floating point conversion module, a calculation module and an output module, the input module is configured to output an initial operand, a first end of the floating point conversion module is connected with the input module, a first end of the calculation module is connected with the floating point conversion module and the input module, and a second end of the calculation module is connected with the output module. The output module is connected with the second end of the calculation module; when the initial operand is an integer, the floating point conversion module is configured to perform floating point conversion on the initial operand to output a first operand; if the calculation module is configured to perform division operation or root extraction operation by adopting the first operand based on the SRT algorithm, the output module is configured to perform mantissa rounding on a division result or a root extraction result after the calculation of the calculation module is completed and output a floating point result, so that the operation accuracy of an integer in the circuit is improved, and the universality of the circuit is improved.
Owner:GUANGDONG LEAPFIVE TECH CO LTD

Methods and circuits for streaming data to processing elements in stacked processor-plus-memory architecture

A stacked processor-plus-memory device includes a processing die with an array of processing elements of an artificial neural network. Each processing element multiplies a first operand—e.g. a weight—by a second operand to produce a partial result to a subsequent processing element. To prepare for these computations, a sequencer loads the weights into the processing elements as a sequence of operands that step through the processing elements, each operand stored in the corresponding processing element. The operands can be sequenced directly from memory to the processing elements or can be stored first in cache. The processing elements include streaming logic that disregards interruptions in the stream of operands.
Owner:RAMBUS INC

Floating point arithmetic device and method of operating the same

A floating point arithmetic device with two floating point operands and its operation method are disclosed. The floating point arithmetic device includes an exponent subtraction circuit, an exponent calculation circuit, a mantissa calculation circuit, and a conversion circuit. The exponent subtraction circuit calculates the difference between the exponents of the two operands and generates a sign bit and an exponent difference. The exponent calculation circuit generates the post-operation exponent bits according to the larger one of the exponents of the two operands. The mantissa calculation circuit aligns the mantissa bits of the two operands and performs one of addition and subtraction on the aligned mantissa bits. To improve the calculation efficiency and reduce the power consumption, the floating point arithmetic device can complete the floating point addition or subtraction operation in one step (one clock cycle) without moving the intermediate floating point data between the registers and the functional circuit units as in the multi-step operation.
Owner:XINLIJIA INTEGRATED CIRCUIT (SHANGHAI) CO LTD

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Loading storage circuit and graphics processor

The invention provides a loading storage circuit and a graphics processor, and relates to the technical field of graphics processing. The loading storage circuit comprises an instruction scheduling module, an address generation module and a data service module, and is provided with a buffer area module which comprises an operand buffer area and an effective data buffer area and is used for caching information required by address calculation and data access; the instruction scheduling module is used for collecting a data access instruction and outputting instruction information; the address generation module calculates a target access address according to the instruction information and the operand; and the data service module executes corresponding data loading or storage operation on the target cache unit. According to the scheme, the operands and the valid data are pre-cached, so that overflow and stagnation caused by inconsistent processing rhythms among modules can be avoided, and the parallel processing capability of an assembly line is improved; through the independent address calculation and data access process, the stability of the access time sequence and the data processing efficiency can be improved.
Owner:MOORE THREADS TECH CO LTD

Efficient execution of group-sparsified neural networks

Creating neural network (NN) code may include for each row in a kernel matrix, finding the first non-zero element; and creating a loop including multiply and add instructions. On each iteration of the loop, the multiply and add instructions may be executed, and the position of the kernel matrix operand operated on by each multiply and add may be correlated to the loop iteration number. Instructions may be issued or created to be executed in the loop. A method may execute a NN by executing a loop including a series of multiply and add instructions to multiply a kernel matrix A by an input, such that on each iteration of the loop the series of multiply and add instructions are executed; and the position of the matrix A operand operated on by each multiply and add instruction in the series is correlated to the iteration number of the loop.
Owner:RED HAT INC

ALU operation fusion processing module and method suitable for neural network

The invention discloses an ALU operation fusion processing module and method suitable for a neural network, and the module comprises a control unit which is used for receiving and decoding a machine instruction, managing the execution processes of internal and external circulation and microinstruction circulation, and generating a control signal of each stage of a microinstruction assembly line; the microinstruction buffer area is used for storing a microinstruction sequence pre-generated by the neural network compiler; the register file is used for storing source operands and results of ALU operation; each entry of the register file is composed of a valid bit, a tag bit and a data bit; the ALU computing core adopts an SIMD (Single Instruction Multiple Data) architecture and comprises a plurality of paths of parallel arithmetic logic function units; the Load / Store unit is used for processing data exchange between a register file and a local buffer area; and the data selection interface is used for selecting a data source or a target buffer area according to the storage tag field of the microinstruction. According to the method, the high efficiency and the flexibility of the ALU in the neural network hardware accelerator can be effectively considered.
Owner:ZHEJIANG UNIV

Universal computing unit and instruction scheduling method

The invention provides a general purpose computing unit and an instruction scheduling method for the general purpose computing unit. The general-purpose computing unit comprises a plurality of execution units, wherein each execution unit is used for executing an operation instruction by taking a thread bundle as a unit; and an instruction scheduler for determining a priority of the plurality of operation instructions based on whether the multiplexing flag exists, and scheduling each operation instruction to one of the plurality of execution units based on the priority; wherein each execution unit comprises one or more reuse registers, each reuse register is used for registering an operand of a previous operation instruction and a label, and the label is used for indicating an address and a thread bundle of the operand; and the instruction decoding unit is used for decoding the received operation instruction to determine the reading position of the operand of the operation instruction. The instruction is preferentially scheduled by adding a reuse mark to the whole operation instruction, so that the hit rate of the operand is improved, and excessive instruction bit fields do not need to be occupied.
Owner:SHANGHAI BIREN TECH CO LTD

Convolutional neural network hardware acceleration method and system

The invention relates to the technical field of hardware acceleration, and provides a convolutional neural network hardware acceleration method and system, and the method comprises the steps: obtaining a convolutional neural network instruction which comprises an img2col instruction, a systolic array multiplication instruction and a col2img instruction; analyzing an operation instruction code and a function code field of each instruction, determining a target register, a source register and calculation parameters, determining an instruction function and generating a control signal; data to be operated are taken out from the source register and transmitted to the corresponding operation module together with the control signal and the calculation parameter to be calculated, and then the data are written back to the target register; through hardware optimization strategies such as multi-level address superposition, systolic array data multiplexing, double-buffer window construction and a parallel comparison tree, extra power consumption caused by multiple times of data loading and storage and intermediate result calculation is reduced, and data handling overhead, calculation delay and control complexity are greatly reduced.
Owner:SHANDONG UNIV

Instant instruction fusion method and device, electronic equipment and computer program product

PendingCN121277564AMachine execution arrangementsIndex registerRename
The invention provides an immediate operand instruction fusion method, an immediate operand instruction fusion device, electronic equipment and a computer program product. In a renaming stage, aiming at a RAW instruction having data read-after-write RAW correlation with a first loading immediate operand instruction, according to a source logic register number index register mapping table of the RAW instruction, the first loading immediate operand instruction is subjected to data read-after-write RAW correlation; obtaining a first status bit and a first table item of a source operand of the RAW instruction; and obtaining a fusion instruction by combining the RAW instruction with a first immediate value of a first loading immediate number instruction stored in the first table item under the condition that the table item storage content corresponding to the source logic register number is determined to be the immediate value according to the first status bit. Therefore, the immediate operand is directly fused to the RAW instruction in the renaming stage by utilizing the extended register mapping table, so that the dependence of instruction fusion on the adjacent program sequence is broken, a consumer instruction can obtain the operand without waiting for the execution of the immediate operand loading instruction, and the instruction execution delay is remarkably reduced.
Owner:GUANGDONG LEAPFIVE TECH CO LTD

Adaptive selection of a configuration-dependent operand

Described apparatuses and methods provide adaptive selection of a configuration-dependent operand. Adaptive selection enables a die to automatically detect its configuration and dynamically read from or write to configuration-dependent operands of one or more mode registers based on the configuration. In this manner, a memory device with dies having a first byte width (e.g., 1 byte) can be transparent to a memory channel having a second byte width that is larger than the first byte width (e.g., two bytes). For example, two 8-bit dies enabled with aspects of adaptive selection may be coupled to a 16-bit memory channel and each automatically detect its byte position (e.g., upper byte or lower byte). Based on the detected byte position or byte indicator, the 8-bit dies may be configured for use through respective mode registers that correspond to the byte position or assignment of the die.
Owner:MICRON TECHNOLOGY INC

Processing unit based on redundant remainder system, in-memory processing device and in-memory processing system

A processing unit includes a redundant remainder generation circuit that converts first operand data and second operand data into a plurality of first operand redundant remainder sets and a plurality of second operand redundant remainder sets based on a redundant remainder system (RRNS) using first to Tth moduli. T is a natural number equal to or greater than 4. The processing unit includes a plurality of arithmetic circuits that perform operations using the plurality of first operand redundant remainder sets and the plurality of second operand redundant remainder sets, and a reconstruction circuit that restores the RRNS-based result data to the weighted coefficient-based result data. An in-memory processing (PIM) apparatus and an in-memory processing (PIM) system are also provided.
Owner:SK HYNIX INC

Fully homomorphic encryption (FHE) operations on a unified FHE accelerator

PendingUS20260189363A1Theoretical computer scienceModular arithmetic
Techniques for fully homomorphic encryption are described. In some examples, a fully homomorphic encryption includes a butterfly compute circuitry to support a polynomial integer multiplication in response to an instance of a single instruction of a first type, wherein the instance of the single instruction is at least to include one or more fields for a register file address for a first source operand of the integer multiplication, one or more fields for a register file address for a second source operand of the integer multiplication, and one more fields for a register file address for a result of the integer multiplication, wherein the butterfly compute circuitry is to additionally support modular arithmetic operations.
Owner:INTEL CORP

Systems and methods for implementing directional operand broadcast and multiply-accumulate execution using a configurable patch mesh in a multi-core processing array of an integrated circuit

A technique is disclosed for operand propagation and accumulation within a processing array of an integrated circuit using overlapping patch regions. The system includes an interconnecting processing patch defined over a rectilinear subset of processing elements, with an origin processing element broadcasting operand data to the remaining elements in a directionally constrained, time-staggered wavefront pattern. A logical processing patch is separately defined over a second rectilinear subset of processing elements. The interconnecting processing patch and the logical processing patch partially overlap to form an interconnecting patch mesh comprising a common set of processing elements. Operand data is propagated from the origin of the interconnecting patch to the common processing elements within the patch mesh, enabling operand handoff or accumulation across patch boundaries. The architecture supports fine-grained, localized data movement and patch-level execution coordination across a mesh of processing elements to optimize compute reuse, operand locality, and execution throughput.
Owner:QUADRIC IO INC

Compute in memory architecture that allows flexible floating-point operations

PCT designated stageWO2026178063A1Floating pointOperand
The disclosure describes systems and methods for mixed‑mode, mixed‑precision computation using a configurable compute engine for vector–matrix operations. A processing system selects an operating mode including at least a floating‑point mode or an integer mode for the compute engine. Operand storage and format preprocessing route differently sized element representations through shared lanes. In floating‑point mode, compute circuitry generates a mantissa product and a product exponent, and alignment applies power‑of‑two control through a selection network, including as part of a fused multiply‑accumulate path. Parallel adder trees feed a hierarchical output accumulator sized to the selected precision. Normalization operates at a wider intermediate precision, applies block‑level scaling with dequantization, and, in lower‑precision operation, performs groupwise normalization with an intermediate‑precision combine to preserve a constant compute‑block size. Implementations integrate the compute engine within a memory macro with double buffering or deploy the compute engine as a standalone accelerator.
Owner:OPENAI OPCO LLC

Non-rectangular matrix computations and data pattern processing using tensor cores

Matrix multiplication operations can be implemented, at least in part, on one or more tensor cores of a parallel processing unit. An efficiency of the matrix multiplication operations can be improved in cases where one of the input operands or the output operand of the matrix multiplication operation is a square matrix having a triangular data pattern. In such cases, the number of computations performed by the tensor cores of the parallel processing unit can be reduced by dropping computations and / or masking out elements of the square matrix input operand on one side of the main diagonal of the square matrix. In other cases where the output operand exhibits the triangular data pattern, computations can be dropped or masked out for the invalid side of the main diagonal of the square matrix. In an embodiment, a library implementing the matrix multiplication operations is provided.
Owner:NVIDIA CORP

Application programming interface for indicating stepped memory

The invention relates to an application programming interface for indicating strided memory. Apparatuses, systems, and methods for performing thread memory addressing are disclosed. In at least one embodiment, a processor includes one or more circuitry to execute an application programming interface (API) to cause execution of one or more instructions based at least in part on one or more API parameters, the one or more API parameters indicate a size of one or more operands to be used by the one or more instructions.
Owner:NVIDIA CORP

Processor operand management using fused buffers

Techniques relating to operand management using a fused buffer are disclosed. A processor includes an operand management circuit, wherein the operand management circuit includes a fusion buffer; and an execution circuit. In one embodiment, operand management circuitry is configured to: detect a first store instruction operation executable to store an operand value usable by one or more consumer instruction operations; and storing the first store instruction operation in the fusion buffer. In response to detecting a drop condition associated with the first store instruction operation, the operand management circuitry is configured to remove the first store instruction operation from the fusion buffer without forwarding the first store instruction operation for execution. The operand management circuitry is configured to forward the first stored instruction operation for execution by the execution circuitry in response to detecting a buffer emptied condition and not detecting a discard condition.
Owner:APPLE INC

Performance analysis method and device for source address verification, equipment and storage medium

The invention provides a performance analysis method and device for source address verification, equipment and a storage medium, and relates to the technical field of network security. The method comprises the following steps: receiving a source address verification processing program written through a description language; loading an algorithm configuration file; wherein the algorithm configuration file at least comprises a performance model configured for a plurality of preset algorithms, and the performance model is used for calculating the consumed time length of the operation statement based on the operand length of each operation statement; based on the source address verification processing program and the performance model, calculating the consumed time length of each operation statement; and generating a performance analysis report of the source address verification processing program according to the time-consuming duration of all the operation statements in the source address verification processing program. According to the embodiment of the invention, automatic and theoretical performance analysis can be carried out on the source address verification processing flow based on the formalized description language and the predefined algorithm configuration file under the condition of separating from the dependence of specific hardware.
Owner:TSINGHUA UNIVERSITY

Systems, methods, and apparatuses for tile transpose

Embodiments detailed herein relate to matrix operations. In particular, support for a matrix transpose instruction is detailed. In some embodiments, decode circuitry to decode an instruction having fields for an opcode, a source matrix operand identifier, and a destination matrix operand identifier; and execution circuitry to execute the decoded instruction to transpose each row of elements of the identified source matrix operand into a corresponding column of the identified destination matrix operand are detailed.
Owner:INTEL CORP

A coprocessor-based method and system for taint propagation

The application discloses a kind of based on coprocessor's taint propagation method and system.The steps of the method include:1) pre-analysis instruction semantics, construct the taint propagation mask table of instruction;2) by coprocessor, intervene CPU instruction pipeline decoding, write-back process;3) in decoding stage, the opcode and operand of current instruction are acquired in real time quickly, and the taint propagation mask corresponding to instruction is found in taint propagation mask table by coprocessor;4) in write-back stage, the mask table corresponding to current instruction is searched in real time quickly, and taint propagation calculation is implemented.The application can analyze the semantics of instruction, calculate taint propagation process by using the characteristics of hardware fast and accurate when CPU instruction is executed by configuring the input function return value of monitoring.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Reduced size arithmetic unit for multiplication and division

An arithmetic unit implemented as an integrated circuit includes control logic configured to receive a control input representing a multiplication operation or a division operation and configure significand logic to perform a selected one of the multiplication operation and the division operation. The significand logic is configured to receive a first operand and a second operand and perform the selected operation on at least a portion of the first operand and at least a portion of the second operand.
Owner:NORTHROP GRUMMAN SYSTEMS CORP

A maximum value search device, a dot product operation device, and a tensor operation system

ActiveCN119473214BOperandLeading zero
The application provides a maximum value searching device for searching a maximum index in a plurality of integer source operands, the device comprising a first maximum value searching unit, wherein the first maximum value searching unit comprises: a first one-hot encoder for obtaining a plurality of source operands and generating a one-hot code corresponding to each source operand; a first bit-by-bit OR module for performing a bit-by-bit OR operation on the one-hot codes of the plurality of source operands and obtaining an intermediate result, wherein the bit-by-bit OR module is consistent with the code bit width of the one-hot encoder; a first leading zero searching module for searching the intermediate result to obtain a leading zero searching result, wherein the format of the leading zero searching result is a binary data format and the data bit width is consistent with the bit width corresponding to the plurality of source operands; and a first bit-by-bit NOT module for performing a bit-by-bit NOT operation on the leading zero searching result to obtain a maximum value in the plurality of source operands. The application can better meet the timing and area requirements.
Owner:T-HEAD (SHANGHAI) SEMICON CO LTD +1

Binary program vulnerability detection method and related device

The invention belongs to the technical field of power system security protection, and discloses a binary program vulnerability detection method and related device.The binary program vulnerability detection method comprises the steps that a VEX instruction of a to-be-detected binary program is obtained, and a program dependency graph is generated according to the VEX instruction; disassembling the to-be-detected binary program to obtain an assembly instruction, traversing the program dependency graph based on a preset sensitive operation instruction, and performing operand granularity slicing of the assembly instruction to obtain assembly instruction slices; a pre-trained encoder model is adopted to obtain feature vectors of the assembly instruction slices; and according to the feature vector of the assembly instruction slice, based on a pre-trained vulnerability detection model, obtaining a vulnerability detection result of the assembly instruction slice. According to the method, accurate dependency analysis is carried out around operands, a large number of irrelevant instruction interferences in binary program codes are effectively filtered out, most sensitive operations generating vulnerabilities can be covered based on preset sensitive operation instructions, the noise influence of redundant information on a model is reduced, and the robustness and generalization ability of detection are improved.
Owner:CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +1

Apparatuses, methods, and systems for instructions for structured-sparse tile matrix fma

Systems, methods, and apparatuses relating sparsity based FMA. In some examples, an instance of a single FMA instruction has one or more fields for an opcode, one or more fields to identify a source / destination matrix operand, one or more fields to identify a first plurality of source matrix operands, one or more fields to identify a second plurality of matrix operands, wherein the opcode is to indicate that execution circuitry is to select a proper subset of FP8 data elements from the first plurality of source matrix operands based on sparsity controls from a first matrix operand of the second plurality of matrix operands and perform a FMA.
Owner:INTEL CORP

Branch prediction based on sampled values

An apparatus comprises sampled state storage to store sampled register values of a register operand sampled at a sampling point in program flow, and prediction circuitry. In response to a semantic branch trigger indicating that there is a future branch instruction at a point in program flow later than the sampling point which satisfies a semantic branch condition, the prediction circuitry is configured to make a determination of a branch outcome of the future branch instruction based on a sampled register value of a particular register operand. The semantic branch condition is satisfied by a given future branch instruction for which a branch outcome is dependent on a sampled register value stored in the sampled state storage, and a value that the given register operand will have when program flow reaches the given future branch instruction can be calculated deterministically based on the sampled register value of the given register operand.
Owner:ARM LTD

In-memory bitwise operation circuit and method thereof

An in-memory bitwise operation circuit and method thereof are provided. The circuit includes a memory array, storing page data each having strings; a page buffer, storing the strings and selecting a portion of the strings as operands; a pop-count counter, receiving the operands and an operator flag, counting the number of bit 1 at corresponding bit positions of the operands, and generating a bitwise operation result corresponding to the operator; a bitwise operation processing unit, receiving and temporally storing the bitwise operation result, and outputting a final bitwise operation result when receiving a final result flag. In case that the operator is a last operator, the bitwise operation result is output as the final bitwise operation result. The circuit and method are suitable for 3D NAND flash memory that has high capacity and high performance.
Owner:MACRONIX INTERNATIONAL CO LTD

An optimization method for fusing floating-point multiplication-addition algorithm in single-precision floating-point multiplier-adder

PendingCN122308781ASign bitComputation process
This invention provides an optimized method for integrating floating-point multiply-accumulate algorithms in a single-precision floating-point multiply-accumulate unit. During the single-precision floating-point multiply-accumulate calculation, the pipeline is implemented as follows: P0 stage: The first pipeline stage receives three operands a, b, and c as input; it performs special value judgment, product mantissa calculation, product leading zero pre-statistics, and exponent difference and comparison logic; P1 stage: The second pipeline stage receives the multiplication result, compares the product exponent with the exponent of operand c, compares the exponent difference with the product leading zero, and performs exponent difference and product leading zero comparison; after exponent alignment, the mantissa of the aligned result is summed; P2 stage: The third pipeline stage receives the mantissa summed result, adjusts the exponent data, the sign bit to be operated on, and compares the results; after completion of normalization and rounding operations, the final result is adjusted accordingly, and abnormal states are judged. A single data path is used to ensure reuse of the same logic module within the pipeline, reducing area overhead.
Owner:HEFEI JUNZHENG TECH CO LTD

Granular Source Read Scheduling for Instruction Execution

PendingUS20260079715A1Resource allocationConcurrent instruction executionComputer architectureSingle instruction, multiple threads
Techniques are disclosed relating to accessing source data in single-instruction multiple-thread (SIMT) pipelines. In some embodiments, multiple categories of operand resource circuits are configured to provide operands for instructions executed by processor pipeline circuitry. Per-resource arbitration circuitry may arbitrate between the SIMT execution slots for access to different operand resources. Source access circuitry may access operand data from operand resources based on source capture commands and source control circuitry may prior to a first SIMT group winning arbitration for all its operands, send a source capture command to the source access circuitry in response to the first SIMT group winning arbitration at the per-resource arbitration circuitry for a first operand resource. Instruction control circuitry may send an instruction release command down the processor pipeline circuitry for the first SIMT group, in response to the first SIMT group winning arbitration for all its operands.
Owner:APPLE INC