Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

18 results about "Vector operations" patented technology

Vector operations, Extension of the laws of elementary algebra to vector s. They include addition, subtraction, and three types of multiplication. The sum of two vectors is a third vector, represented as the diagonal of the parallelogram constructed with the two original vectors as sides.

Neural processing unit for performing RMS norm operation and control method thereof

A neural processing unit for performing inference operations of a large-scale language model based on an artificial neural network is disclosed. The neural processing unit according to the present disclosure includes a processing element core configured to perform an attention mechanism-based operation based on input data in vector format to output an operation result, a special function unit comprising a plurality of arithmetic circuits including at least one vector-dedicated arithmetic circuit that exclusively performs vector operations and at least one mixed arithmetic circuit capable of performing both vector and scalar operations, and configured to perform a special function operation on the operation result, and a controller configured to, upon receiving an RMS normalization operation execution command, activate at least one of the plurality of arithmetic circuits to control the special function unit to perform an operation of converting at least one of the operation result or the input data into a normalized vector whose magnitude is adjusted based on a root mean square (RMS), wherein the operation result may include an attention score for the input data.
Owner:DEEPX CO LTD

Alignment in hardware accelerators

Systems, apparatuses, and methods are disclosed for improved matrix–vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values in floating point formats. The CIM macro has a functional block configured to align mantissa bits of primitive products between the activation values and the weight values by shifting the mantissa bits and an adder tree configured to output an accumulation value of the primitive products in an integer format by adding the shifted mantissa bits.
Owner:OPENAI OPCO LLC

Alignment in hardware accelerators

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values in floating point formats. The CIM macro has a functional block configured to align mantissa bits of primitive products between the activation values and the weight values by shifting the mantissa bits and an adder tree configured to output an accumulation value of the primitive products in an integer format by adding the shifted mantissa bits.
Owner:OPENAI OPCO LLC

Block quantization techniques for processing-in-memory devices

PCT designated stageWO2026135899A1Code conversionCoding detailsComputational scienceBinary multiplier
A processing-in-memory (PIM) device implements block quantization techniques for matrix-vector operations. The PIM device performs matrix-vector operations between portions of a weight matrix and an input vector, and copies results to a register. A read operation retrieves the copied results while additional matrix-vector operations are performed in parallel. The device may apply scaling factors to the results using multipliers within the PIM device. In some implementations, the weight matrix includes data columns and scaling factor columns interspersed at regular intervals. The scaling factors may be applied to accumulated results using parallel multiplication operations. Disclosed techniques enable efficient implementation of block quantization for applications such as Large Language Models while managing computational resources within the PIM architecture.
Owner:QUALCOMM INC

Method and apparatus for performing deep learning operations

ActiveCN114595811BAccumulator (computing)Binary multiplier
Methods and apparatuses for performing deep learning operations are provided. The computing apparatus includes an adder-tree-based tensor kernel and a multiplier-accumulator (MAC)-based vector kernel. The adder-tree-based tensor kernel is configured to perform tensor operations, and the multiplier-accumulator (MAC)-based vector kernel is configured to perform vector operations using the output of the tensor kernel as input.
Owner:SAMSUNG ELECTRONICS CO LTD

Vector processing dual-stream prefetch hardware system and vector operation method

This invention provides a dual-stream prefetch hardware system and vector operation method for vector processing. The system includes a computational path for performing regular vector operations in response to vector operation instructions, a computational path for performing streaming vector operations in response to predefined streaming operation instructions, and a storage unit. The system includes a streaming instruction control unit and a vector streaming processing core. The streaming instruction control unit is used to identify predefined streaming operation instructions and is configured to close the computational path for regular vector operations when a predefined streaming operation instruction is detected to be fetched. The vector streaming processing core is used to perform streaming vector operations in response to predefined streaming operation instructions. This solution achieves hardware decoupling of dense layout operations and non-dense sparse streaming operations by building dual independent computational paths for regular vector operations and streaming vector operations, thereby effectively improving the overall operation efficiency and real-time performance.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

Impedance line automatic adjustment method, device and medium

This invention discloses an automatic impedance line adjustment method, device, and medium, relating to the field of impedance line design technology. The method includes: identifying whether two input target lines are impedance lines; if so, acquiring their spatial constraint data on the circuit board plane; determining the relative positional relationship of the two impedance lines through vector operations to distinguish the upper and lower lines; extracting the outer spatial constraints of both lines from the spatial constraint data based on this relationship, and determining the available outer space; calculating the current actual spacing between the impedance lines, and calculating the target spacing and movement compensation amount based on impedance model parameters; determining the movement direction and distance of the upper and / or lower lines based on the available outer space and movement compensation amount, and controlling their movement; after the movement is completed, performing operations to merge adjacent line segments and delete redundant line segments on the line graphic to optimize the layout. This invention achieves efficient, accurate, and adaptive adjustment of impedance line spacing, effectively avoiding spatial interference and adapting to constraints in multiple scenarios.
Owner:SHENZHEN PARTNER INFORMATION TECH

A non-contact measuring device for detecting the alignment of shafts in large turbines

This invention belongs to the field of mechanical measurement and automated inspection technology, and discloses a non-contact measuring device for the alignment detection of large turbine shafts. The device includes a diameter measuring probe, a measuring rod unit, and a processor. The diameter measuring probe captures the position of a reference laser on its incident and exit surfaces. The measuring rod unit is installed at the bottom of the diameter measuring probe and contacts the inner wall of the turbine. The processor is electrically connected to the diameter measuring probe and is installed inside the probe. The processor processes, analyzes, and calculates the data measured by the diameter measuring probe to obtain the measurement result. This invention employs a non-contact optical measurement method, avoiding the contact error and wire bending deformation problems of the traditional wire drawing method. Combined with the precise calculation logic of vector operations, it significantly improves the accuracy of turbine shaft alignment detection.
Owner:CHINA ENERGY ENG GRP TIANJIN ELECTRIC POWER CONSTR CO LTD

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix–vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

GEMV operation method, system, device and medium of near-memory computing architecture

The application provides a GEMV operation method, system, device and medium of a near-memory computing architecture, the method is applied to an AI accelerator architecture integrated with M independent GEMV accelerators, each GEMV accelerator internally contains N parallel computing channels, and each GEMV accelerator is separately connected with a dedicated 3D DRAM storage module, and each computing channel directly interacts with the 3D DRAM storage module corresponding to the GEMV accelerator. The near-memory computing AI accelerator architecture provided by the application closely couples the GEMV accelerator with the 3D DRAM storage module, precisely matches the characteristics of the GEMV operation in the output vector length without data dependency with the high-bandwidth hardware advantage of the 3D DRAM, so that each accelerator can independently read data from the local 3D DRAM, completely avoids the data transfer overhead across the storage blocks, significantly shortens the data transmission path, and realizes the dual optimization of computing efficiency and storage bandwidth utilization. The architecture is particularly suitable for AI model inference and other scenarios requiring high-density matrix vector operation.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Path planning

To determine a path through a pose configuration space, trajectories of poses may be evaluated in parallel based at least on translating the trajectories along at least one axis of the pose configuration space (e.g., an orientation axis). A trajectory may include at least a portion of a turn having a fixed turn radius. Turns or turn portions that have the same turn radius and initial orientation can be translatively shifted along and processed in parallel along the orientation axis as they are translated copies of each other, but with different starting points. Trajectories may be evaluated based at least on processing variables used to evaluate reachability as bit vectors with threads effectively performing large vector operations in synchronization. A parallel reduction pattern may be used to account for dependencies that may exist between sections of a trajectory for evaluating reachability, allowing for the sections to be processed in parallel.
Owner:NVIDIA CORP

Inference result acquisition method, electronic device, and computer-readable storage medium

The embodiments of the present application provide an inference result acquisition method, an electronic device, and a computer-readable storage medium. The inference result acquisition method comprises: after acquiring query text, an electronic device converting the query text into a token sequence, then acquiring, according to the token sequence, a first query matrix, a first key matrix, and a first value matrix in a prefill stage, and performing in parallel a dot product operation, a softmax normalization operation, and a weighted summation operation on the first query matrix, the first key matrix, and the first value matrix; next, in a decode stage, performing in parallel a matrix-vector operation, a softmax normalization operation, and a weighted summation operation on a current result vector, a second key matrix, and a second value matrix; and finally, obtaining a next token of an inference result according to a result of the weighted summation operation in the decode stage. The method uses parallel computing as a means to reduce computational latency and mitigate computational bottlenecks in prefill stages, and to increase bandwidth utilization and mitigate memory bottlenecks in decode stages.
Owner:HUAWEI TECH CO LTD

Processing unit, computing device and instruction processing method

ActiveCN114924793BConditional code generationRegister arrangementsComputer architectureInstruction unit
Embodiments of the present application provide a processing unit, a computing device and an instruction processing method. The present solution is applicable to various chips including CISC instruction set, RISC reduced instruction set (especially RISC-V instruction set) or VLIM instruction set architecture, such as Internet of Things chips, audio / video chips and the like. The processing unit comprises: an instruction fetching unit configured to perform instruction fusion on sequentially adjacent vector configuration instructions and vector operation instructions to obtain fused instructions; an instruction decoding unit configured to decode the fused instructions to obtain first execution information and second execution information; a vector configuration unit configured to execute the vector configuration instructions according to the first execution information, modify a vector control register, and bypass the value of the modified vector control register to a vector operation unit; and the vector operation unit configured to execute the vector operation instructions according to the second execution information and the value of the modified vector control register. The present solution can improve the efficiency of vector operation of the processing unit.
Owner:C SKY MICROSYST CO LTD

Method, device, equipment and medium for realizing causality mask in large model inference

The application discloses a method and device for realizing causality mask in large model inference, equipment and medium, relates to the technical field of self-attention mechanism, and the method comprises the following steps: after a vector-matrix operation is completed by VMPE, the operation result is sent to SFU; the input data is acquired by SFU receiving the cmd command sent by FW; the acquired input data is processed based on preset bit operation mask operation, and then the Softmax operation is performed; wherein, the hardware implementation architecture comprises VMPE responsible for processing vector-matrix operation, SFU responsible for processing other vector operation, and FW running on RISC-V CPU and used for controlling the work of DMA, VMPE and SFU. According to the application, the mask matrix does not need to be constructed in advance, so that the storage space is effectively saved.
Owner:SIENGINE TECH CO LTD

Sparse data processing method and device of neural network processor

The present disclosure provides a sparse data processing method and device of a neural network processor, the neural network processor comprising a basic computing unit, the method comprising: obtaining a plurality of groups of weight sub-vectors, wherein the weight sub-vectors are obtained by sparse processing of a to-be-computed weight vector based on an information unit supported by the basic computing unit; determining a to-be-computed feature vector corresponding to the to-be-computed weight vector; controlling the basic computing unit to perform vector inner product operation on each group of weight sub-vectors and the to-be-computed feature vector to obtain vector operation results; and performing shift operation on part of the group vector operation results, adding the vector operation results obtained by the shift operation to the remaining group vector operation results, and taking the addition result as a sparse data processing result, wherein the part of the group vector operation results and the remaining group vector operation results together constitute the plurality of group vector operation results. Through the present disclosure, the distribution of weights can be fully utilized, the data processing accuracy of the neural network processor can be effectively guaranteed, the hardware unit cost and the execution power consumption can be effectively reduced, and the weight storage space can be effectively reduced.
Owner:AXERA TECH (BEIJING) CO LTD