Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

78 results about "Vector operations" patented technology

Vector operations, Extension of the laws of elementary algebra to vector s. They include addition, subtraction, and three types of multiplication. The sum of two vectors is a third vector, represented as the diagonal of the parallelogram constructed with the two original vectors as sides.

High-speed NTT method and system based on GPU matrix-thread collaborative optimization

The invention provides a high-speed NTT method and system based on GPU matrix-thread collaborative optimization. The method comprises the steps that an input twiddle factor is converted into a matrix form; transmitting the updated twiddle factor matrix and the input array to a graphics processor; according to the input array and the size of the twiddle factor matrix, the number of blocks in the grid and the number of threads included in each block are set, so that each block is responsible for the operation of multiplying one row by one column, and each thread in the block completes the operation of multiplying one pair of data in the corresponding row-by-column process to obtain a group of intermediate results; and obtaining output data corresponding to the block according to an intermediate result in the block, and sending the output data to a corresponding position of an output matrix to serve as a multiplication result. According to the method, NTT calculation is reconstructed into parallel matrix-vector operation, the maximum utilization of GPU calculation resources is realized through the combination of an efficient thread allocation strategy and an optimized memory access mode, and the calculation efficiency is remarkably improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Data processing method and device, electronic equipment and storage medium

Embodiments of the invention provide a data processing method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a unit vector operation width supported by hardware; obtaining a source address to be aligned and the number of object data to be operated; in response to the fact that the width of the object data is larger than or equal to a first threshold value, the remainder obtained after the to-be-aligned source address is aligned with the unit vector operation width is calculated so as to obtain the data width of which the addresses are not aligned in the object data, and the width of the object data is equal to the product of the unit data width of the object data and the number of the object data; the width of the address misalignment data is equal to the difference value between the unit vector operation width and the remainder; determining the number of the address misalignment data according to the unit data width of the object data and the width of the address misalignment data, and determining first mask information based on the number of the address misalignment data and the unit vector operation width; and according to the first mask information, loading and processing the data of which the addresses are not aligned, and the method improves the data processing performance.
Owner:HYGON INFORMATION TECH CO LTD

Data processing apparatus, processor, board card, and data processing method

A data processing apparatus, a processor, a board card, and a data processing method. The data processing apparatus may be comprised by a combined processing apparatus (20), the combined processing apparatus (20) comprising a computing apparatus (201), an interface apparatus (202), a processing apparatus (203) and a storage apparatus (204). The computing apparatus (201) is configured to execute an operation specified by a user, so as to execute deep learning or machine learning calculation. The computing apparatus (201) may interact with the processing apparatus (203) by means of the interface apparatus (202), so as to jointly complete the operation specified by the user. The interface apparatus (202) is used to transmit data and a control instruction between the computing apparatus (201) and the processing apparatus (203). The storage apparatus (204) is used to store data of the computing apparatus (201) and the processing apparatus (203). The data processing apparatus provides a vector computation scheme integrating different types of fine-grained quantization, so that processing can be simplified, and the advantages of a low bit width operation and small memory occupation of a fine-grained quantization format are fully utilized.
Owner:SHANGHAI CAMBRICON INFORMATION TECH CO LTD

Compressing and transforming vector operations in an ai model

Techniques are described herein that are capable of compressing and transforming vector operations in an AI model. First output multi-bit elements (MBEs) are generated by combining input single-bit components (SBCs) representing an input token in an AI prompt and first SBCs representing a first layer of the AI model using an exclusive-or operation. The first output MBEs are transformed into first output single-bit elements (SBEs) using a random probability distribution. Second output MBEs are generated by combining intermediate SBEs corresponding to intermediate MBEs derived from the first output SBEs and second SBCs representing a second layer of the AI model using the exclusive-or operation. A response to the AI prompt is generated to include an output token corresponding to a combination of a norm of the intermediate MBEs, a norm of second multi-bit components from which the second SBCs are derived, and a representation of the second output MBEs.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Implementing vector index build and search using query operators

Methods, systems, and computer program products are provided that implement vector-related requests using query operators. For example, a system includes a database, a parser, a converter, an optimizer, and an execution engine. The database is associated with query operators that are non-vector-specific. The database stores a table with vector embeddings. The parser is configured to parse a request indicating a vector operation associated with the vector embeddings. The converter is configured to convert the vector operation into a logical operator tree comprising a representation of the request as a logical flow of the query operators, enabling the vector operation without vector-specific executable code or operators. The optimizer is configured to convert the logical operator tree into an executable plan. The execution engine is configured to execute the executable plan against the table with vector embeddings.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Processor, method, device and storage medium for data processing

A processor and a method, device and storage medium for data processing are provided. The processor includes an instruction decoder configured to decode a target instruction for a vector operation. The target instruction involves a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location in the memory for reading to-be-processed data. The target operand specifies a target storage location in the memory for writing a processed result. The processor further includes an arithmetic logic unit configured to: read the to-be-processed data from the source storage location of the memory; perform, on the to-be-processed data, an arithmetic logic operation associated with the vector operation specified by the target instruction; and write the processed result to the target storage location of the memory.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Instruction processing method and device, chip product, computer equipment and storage medium

The invention provides an instruction processing method and device, a chip product, computer equipment and a storage medium. The method comprises the following steps: in response to a received to-be-processed vector instruction, splitting the vector instruction into a plurality of microoperations respectively corresponding to a plurality of destination registers; the micro-operation is used for acquiring data access information required by the destination register corresponding to the micro-operation; executing each item of microoperation in parallel; triggering an execution unit to execute a calculation operation under the condition of determining that all the micro-operations are executed; wherein the execution unit, in response to the calculation operation, performs a vector operation on the plurality of destination registers based on the data access information acquired by all the micro-operations. The number of micro-operations can be greatly reduced, and the problem of pipeline blockage caused by the fact that a transmitting queue is filled is avoided.
Owner:SOPHGO TECH LTD

Method for performing FFT, processor, and computing device

Embodiments disclosed in this application pertain to the field of computer technologies, and in particular, to a method for performing FFT, a processor, and a computing device. The method includes: A processor responds to an execution request of fast Fourier transform FFT calculation of an application, and decomposes the FFT calculation into a plurality of calculation stages. The processor sequentially executes the plurality of calculation stages, where when a target calculation stage is executed, a vector operation circuit performs rotation factor calculation, and a matrix operation circuit performs DFT calculation. After the execution of the plurality of calculation stages is completed, the processor determines an execution result of the FFT calculation based on an execution result of a last calculation stage, and returns the execution result to the application. vector operation circuit matrix operation circuit
Owner:HUAWEI TECH CO LTD +1

A method of synthesizing a rotating phasor

The application relates to a synthesis method of a rotating phasor, which is suitable for the technical field of electrical rotating phasor conversion and is particularly directed to the decomposition and synthesis of an electrical rotating phasor. The method adjusts the initial phase of two amplitude-constant sub-rotating phasors, and then synthesizes a target rotating phasor with adjustable amplitude and initial phase, and has the characteristics of easy generation of the sub-rotating phasor, good flexibility and high reliability. The application mainly utilizes the parallelogram rule of vector operation, reduces the generation difficulty of the sub-rotating phasor by reducing the generation problem of the target rotating phasor with adjustable amplitude and initial phase to the synthesis problem of the two sub-rotating phasors with constant amplitude and adjustable phase, and solves the problem of synthesizing a complex rotating phasor from a simple rotating phasor. The technical scheme provided by the application forms a simple sub-rotating phasor through electromagnetic conversion, and provides a new idea for the synthesis of a complex rotating phasor.
Owner:NORTH CHINA ELECTRIC POWER UNIV +1

Neural processing unit for performing RMS norm operation and control method thereof

A neural processing unit for performing inference operations of a large-scale language model based on an artificial neural network is disclosed. The neural processing unit according to the present disclosure includes a processing element core configured to perform an attention mechanism-based operation based on input data in vector format to output an operation result, a special function unit comprising a plurality of arithmetic circuits including at least one vector-dedicated arithmetic circuit that exclusively performs vector operations and at least one mixed arithmetic circuit capable of performing both vector and scalar operations, and configured to perform a special function operation on the operation result, and a controller configured to, upon receiving an RMS normalization operation execution command, activate at least one of the plurality of arithmetic circuits to control the special function unit to perform an operation of converting at least one of the operation result or the input data into a normalized vector whose magnitude is adjusted based on a root mean square (RMS), wherein the operation result may include an attention score for the input data.
Owner:DEEPX CO LTD

Alignment in hardware accelerators

Systems, apparatuses, and methods are disclosed for improved matrix–vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values in floating point formats. The CIM macro has a functional block configured to align mantissa bits of primitive products between the activation values and the weight values by shifting the mantissa bits and an adder tree configured to output an accumulation value of the primitive products in an integer format by adding the shifted mantissa bits.
Owner:OPENAI OPCO LLC

Method, system and equipment for calculating water quantity and water rate of water meter based on vector calculation and medium

The invention discloses a vector calculation-based water meter water quantity and water rate calculation method, and aims to solve the technical problems of I / O bottleneck, high calculation delay and poor expansibility in existing large-scale water meter charging. The method comprises the following steps: constructing a water meter data column vector, calculating an original difference vector based on a current-period reading vector and a previous-period reading vector, and performing Boolean mask processing to obtain a meter reading water volume vector; actual water consumption accounting, total and score table allocation and allocated water amount calculation are achieved through vector operation; a billing water volume vector is obtained in combination with a preferential strategy, and an actual water fee is finally obtained through mixed water splitting, stepped billing matrix operation and water fee preferential adjustment. According to the method, vectorization and matrix operation are adopted to replace traditional line-by-line calculation, so that the operation efficiency can be improved through parallel calculation, the vector algorithm is designed to adapt to complex water affair billing logic, the processing performance is effectively improved, efficient and large-scale billing is achieved, the intelligent water affair development requirement is met, and practicability and popularization value are achieved.
Owner:CHONGQING SENXINJU INTELLIGENT TECH CO LTD

Cache mode dynamic switching-based AI calculation acceleration method and system

The invention discloses an AI calculation acceleration method and system based on cache mode dynamic switching, and the method comprises the steps: determining a to-be-switched target cache block in a last-stage cache after receiving an AI calculation acceleration request; based on the physical address, the data is quickly refreshed to guarantee the consistency; marking the AI calculation mode as an AI calculation mode and shielding the AI calculation mode in an allocation strategy; executing in-memory calculation by utilizing a storage array and a vector operation unit which are integrated; and after the task is completed, clearing the mark and recovering the available state. The system comprises a cache switching controller, a matrix multiplication scheduler, a data preprocessing module, a cache isolation module and an instruction interface and software scheduling unit. By multiplexing the last level of cache and dynamically switching the working mode of the last level of cache, the AI calculation is efficiently completed while the system stability is ensured, and the computing power per unit area of a chip is remarkably improved.
Owner:SHANGHAI JIAOTONG UNIV

Method and apparatus for controlling input / output operation of vector processor in mixed precision environment

A The present invention relates to a technique for controlling input / output operation of a vector processor, which is designed to optimize vector operation in a mixed precision environment, and to a technique for maximizing data processing performance while minimizing waste of operation resources by improving the data conversion process between the memory and the vector processor.
Owner:SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION

Alignment in hardware accelerators

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values in floating point formats. The CIM macro has a functional block configured to align mantissa bits of primitive products between the activation values and the weight values by shifting the mantissa bits and an adder tree configured to output an accumulation value of the primitive products in an integer format by adding the shifted mantissa bits.
Owner:OPENAI OPCO LLC

Block quantization techniques for processing-in-memory devices

A processing-in-memory (PIM) device implements block quantization techniques for matrix-vector operations. The PIM device performs matrix-vector operations between portions of a weight matrix and an input vector, and copies results to a register. A read operation retrieves the copied results while additional matrix-vector operations are performed in parallel. The device may apply scaling factors to the results using multipliers within the PIM device. In some implementations, the weight matrix includes data columns and scaling factor columns interspersed at regular intervals. The scaling factors may be applied to accumulated results using parallel multiplication operations. Disclosed techniques enable efficient implementation of block quantization for applications such as Large Language Models while managing computational resources within the PIM architecture.
Owner:QUALCOMM INC

Hardware-friendly Transform column balance pruning model compression and efficient deployment method

The invention discloses a hardware-friendly Transform column balance pruning model compression and efficient deployment method. A model compression algorithm, a lightweight parameter storage format, an operation data buffer, a systolic array operation block, a vector operation unit, a nonlinear operator unit, a data flow controller and a DMA unit are included. A model compression algorithm and an efficient deployment architecture are explored according to Transform network software and hardware collaborative reasoning requirements: in a software level, the scale calculation complexity of model parameter quantities is reduced through a fine-grained column balance structured pruning strategy, and parameters are stored in a single-instruction multi-data-stream format and the parameter storage efficiency is optimized through mask code storage; according to the hardware level, an edge computing-oriented Transform special accelerator architecture is designed, so that the architecture can support column balance structured pruning characteristics and a lightweight parameter storage scheme in an original manner. According to the Transform model compression and efficient deployment method, the parameter sparsity after structured pruning is fully utilized, so that the parameter storage pressure of a hardware architecture is reduced, the complex balance of an arithmetic unit is ensured, the operation efficiency of an accelerator is improved, and load balance and efficient reasoning during software and hardware collaborative optimization are realized; the method is widely applicable to efficient deployment scenes of Transform models for edge calculation.
Owner:BEIJING UNIV OF TECH

Method and apparatus for performing deep learning operations

ActiveCN114595811BAccumulator (computing)Binary multiplier
Methods and apparatuses for performing deep learning operations are provided. The computing apparatus includes an adder-tree-based tensor kernel and a multiplier-accumulator (MAC)-based vector kernel. The adder-tree-based tensor kernel is configured to perform tensor operations, and the multiplier-accumulator (MAC)-based vector kernel is configured to perform vector operations using the output of the tensor kernel as input.
Owner:SAMSUNG ELECTRONICS CO LTD

Three-dimensional discrete element flexible undrained loading method based on dynamic geometry

The invention belongs to the technical field of three-dimensional discrete element loading, and aims to solve the technical problems that a hexagonal unit is mostly adopted in the current discrete element three-axis undrained loading and volume monitoring process, adjacent particle pointers need to be repeatedly read through id numbers, a large amount of vector operation is carried out, and the computing power is wasted. The invention provides a three-dimensional discrete element flexible undrained loading method based on dynamic geometry, and the method comprises the steps: simplifying a hexagonal basic unit in a conventional flexible film loading process into a simpler triangular unit through introducing a geometry unit; and in combination with the specific area and normal vector extraction advantages of the geometer poly unit, the load application and volume dynamic monitoring process in the non-drainage loading process of the flexible film is effectively simplified, and the related research of the discrete element three-dimensional flexible film is promoted.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY +2

Software simulation method for vector multiplication

The invention relates to a software simulation method for vector multiplication. The core of the method is to select a bounded or non-bounded instruction to efficiently read data according to the alignment condition of a data access address, and intelligently select different simulation schemes according to the condition that a multiplier is a constant or a variable: for the constant, the constant is decomposed into a power combination of 2, and multiplication is simulated through a shift and addition instruction; for variables, an iteration process is adopted, and multiplication is achieved through least significant bit group extraction, number head zero calculation, shifting and conditional accumulation operation; and finally, selecting a storage instruction according to a destination address alignment condition. According to the method, an existing SIMD instruction set is fully utilized, the hardware vector multiplication function is efficiently simulated in a software mode, the word integer vector operation performance can reach eight times of that of standard quantity operation, and the performance and competitiveness of a domestic processor in the field of data processing are remarkably improved.
Owner:CLP KESHENTAI INFORMATION TECH CO LTD

Vector processing dual-stream prefetch hardware system and vector operation method

This invention provides a dual-stream prefetch hardware system and vector operation method for vector processing. The system includes a computational path for performing regular vector operations in response to vector operation instructions, a computational path for performing streaming vector operations in response to predefined streaming operation instructions, and a storage unit. The system includes a streaming instruction control unit and a vector streaming processing core. The streaming instruction control unit is used to identify predefined streaming operation instructions and is configured to close the computational path for regular vector operations when a predefined streaming operation instruction is detected to be fetched. The vector streaming processing core is used to perform streaming vector operations in response to predefined streaming operation instructions. This solution achieves hardware decoupling of dense layout operations and non-dense sparse streaming operations by building dual independent computational paths for regular vector operations and streaming vector operations, thereby effectively improving the overall operation efficiency and real-time performance.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Method, device and equipment for evaluating quality of title answer and storage medium

The application relates to the computer field, in particular to the artificial intelligence field, and provides a question answering quality evaluation method and device, equipment and a storage medium. The method comprises the following steps: obtaining answering results input by target objects for a same subjective question; obtaining target similarity vectors of the answering results based on the intermediate similarity vectors of each answering result and the intermediate similarity vectors of the answering results, wherein each target similarity vector represents a first similarity degree distribution of the answering result and a second similarity degree distribution between the first similarity degree distributions of the answering results; determining answering quality evaluation features of the corresponding answering results based on the obtained target similarity vectors, and inputting the answering quality evaluation features into a target evaluation model to obtain answering quality evaluation results of the corresponding answering results. The feature extraction is performed in the form of vector operation, the operation amount is small, the delay is low, and the application scenarios are wide.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Search device and search method

A retrieval device and a retrieval method. The retrieval device (10) comprises a first processor (110) and a second processor (120), wherein the first processor (110) is a general central processor, used to perform table lookup operation in the retrieval process; the second processor (120) is a neural network processor, used to perform matrix / vector operation in the retrieval process. In the retrieval method, the steps of the retrieval algorithm (retrieval process) are reasonably allocated to different processors for execution, so as to exert the advantages of different kinds of processors and improve the retrieval efficiency.
Owner:HUAWEI TECH CO LTD

Vector processor-oriented layered bypass forwarding method and system

The invention discloses a hierarchical bypass forwarding method and system for a vector processor, belongs to the technical field of vector processors, and aims to solve the technical problem of how to overcome the defect of low efficiency caused by read-after-write (RAW) data risk in vector operation of the vector processor, effectively reduce pause caused by VRF access delay and improve the reliability of the vector processor. According to the technical scheme, in each vector processing channel of the vector processor, before a front instruction result is written back to a local vector register file of each vector processing channel, a two-stage data forwarding architecture comprising an in-channel bypass unit and a cross-channel bypass network is established by adopting a layered result forwarding network; a front instruction result can be directly transmitted to a subsequent instruction through a two-stage data forwarding architecture comprising an in-channel bypass unit and a cross-channel bypass network; wherein each vector processing channel is internally provided with an in-channel bypass unit.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Fused instruction operation for performing an operation and adjusting a sign of the operation result

Techniques are disclosed involving fusing instruction pairs and executing corresponding fused instruction operations. A processor includes fusion detection circuitry to detect a pair of fetched instructions and fuse the instructions into a fused instruction operation, and execution circuitry to execute the fused instruction operation. In one embodiment, a first instruction is executable to perform an operation and a second instruction is executable to adjust a sign of a result of the operation. In another embodiment, the first instruction is executable to perform an operation and the second instruction is executable to find a maximum or minimum, as compared to a comparison operand, of a result of the operation. In another embodiment, the first instruction is executable to perform a vector operation and the second instruction is executable to read a first element of the vector result and overwrite one or more additional elements of the vector result.
Owner:APPLE INC

Accelerated quantization and inverse quantization data processing method, system, product and terminal for neural network reasoning

The invention provides an accelerated quantization and inverse quantization data processing method and system for neural network reasoning, a product and a terminal, and the method comprises the steps: carrying out the grouping processing of weight data of a current layer, and adaptively determining a corresponding inverse quantization processing mode based on preset configuration; performing zero adjustment processing, product accumulation and scaling processing on the weight data of the current layer; meanwhile, through accumulation processing and re-quantization processing, target result data are obtained and serve as activation data of the next layer. According to the method and the device, efficient utilization of hardware resources is realized through fine control of a quantization and inverse quantization process, DOT operation unit multiplexing and application of an inverse quantization bypass mode, so that the technical problems of low inverse quantization efficiency and high power consumption caused by insufficient utilization rate of a vector operation unit and unreasonable execution time sequence in the existing inverse quantization operation are solved; and the reasoning precision and the calculation efficiency of the neural network are balanced, and the hardware utilization rate and the calculation efficiency are remarkably improved.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Floating point multiplications

Systems, apparatuses, and methods are disclosed for improved matrix-vector operations in accelerators that may be useful or heavy AI training and inference workloads. The disclosed technology provides arrangements that permit more efficient computation by, for example, performing calculations without repeated conversion between numeric domains. In some implementations, a compute-in-memory (CIM) macro is configured to perform a vector matrix multiplication (VMM) operation between a vector of activation values and a matrix of weight values, and a mode decoding unit is configured to provide a mode of the VMM operation according to a first floating point format of the activation values and a second floating point format of the weight values. A column cell of the CIM macro can be configured to output a product between an activation value and a weight value using a half adder and multiple multiplexers that provide selections to a full adder based on control signals.
Owner:OPENAI OPCO LLC

Large model three-valued zero inverse quantization inference method based on position mask coding

The invention discloses a large model three-valued zero inverse quantization reasoning method based on position mask coding, and belongs to the field of artificial intelligence. According to the method, positive and negative mask byte pairs are generated for each 8-element subgroup in a quantization stage aiming at three-valued representation {+ 1, 0,-1} of a large model weight, and memory access is optimized by adopting a mask cross storage layout. In the reasoning stage, conditional accumulation of activation values is guided by directly utilizing a position mask through a hardware bit mask instruction, and an inverse quantization link and element-by-element multiplication calculation in a traditional three-valued quantization model are thoroughly eliminated. According to the method, symbol information is converted into a bit-level control signal which can be directly analyzed by hardware, so that the three-valued weight can be more directly mapped to a vector operation unit of a general CPU, and the reasoning efficiency of a large model on resource-constrained equipment is greatly improved.
Owner:HUNAN UNIV

Techniques for efficient complex vector multiplication

Processing circuitry to perform a vector operation and instruction decoder circuitry to decode an instruction from an instruction set to control the processing circuitry to perform the vector operation specified by the instruction are provided. An array storage device having storage elements for storing data blocks is used to store at least one two-dimensional array of data blocks, the processing circuitry being accessible to the at least one two-dimensional array of data blocks when the vector operation is performed. The instruction set includes a complex numerical outer product instruction specifying a first source operand, a second source operand, and a destination operand, where each of the first source operand and the second source operand is a vector operand including a plurality of source data elements, each source data element being a complex number formed by a real part and an imaginary part, and the destination operand identifies a given two-dimensional array of data blocks within the array storage device. The processing circuitry, in response to the complex value outer product instruction, performs an outer product operation using the source data element of the first source operand and the source data element of the second source operand to produce a plurality of result data elements, where each result data element is a complex number formed by a real part and an imaginary part, and wherein each real part and each imaginary part of each result data element are associated with one of the data blocks in the given two-dimensional array of data blocks and are used to update the value of the associated data block.
Owner:ARM LTD