Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

15 results about "Rotation factor" patented technology

Method for performing FFT, processor, and computing device

Embodiments disclosed in this application pertain to the field of computer technologies, and in particular, to a method for performing FFT, a processor, and a computing device. The method includes: A processor responds to an execution request of fast Fourier transform FFT calculation of an application, and decomposes the FFT calculation into a plurality of calculation stages. The processor sequentially executes the plurality of calculation stages, where when a target calculation stage is executed, a vector operation circuit performs rotation factor calculation, and a matrix operation circuit performs DFT calculation. After the execution of the plurality of calculation stages is completed, the processor determines an execution result of the FFT calculation based on an execution result of a last calculation stage, and returns the execution result to the application. vector operation circuit matrix operation circuit
Owner:HUAWEI TECH CO LTD +1

Methods, processors, and computing devices that perform FFTs

The embodiment disclosed by the application belongs to the technical field of computing, and particularly relates to a method for performing FFT, a processor and a computing device. The method comprises the following steps: in response to an execution request of a fast Fourier transform (FFT) calculation of an application program, a processor decomposes the FFT calculation into multiple calculations; the processor sequentially executes multiple calculation stages, wherein, when a target calculation stage is executed, a rotation factor calculation is executed based on a vector operation unit, and a DFT calculation is executed based on a matrix operation unit; after the execution of the multiple calculation stages is completed, an execution result of the FFT calculation is determined based on an execution result of the last calculation stage, and the execution result is returned to the application program. According to the application, the processor can jointly implement the FFT calculation based on the vector operation unit and the matrix operation unit, and the efficiency of the processor in executing the FFT calculation can be improved.
Owner:HUAWEI TECH CO LTD +1

Data transformation method, apparatus, and storage medium

ActiveCN116186473Breduce complexityomit remainder operationRotation factorData transformation
The embodiment of the application discloses a data transformation method and device and a storage medium, and belongs to the field of data processing. In the embodiment of the application, the first transformation matrix is obtained by transforming an original transformation matrix according to the symmetry of a rotation factor in NTT. Thus, part of the matrix elements in the first transformation matrix has symmetry, so that the data processing results corresponding to the part of the matrix elements in the first transformation matrix are obtained, the data processing results corresponding to other matrix elements are restored, and then each output data is obtained. That is, in the embodiment of the application, the data processing results corresponding to part of the matrix elements are obtained, and all transformation information can be represented, the complexity of data transformation is reduced, and the transformation speed is accelerated.
Owner:HUAWEI TECH CO LTD

A GPU multi-thread parallel-based hawk algorithm acceleration method

ActiveCN122044804BComputational scienceRotation factor
The application provides a Hawk algorithm acceleration method based on GPU multi-thread parallelism. The Hawk algorithm acceleration method based on GPU multi-thread parallelism comprises the following steps: (1) thread resource division and task mapping; (2) kernel fusion and throughput peak detection; (3) NTT / iNTT parallelization reconstruction and memory access optimization; (4) FFT / iFFT structured parallelization optimization and butterfly operator fusion. The application realizes a coarse-grained parallel butterfly operator execution mode without complex address calculation, without shared memory synchronization, and without competition between threads, greatly improving the throughput efficiency on the GPU; meanwhile, the application unifies the parallel execution framework in the forward and inverse transformations of FFT / NTT, wherein the inverse transformation only needs to use the corresponding inverse rotation factor and perform simple normalization processing at the end to complete the overall recovery.
Owner:NANJING UNIV OF POSTS & TELECOMM

A method for reducing loss of large dynamic FFT operation by truncation and number system conversion

The application discloses a method for reducing the loss of large dynamic FFT operation by truncation and number system conversion, and relates to the field of radar signal processing. Firstly, by selecting appropriate rotation factor quantization bits and complex multiplication truncation bits, the full frequency domain precision loss is ensured to be not more than a fixed value under the condition of small signal input; then, whether to retain the bit width increased by complex addition operation is judged according to the order corresponding to the number of FFT points required to be operated; finally, according to the operation bit width truncation rule, after completing the entire FFT operation, the FFT operation result represented by fixed-point numbers is converted into the FFT operation result represented by floating-point numbers. The application can effectively retain the spectrum information under the condition of small signal input, avoid the signal distortion caused by the truncation in the FFT processing process, greatly improve the dynamic range, and provide good data support for subsequent fine processing.
Owner:THE 724TH RESEARCH INSTITUTE OF CHINA STATE SHIPBUILDING CORP LTD

Method and apparatus for constructing processing circuitry for target transform

ActiveCN119719591BComplex mathematical operationsRotation factorSequence transformation
Embodiments of the present specification provide a method for constructing a processing circuit for a target transform. The target transform is a discrete transform or its inverse transform that transforms an input coefficient sequence into an output coefficient sequence based on a rotation factor. The method includes, for a K-point input coefficient sequence to be processed, determining a plurality of candidate decomposition points corresponding to a k-bit coefficient index. Then, starting from a low-bit decomposition point, for each candidate decomposition point, determining the minimum storage cost under each candidate decomposition point according to a decomposition cost evaluation function through recursive index bit number decomposition for several levels, so as to determine a target decomposition mode of the k-bit coefficient index with the minimum storage cost. The decomposition cost evaluation function limits the storage cost to include the cost of a (n-p)th order first transform, the cost of a pth order second transform, and the local cost for inter-level rotation factor multiplication. According to the target decomposition mode, a corresponding memory for the rotation factor is allocated to form a processing circuit.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Rotation factor storage method, reading method and hardware accelerator for fft operation

The application relates to a rotation factor storage method, a reading method and a hardware accelerator for FFT operation. The hardware accelerator comprises a rotation factor storage unit for storing sine components or cosine components in rotation factor original data; an index calculation unit for mapping rotation factor original indexes into look-up table indexes of a rotation factor storage table and calculating sine component indexes and cosine component indexes; a folding judgment unit for judging whether the sine component indexes and the cosine component indexes exceed the index range of the rotation factor storage table; a data reading unit for respectively obtaining imaginary part readout values and real part readout values from the rotation factor storage table; and a sign correction unit for performing sign bit correction on the imaginary part readout values and the real part readout values to obtain real parts and imaginary parts of the rotation factors. The embodiment of the application reduces the storage resource occupation of the rotation factors, realizes the compatibility of a single rotation factor table to different point number FFTs, has high data reading precision and small hardware cost.
Owner:JINGYIN ELECTRONIC TECH (SHANGHAI) CO LTD

GPU multi-thread parallel Hawk algorithm acceleration method

The invention provides a Hawk algorithm acceleration method based on GPU (Graphics Processing Unit) multi-thread parallel. The GPU multi-thread parallel-based Hawk algorithm acceleration method comprises the following steps of (1) thread resource division and task mapping; (2) kernel fusion and throughput peak detection; (3) NTT / iNTT parallel reconstruction and memory access optimization are carried out; and (4) carrying out FFT / iFFT structured parallel optimization and butterfly operator fusion. According to the method, a coarse-grained parallel butterfly operator execution mode which does not need complex address calculation, does not need shared memory synchronization and does not compete among threads is realized, and the throughput efficiency on the GPU is greatly improved; meanwhile, a parallel execution framework is unified in forward and reverse transformation of the FFT / NTT, and overall recovery can be completed only by using a corresponding reverse twiddle factor and performing simple normalization processing at the tail end in the reverse transformation.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method and acceleration hardware for performing target transformation

The embodiment of the invention provides a method for executing target transformation and acceleration hardware. Target transformation comprises n rounds of transformation operation, and the method comprises the step of executing a non-first round of transformation operation by utilizing a qth circuit in acceleration hardware, which comprises the following steps: a bit exchange unit in the qth circuit exchanges two specified bits in each data position index of a (q-1) th round of output sequence to obtain a reordered first sequence; when the power distribution of the left twiddle factor in the q-th round is not zero, a left multiplication unit in the q-th circuit multiplies each data of the first sequence by each corresponding twiddle factor indicated by the power distribution of the left twiddle factor to obtain a second sequence; a butterfly operation unit in the qth circuit respectively carries out butterfly operation on specified data pairs in the first sequence or the second sequence to obtain a third sequence; and when the power distribution of the right twiddle factor in the q-th round is not zero, a right multiplication unit in the q-th circuit multiplies each data of the third sequence by each corresponding twiddle factor indicated by the power distribution of the right twiddle factor to obtain a fourth sequence.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

A hardware accelerator for key switching algorithm in CKKS homomorphic encryption algorithm

The application discloses a hardware accelerator for a key switching algorithm in a CKKS homomorphic encryption algorithm, and belongs to the field of privacy calculation and fully homomorphic encryption hardware acceleration. The hardware accelerator comprises an NTT / INTT circuit and a peripheral polynomial operation circuit. The NTT / INTT circuit is used for realizing domain transformation of a polynomial, and comprises a plurality of parallel butterfly operation cores and an NTT control module for generating a memory address and a data flow control. The butterfly operation core comprises a storage array, a calculation array and a peripheral digital circuit. The storage array is used for storing a rotation factor, and the calculation array is used for calculating a partial product of the rotation factor and polynomial coefficient multiplication. The hardware accelerator can realize highly parallel key switching operation, solves the problem of excessive fully pipelined data flow storage overhead caused by data dependency in other works, and effectively reduces the area and power consumption of the circuit.
Owner:ZHEJIANG UNIV

Data processing method and related apparatus

PCT designated stageWO2026109078A1Complex mathematical operationsFast Fourier transformRotation factor
Provided in the embodiments of the present application are a data processing method and a related apparatus. In the present application, a matrix multiplication operation is performed on a rotation factor matrix and an input signal by means of cube units, so as to accelerate in combination with the cube units a fast Fourier transform (FFT) operation performed on the input signal. The method comprises: on the basis of the length of a first input signal, generating a rotation factor matrix, wherein the rotation factor matrix comprises a plurality of rotation factors, which are used for performing an FFT operation on the first input signal; and performing a matrix multiplication operation on the rotation factor matrix and the first input signal, so as to obtain a first matrix-product matrix, wherein the first matrix-product matrix comprises a result of performing the FFT operation on the first input signal.
Owner:HUAWEI TECH CO LTD

Method and apparatus for storing or reading rotation factors in target transformation

ActiveCN119719590BComplex mathematical operationsRotation factorSequence transformation
Embodiments of the present specification provide a method for storing or reading a rotation factor in a target transform. The target transform is a discrete transform or its inverse transform that transforms an input coefficient sequence into an output coefficient sequence based on a rotation factor. The method includes determining a target rotation factor to be stored, whose power is a result of a multiplication of a first factor of a first number of bits and a second factor of a second number of bits modulo a target value; and storing the target rotation factor by a two-level storage manner, the two-level storage manner including storing the target rotation factor at a first address in a rotation factor storage, wherein the rotation factor storage includes a first number of first storage units, the first number being a number of different modulo multiplication results generated by the first factor of the first number of bits and the second factor of the second number of bits; and storing an index value pointing to the first address at a second address in an index storage, wherein the index storage includes a target value of second storage units, the second address corresponding to the power of the target rotation factor; and the target value is greater than the first number.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

FFT calculation method, system and device based on angle precoding circumference system CORDIC and medium

The invention is applicable to the field of digital signal processing, and discloses an FFT (Fast Fourier Transform) calculation method, system, equipment and medium based on CORDIC (Coordinated Rotation Digital Integrated Circuit) of an angle precoding circumference system, the method comprises the following steps: acquiring an input angle to be calculated, and obtaining a vector and a residual angle by inquiring a predefined distribution table; based on residual angle block coding, obtaining a non-scaling twiddle factor corresponding to the coded angle through Taylor expansion, and rotating the vector to obtain a rotated vector and a micro angle; obtaining a trigonometric function value corresponding to the input angle by combining approximate processing with multiply-add operation based on the rotated vector and the micro angle; and according to the received sampling sequence data, applying the trigonometric function value as a twiddle factor, and executing complex multiplication and data exchange in each level of butterfly operation unit of the FFT to obtain FFT spectrum data. The CORDIC algorithm with few iterations, high calculation precision and low hardware resource consumption is realized, and the FFT calculation efficiency and performance are improved.
Owner:GUIZHOU POWER GRID CO LTD

NTT hardware system and control method based on double coefficient folding storage and adaptive scheduling

The application discloses a kind of NTT hardware system and control method based on double coefficient folding storage and adaptive scheduling, it is related to post quantum cryptography hardware acceleration technical field.Its hardware system includes input layer and bit inversion module, double coefficient folding unit module, single BRAM storage module, adaptive conflict-free scheduling subsystem, rotation factor processing path and BFU butterfly computation unit;Its control method is compressed by double coefficient folding storage to splice logical coefficient, to realize high-bandwidth storage with the theoretical lower limit physical depth single body BRAM, and dynamically switches inner layer and interlayer access mode according to calculation level by adaptive conflict-free scheduling subsystem, eliminates variable step size memory conflict, realizes full-pipeline bubble-free operation.The hardware system of the application significantly reduces the on-chip storage area consumption, improves the calculation throughput and timing convergence, and is suitable for hardware acceleration implementation of post quantum cryptography algorithm.
Owner:NANJING UNIV OF INFORMATION SCI & TECH