processing unit

CN122672746APending Publication Date: 2026-09-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611152801.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

随着各类高端场景对信号处理实时性与数据吞吐量的要求持续提升,传统快速傅里叶变换实现方案逐渐显露出性能与硬件资源难以兼顾的问题

Benefits of technology

[0004]根据本申请的实施例,将M级计算电路中的第M级计算电路之前的计算电路配置为响应于接收到第m级控制信号,对计算电路的多个输入端接收到的多个操作数执行计算,得到多个计算结果,并将多个计算结果在计算电路的多个输出端输出;将M级计算电路中的第M级计算电路配置为响应于接收到第M级控制信号,对计算电路的多个输入端接收到的多个操作数执行计算,得到多个计算结果,对多个计算结果进行规格化,并将规格化的多个计算结果在计算电路的多个输出端输出。相比于相关技术中将每一级计算电路都配置为对计算电路的多个计算结果进行规格化,输出规格化的多个计算结果,本申请实施例的处理单元通过参数配置选择是否进行规格化操作,将对多个计算结果进行规格化的操作从多级计算电路中省去,并在最后一级计算电路中进行规格化操作,降低了蝶形计算对硬件资源的消耗。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122672746A_ABST
    Figure CN122672746A_ABST
Patent Text Reader

Abstract

This application discloses a processing unit, relating to the field of computing circuit technology. The processing unit includes M-level computing circuits. Computing circuits preceding the M-level computing circuit, responding to a received m-level control signal, perform calculations on multiple received operands, obtain multiple calculation results, and output these results at multiple output terminals. The M-level computing circuit, responding to the received M-level control signal, performs calculations on the received operands, obtains multiple calculation results, normalizes these results, and outputs the normalized results at multiple output terminals. Compared to normalizing multiple calculation results at each level of computing circuit, this application selects whether to perform normalization through parameter configuration, eliminating the normalization operation from multiple levels of computing circuits and performing it in the last level of computing circuit, thus reducing the hardware resource consumption of butterfly computing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computing circuit technology, and more particularly to a processing unit. Background Technology

[0002] The Fast Fourier Transform (FFT) is a core algorithm in digital signal processing, widely used in key areas such as 5G (5th Generation Mobile Communication Technology), radar detection, and medical imaging. As the demands for real-time signal processing and data throughput continue to increase in various high-end scenarios, traditional FFT implementations are increasingly revealing the difficulty of balancing performance and hardware resources. While high-speed parallel FFT operations can meet high-performance processing requirements, they incur significant hardware resource consumption and area overhead, making them unsuitable for lightweight, high-efficiency engineering applications. Summary of the Invention

[0003] This application provides a processing unit for performing butterfly calculations in Fast Fourier Transform. The processing unit includes: an M-level computing circuit, wherein multiple output terminals of the m-th level computing circuit are cross-connected with multiple input terminals of the (m+1)-th level computing circuit to form a butterfly network, where M is an integer greater than 1, and m = 1, 2, …, M-1; wherein the computing circuit further has a control terminal, wherein the control terminal of the m-th level computing circuit is configured to receive a m-th level control signal, and the control terminal of the M-th level computing circuit is configured to receive an M-th level control signal; the computing circuit is configured to: in response to the control terminal of the computing circuit receiving the M-th level control signal, perform calculations on multiple operands received at the multiple input terminals of the computing circuit to obtain multiple calculation results, normalize the multiple calculation results, and output the normalized multiple calculation results at the multiple output terminals of the computing circuit; and in response to the control terminal of the computing circuit receiving the m-th level control signal, perform calculations on multiple operands received at the multiple input terminals of the computing circuit to obtain multiple calculation results, and output the multiple calculation results at the multiple output terminals of the computing circuit.

[0004] According to embodiments of this application, the computing circuits preceding the M-th stage in an M-stage computing circuit are configured to perform calculations on multiple operands received at multiple input terminals of the computing circuit in response to receiving a control signal from the m-th stage, obtaining multiple calculation results, and outputting the multiple calculation results at multiple output terminals of the computing circuit; the M-th stage computing circuit is configured to perform calculations on multiple operands received at multiple input terminals of the computing circuit in response to receiving a control signal from the M-th stage, obtaining multiple calculation results, normalizing the multiple calculation results, and outputting the normalized multiple calculation results at multiple output terminals of the computing circuit. Compared to related technologies where each stage of the computing circuit is configured to normalize the multiple calculation results and output normalized multiple calculation results, the processing unit of this application selects whether to perform a normalization operation through parameter configuration, thus eliminating the normalization operation of multiple calculation results from the multi-stage computing circuits and performing the normalization operation in the last stage computing circuit, reducing the consumption of hardware resources by butterfly computing. Attached Figure Description

[0005] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0006] Figure 1 This is a schematic diagram of the architecture for radix-2 butterfly computation in Fast Fourier Transform.

[0007] Figure 2 This is a schematic diagram of the architecture for radix-4 butterfly computation in Fast Fourier Transform.

[0008] Figure 3 A schematic diagram of the processing flow of an adder or subtractor for radix-4 butterfly computation.

[0009] Figure 4 This is a schematic diagram of the processing unit in an embodiment of this application.

[0010] Figure 5 This is a schematic diagram of the processing unit in an embodiment of this application.

[0011] Figure 6 This is a schematic diagram of the processing flow of the processing unit according to an embodiment of this application.

[0012] Figure 7 This is a schematic diagram of an adder / subtractor of a first type according to an embodiment of this application.

[0013] Figure 8 This is a schematic diagram of the processing flow of the processing unit in an embodiment of this application.

[0014] Figure 9 This is a schematic diagram of the adder / subtractor according to an embodiment of this application.

[0015] Figure 10 This is a schematic diagram of the processing unit in an embodiment of this application.

[0016] Figure 11 This is a schematic diagram of the processing unit in an embodiment of this application.

[0017] Figure 12 This is a schematic diagram of the processing unit in an embodiment of this application.

[0018] Figure 13 This is a schematic diagram of the processing unit in an embodiment of this application.

[0019] Figure 14 This is a schematic diagram of a 4-point radix-2 FFT butterfly calculation.

[0020] Figure 15 This is a schematic diagram of an 8-point radix-2 FFT butterfly calculation.

[0021] Figure 16 This is a schematic diagram of a 16-point radix-4 FFT butterfly calculation.

[0022] Figure 17 This is a schematic diagram of the signal transformation and processing flow of Fast Fourier Transform.

[0023] Figure 18 This is a schematic diagram illustrating the decomposition principle of multidimensional FFT.

[0024] Figure 19 This is a schematic diagram of a first application scenario of the processing unit in an embodiment of this application.

[0025] Figure 20 This is a schematic diagram of a second application scenario of the processing unit in an embodiment of this application.

[0026] Figure 21 This is a schematic diagram of a third application scenario of the processing unit in an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0028] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0029] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] The Fast Fourier Transform (FFT) is a commonly used algorithm in modern digital signal processing. With continuous technological iteration and upgrades, modern communications, radar systems, medical imaging, and other fields are increasingly demanding higher throughput and real-time processing capabilities. In practical engineering applications, scenarios such as 5G communication, radar detection, and medical imaging heavily rely on the computational performance of the FFT. To meet the ultra-high downlink throughput requirement of approximately 30 Gbps per cell, 5G communication systems employ Orthogonal Frequency Division Multiplexing (OFDM) multi-carrier technology, transmitting data in parallel through multiple mutually orthogonal subcarriers. Since OFDM relies on the Fast Fourier Transform algorithm to construct an orthogonal basis for the subcarriers, the data processing throughput of the Fast Fourier Transform directly determines the upper limit of the 5G communication transmission rate. In radar systems, equipment detects targets by emitting detection pulses and receiving target echoes. Signal processing operations, such as filtering based on Fast Fourier Transform (FFT), must be completed within a very short time. It is essential to ensure that each FFT operation finishes before the next detection pulse arrives, allowing sufficient time for subsequent system decisions and execution. In medical imaging scenarios such as Magnetic Resonance Imaging (MRI), the FFT algorithm is primarily used for image reconstruction of raw acquired data. Its processing speed directly impacts clinical diagnostic efficiency and patient experience. The industry has clear standards for image reconstruction rates in these scenarios: 3000 frames per second for 128×128 resolution images and 800 frames per second for 256×256 resolution images.

[0031] Various engineering scenarios have high standards for FFT computation speed and data throughput, and radix-4 butterfly networks are the mainstream structure for achieving efficient FFT computation. Taking a 16-point radix-4 FFT as an example, the overall structure includes two computation levels, with four butterfly computations per level, for a total of eight butterfly computations. Each butterfly unit needs to complete three sets of complex multiplications and eight sets of complex addition and subtraction operations. Although multi-path parallel architecture can effectively improve the FFT computation speed, it will significantly increase hardware resource consumption and chip area. Therefore, reducing the resource consumption of butterfly computations is the current focus of FFT hardware optimization.

[0032] This application provides a processing unit for performing butterfly calculations in the Fast Fourier Transform.

[0033] The Fast Fourier Transform (FFT) can transform a digital discrete signal from the time domain to the frequency domain. This is possible if the number of sampling points in the digital discrete signal satisfies the following condition. (where n is an integer and the signal length is a power of 2), a radix-2 FFT butterfly calculation can be used. The time-domain sequence of the digital discrete signal is decimated according to the parity of the sampling number, and split into two groups of length... The subsequences are processed by performing a decimation-in-time (DIT) Fourier transform on both sets of subsequences to obtain the calculation results of the first set of subsequences corresponding to the even-position subsequences. The calculation results of the second group of subsequences corresponding to the odd position subsequences The two sets of calculation results are then combined using a fixed cross formula, which is called radix-2 FFT butterfly calculation.

[0034] The butterfly equations of the architecture of radix-2 butterfly computation:

[0035] (1);

[0036] in, , The indexes are respectively The two inputs of the butterfly unit; , The indexes are respectively Two-way output of the butterfly unit; twitch factor This can be used to compensate for the time-domain delay phase difference between even-position and odd-position subsequences at one sampling point. The rotation factor has periodicity and symmetry, and is the mathematical basis for the separation of addition and subtraction branches in butterfly calculation.

[0037] Figure 1 This is a schematic diagram of the architecture for radix-2 butterfly computation in Fast Fourier Transform.

[0038] like Figure 1 As shown, This represents complex number multiplication. This represents complex number addition. This represents the operation of subtraction of complex numbers.

[0039] The architecture of radix-2 butterfly computing includes two inputs and two outputs. The architecture of radix-2 butterfly computing includes one complex multiplier, one complex adder and one complex subtractor. The signal lines intersect to form a butterfly network.

[0040] Results of subsequence calculation for group 1 The input is divided into two paths: one for a complex adder and the other for a complex subtractor. The result of the second subsequence calculation. Entering the complex multiplier and the rotation factor Multiply, we get This corresponds to the multiplication term in formula (1). The complex adder calculates the result of the first subsequence. Multiplication terms Perform a summation operation and output the summation result. The complex subtractor calculates the result for the first subsequence. Multiplication terms Perform a difference operation and output the result. The wiring pattern of the two signals branching and being cross-connected to the complex adder and complex subtractor resembles the wings of a butterfly, hence the name butterfly network for this arithmetic unit.

[0041] If the number of sampling points contained in the digital discrete signal satisfies (where n is an integer and the signal length is a power of 4), a radix-4 FFT butterfly calculation can be used. The time-domain sequence of the digital discrete signal is decimated modulo 4 according to its index, and split into 4 groups of length... The first subsequence is calculated by performing Discrete Fourier Transform on each of the four subsequences. Results of the second subsequence calculation Results of the calculation of the third subsequence The results of the calculation of the fourth subsequence The four sets of calculation results are then merged using a fixed 4-way cross-calculation formula, which is called radix-4 FFT butterfly calculation.

[0042] The butterfly equations of the architecture of radix-4 butterfly computation:

[0043] (2);

[0044] in, , , , The indexes are respectively The four inputs of the butterfly unit; , , , The indexes are respectively The butterfly unit has four outputs.

[0045] Figure 2 This is a schematic diagram of the architecture for radix-4 butterfly computation in Fast Fourier Transform.

[0046] like Figure 2 As shown, the architecture of the radix-4 butterfly computation includes three levels of addition and subtraction units, corresponding to... Figure 2 The three columns of operation symbols are: Level 1 addition and subtraction unit includes 3 complex multipliers, each containing an adder and a subtractor; Level 2 addition and subtraction unit includes 2 complex adders and 2 complex subtractors; Level 3 addition and subtraction unit includes 2 complex adders and 2 complex subtractors, each containing an adder and a subtractor.

[0047] Results of subsequence calculation for group 1 It is divided into two paths, which are respectively input to the first complex number adder and the first complex number subtractor of the second-level addition and subtraction calculation unit.

[0048] Results of subsequence calculation for group 2 The result of the calculation of the second complex multiplier and the second subsequence in the first-level addition and subtraction calculation unit. rotation factor Multiplying them together gives the first multiplication term. The first multiplication term It is divided into two paths, which are respectively input to the second complex number adder and the second complex number subtractor of the second-level addition and subtraction calculation unit.

[0049] Results of subsequence calculation for group 3 The result of the calculation of the first complex multiplier and the third subsequence in the first-level addition and subtraction calculation unit. rotation factor Multiplying them together gives us the second multiplication term. The second multiplication term It is divided into two paths, which are respectively input to the first complex number adder and the first complex number subtractor of the second-level addition and subtraction calculation unit.

[0050] Results of subsequence calculation for group 4 The result of the calculation of the third complex multiplier and the fourth subsequence in the first-level addition and subtraction calculation unit. rotation factor Multiplying them together gives us the third multiplication term. The third multiplication term It is divided into two paths, which are respectively input to the second complex number adder and the second complex number subtractor of the second-level addition and subtraction calculation unit.

[0051] The calculation result of the first complex adder in the second-level addition and subtraction unit is input into the first complex adder and the first complex subtractor in the third-level addition and subtraction unit via two separate inputs. The calculation result of the first complex subtractor in the second-level addition and subtraction unit is input into the second complex adder and the second complex subtractor in the third-level addition and subtraction unit via two separate inputs. The calculation result of the second complex adder in the second-level addition and subtraction unit is input into the first complex adder and the first complex subtractor in the third-level addition and subtraction unit via two separate inputs. The calculation result of the second complex subtractor in the second-level addition and subtraction unit is input into the second complex adder and the second complex subtractor in the third-level addition and subtraction unit via two separate inputs. The third-level addition and subtraction unit outputs four calculation results.

[0052] Figure 3 A schematic diagram of the processing flow of an adder or subtractor for radix-4 butterfly computation.

[0053] Figure 3 This represents a three-stage addition and subtraction unit for radix-4 butterfly computation. The adders and subtractors in each stage of the addition and subtraction unit need to perform alignment, mantissa processing, and normalization on the two input signals in sequence to complete the addition or subtraction operation on the two input signals.

[0054] This application proposes a processing unit. The processing unit includes M-level computing circuits, wherein multiple output terminals of the m-th level computing circuit are cross-connected with multiple input terminals of the (m+1)-th level computing circuit to form a butterfly network, where M is an integer greater than 1, and m = 1, 2, …, M-1. The computing circuit also has a control terminal, wherein the control terminal of the m-th level computing circuit is configured to receive the m-th level control signal, and the control terminal of the M-th level computing circuit is configured to receive the M-th level control signal.

[0055] The computing circuit is configured to: in response to the control terminal of the computing circuit receiving the Mth level control signal, perform calculations on multiple operands received at multiple input terminals of the computing circuit, obtain multiple calculation results, normalize the multiple calculation results, and output the normalized multiple calculation results at multiple output terminals of the computing circuit; and in response to the control terminal of the computing circuit receiving the mth level control signal, perform calculations on multiple operands received at multiple input terminals of the computing circuit, obtain multiple calculation results, and output the multiple calculation results at multiple output terminals of the computing circuit.

[0056] Figure 4 This is a schematic diagram of the processing unit in an embodiment of this application.

[0057] In one example, m=1, 2, M=3, and the processing unit in this embodiment includes a 3-level computing circuit. For example... Figure 4As shown, the first output terminal B1 and the second output terminal B2 of the first-level computing circuit 10_1 are connected to the second input terminal A2 and the first input terminal A1 of the second-level computing circuit 10_2, respectively. The first output terminal B1 and the second output terminal B2 of the second-level computing circuit 10_2 are connected to the second input terminal A2 and the first input terminal A1 of the third-level computing circuit 10_3, respectively, forming a butterfly network.

[0058] The first-level computing circuit 10_1 responds to the first-level control signal received at its control terminal C, performs calculations on the multiple operands received at the multiple input terminals of the first-level computing circuit 10_1, obtains multiple calculation results, and outputs the multiple calculation results at the multiple output terminals of the first-level computing circuit 10_1.

[0059] The second-level computing circuit 10_2 receives multiple unnormalized calculation results from the first-level computing circuit 10_1 at multiple output terminals. Responding to a second-level control signal received at its control terminal C, the second-level computing circuit 10_2 performs calculations on the received unnormalized calculation results from the first-level computing circuit 10_1, obtaining its own calculation results, and then outputs these results.

[0060] The third-level computing circuit 10_3 receives multiple unnormalized calculation results from the second-level computing circuit 10_2 at multiple output terminals of the third-level computing circuit 10_2. Responding to a third-level control signal received at its control terminal C, the third-level computing circuit 10_3 performs calculations on the received unnormalized calculation results from the second-level computing circuit 10_2, obtaining multiple calculation results for the third-level computing circuit 10_3. It then normalizes these calculation results and outputs them at multiple output terminals of the third-level computing circuit 10_3.

[0061] According to embodiments of this application, the computing circuits preceding the M-th stage in an M-stage computing circuit are configured to perform calculations on multiple operands received at multiple input terminals of the computing circuit in response to receiving a control signal from the m-th stage, obtaining multiple calculation results, and outputting the multiple calculation results at multiple output terminals of the computing circuit; the M-th stage computing circuit is configured to perform calculations on multiple operands received at multiple input terminals of the computing circuit in response to receiving a control signal from the M-th stage, obtaining multiple calculation results, normalizing the multiple calculation results, and outputting the normalized multiple calculation results at multiple output terminals of the computing circuit. Compared to related technologies where each stage of the computing circuit is configured to normalize the multiple calculation results and output normalized multiple calculation results, the processing unit of this application selects whether to perform a normalization operation through parameter configuration, thus eliminating the normalization operation of multiple calculation results from the multi-stage computing circuits and performing the normalization operation in the last stage computing circuit, reducing the consumption of hardware resources by butterfly computing.

[0062] According to an embodiment of this application, the computing circuit includes multiple adders and subtractors. The first and second input terminals of the adders and subtractors serve as input terminals of the computing circuit, and the output terminals of the adders and subtractors serve as output terminals of the computing circuit. Each adder and subtractor also has a normalization control terminal, which is connected to the control terminal of the computing circuit. The adder and subtractor are configured to: in response to the normalization control terminal receiving an M-th level control signal, perform addition and / or subtraction operations on the operands received at the first and second input terminals of the adder and subtractor to obtain a calculation result; normalize the calculation result; and output the normalized calculation result at the output terminal of the adder and subtractor; and in response to the normalization control terminal receiving an m-th level control signal, perform addition and / or subtraction operations on the operands received at the first and second input terminals of the adder and subtractor to obtain a calculation result, and output the calculation result at the output terminal of the adder and subtractor.

[0063] Figure 5 This is a schematic diagram of the processing unit in an embodiment of this application.

[0064] In one instance, m=1, M=2. For example... Figure 5As shown, the first-stage calculation circuit 10_1 includes multiple adders and subtractors 11. The first input terminal c1 and the second input terminal c2 of the two adders and subtractors 11 of the first-stage calculation circuit 10_1 serve as the first input terminal A1, the second input terminal A2, the third input terminal A3, and the fourth input terminal A4 of the first-stage calculation circuit 10_1. The output terminal c3 of the two adders and subtractors 11 of the first-stage calculation circuit 10_1 serves as the first output terminal B1 and the second output terminal B2 of the first-stage calculation circuit 10_1. Each of the two adders and subtractors 11 of the first-stage calculation circuit 10_1 has a normalization control terminal c0. The normalization control terminal c0 of each of the two adders and subtractors 11 of the first-stage calculation circuit 10_1 is connected to the control terminal C of the first-stage calculation circuit 10_1 to receive the first-stage control signal. Under the control of the first-level control signal, the multiple adders and subtractors 11 of the first-level calculation circuit 10_1 perform addition and / or subtraction operations on the operands received by the first input terminal c1 and the second input terminal c2 of the multiple adders and subtractors 11 of the first-level calculation circuit 10_1, respectively, to obtain the calculation result, and output the calculation result at the output terminal c3 of the multiple adders and subtractors 11 of the first-level calculation circuit 10_1.

[0065] The normalization control terminal c0 of each of the multiple adders and subtractors 11 in the second-level computing circuit 10_2 is connected to the control terminal C of the second-level computing circuit 10_2 to receive the second-level control signal. Under the control of the second-level control signal, the multiple adders and subtractors 11 in the second-level computing circuit 10_2 perform addition and / or subtraction operations on the operands received by each of the multiple adders and subtractors 11 in the second-level computing circuit 10_2 at their respective first input terminal c1 and second input terminal c2, respectively, to obtain the calculation result. The calculation result is then normalized, and the normalized calculation result is output at the respective output terminal c3 of each of the multiple adders and subtractors 11 in the second-level computing circuit 10_2.

[0066] According to embodiments of this application, control signals are configured so that multiple adders and subtractors in different levels of the computing circuit can perform addition and / or subtraction operations on the received operands under the control of the m-th level control signal to obtain the calculation result; and can perform addition and / or subtraction operations on the received operands under the control of the M-th level control signal to obtain the calculation result, and normalize the calculation result and output the normalized calculation result. This realizes flexible control over whether to perform normalization operations, and thus enables selective omission of normalization operations in the butterfly calculation of Fast Fourier Transform, reducing the consumption of hardware resources in the butterfly calculation of Fast Fourier Transform.

[0067] According to an embodiment of this application, the adder / subtractor includes: an alignment module connected to a first input terminal and a second input terminal of the adder / subtractor, used to align the first operand received at the first input terminal with the second operand received at the second input terminal, such that the exponent of the first operand is equal to the exponent of the second operand; a mantissa operation module connected to the alignment module, used to perform addition and / or subtraction operations on the mantissas of the first operand and the second operand output by the alignment module to obtain a calculation result; and a normalization module connected to the mantissa operation module, used to perform carry or normalization processing on the calculation result of the mantissa operation module, wherein when the calculation result generates a carry to the integer part, a carry is performed, and when the calculation result generates leading zeros, normalization processing is performed; wherein the normalization module is connected to the normalization control terminal of the adder / subtractor, and the normalization module is configured to enable normalization processing in response to the normalization control terminal receiving an M-th level control signal, and to disable normalization processing in response to the normalization control terminal receiving an m-th level control signal.

[0068] According to embodiments of this application, the multiple operands received by the multiple input terminals of the computing circuit are floating-point operands, which include a sign bit, an exponent bit, and a mantissa bit. The floating-point operands can be configured as 16-bit half-precision, 32-bit single-precision, or 64-bit double-precision floating-point numbers, and can be adapted to different precision floating-point operations by adjusting the bit width of the exponent and mantissa bits.

[0069] In one example, floating-point operands can use a 32-bit storage width, which includes 1 sign bit, 8 exponent bits, and 23 mantissa bits. The calculation rules for single-precision floating-point numbers are shown in Table 1, and their numerical expressions are as follows: To ensure that the output of complex number operations conforms to the standard floating-point format, the calculation results need to be normalized.

[0070] Table 1

[0071]

[0072] Taking the 32-bit hexadecimal floating-point number C1C6_0000 as an example, the complete calculation process is as follows: First, the hexadecimal value is converted into the corresponding 32-bit binary sequence, and then it is split according to the bit field, with the highest bit being the sign bit. A value of 1 indicates that the final result is negative; the middle 8 bits are the exponent bits. The corresponding binary sequence is 1000_0011, which translates to 131 in decimal. Since single-precision floating-point exponents use a 127 offset encoding, the actual exponent in the exponentiation operation needs to be obtained by subtracting 127 from the offset exponent value. The remaining lower 23 bits of the original mantissa are represented by the binary sequence 100_0110_0000_0000_0000_0000. In the definition of normalized single-precision floating-point numbers, the mantissa portion has a hidden integer bit 1 by default, which does not occupy storage bits; therefore, the complete mantissa expression needs to be supplemented. In the original binary mantissa, only the first, fifth, and sixth bits are at a level of 1, corresponding to the negative power term. The complete mantissa expression after integration is: Substituting the sign factor, complete mantissa, and true exponent into the general formula for floating-point numerical calculation yields the complete arithmetic expression. .

[0073] Complex multipliers, adders, and subtractors use fixed-point numbers to perform multiplication, addition, and subtraction operations. Fixed-point numbers are binary numbers with a fixed decimal point position. Multi-level iterative complex number operations are prone to numerical overflow and loss of effective precision. After normalization, the calculation results can be converted into standard 32-bit single-precision floating-point numbers.

[0074] The processing unit in this embodiment is designed to simplify normalization processing and integrate addition and subtraction operations, thereby reducing redundant calculations in the order-alignment module and forming a simplified FFT hardware processing unit, taking into account the characteristics of the operands in the butterfly calculation of the Fast Fourier Transform.

[0075] Figure 6 This is a schematic diagram of the processing flow of the processing unit according to an embodiment of this application.

[0076] Figure 6 The processing flow of a three-stage computing circuit of a processing unit according to an embodiment of this application is illustrated. In the first-stage computing circuit, each of the two first-type adders / subtractors can, under the control of a first-stage control signal, perform alignment and mantissa processing on the multiple operands received by each of the two first-type adders / subtractors in the first-stage computing circuit. In the second-stage computing circuit, each of the two first-type adders / subtractors can, under the control of a second-stage control signal, perform alignment and mantissa processing on the multiple operands received by each of the two first-type adders / subtractors in the second-stage computing circuit. In the third-stage computing circuit, each of the two first-type adders / subtractors can, under the control of a third-stage control signal, perform alignment, mantissa processing, and normalization processing on the multiple operands received by each of the two first-type adders / subtractors in the third-stage computing circuit. The first-type adders / subtractors can perform normalization selectively by configuring control signals; therefore, the first-type adders / subtractors can also be called simplified adders / subtractors.

[0077] Figure 7 This is a schematic diagram of an adder / subtractor of a first type according to an embodiment of this application.

[0078] According to embodiments of this application, such as Figure 7 As shown, the first type of adder / subtractor 11_1 includes an alignment module 21, a mantissa operation module 22, and a normalization module 23.

[0079] The alignment module 21 can be connected to the first input terminal and the second input terminal of the adder / subtractor 11, and is used to perform alignment processing on the first operand received by the first input terminal and the second operand received by the second input terminal, so that the exponent of the first operand is equal to the exponent of the second operand.

[0080] In one example, using a smaller-to-larger-order alignment rule can reduce computational errors caused by mantissa shifting. The difference between the exponents of the first and second operands is taken, and the mantissa of the operand with the smaller exponent is shifted right by a number of bits equal to the exponent difference, with leading zeros padded during the shift. For example, if the difference between the exponents of the first and second operands is 5, then the mantissa of the operand with the smaller exponent is shifted right by 5 bits, with leading zeros padded.

[0081] The mantissa operation module 22 is connected to the alignment module 21. The mantissa operation module 22 is used to perform addition and / or subtraction operations on the mantissas of the first operand and the second operand output by the alignment module 21 to obtain the calculation result.

[0082] After aligning the first and second operands, the exponents of the first and second operands become equal. The addition and subtraction of the first and second operands essentially transforms into signed addition and subtraction of the mantissas of the first and second operands. The subtraction instruction is marked as 1, and the addition instruction is marked as 0. If a subtraction operation is performed (instruction = 1), and the second operand is positive (sign bit S = 0), the combination is marked as 1, and subtraction is actually performed. If a subtraction operation is performed (instruction = 1), and the second operand is negative (sign bit S = 1), the combination is marked as 0, and addition is actually performed. The combined sign bit is then padded to the highest bit of each of the two mantissas. By performing addition and subtraction on the two sets of signed mantissas, the calculation result of the mantissa operation module 22 is obtained.

[0083] The normalization module 23 is connected to the mantissa operation module 22. The normalization module 23 is used to perform carry or normalization processing on the calculation result of the mantissa operation module 22. Specifically, when the calculation result generates a carry to the integer part, the carry is performed, and when the calculation result generates a leading zero, the normalization processing is performed.

[0084] The value of the calculation result from the mantissa operation module 22 can be either amplified or reduced. If the value of the calculation result from the mantissa operation module 22 increases, the normalization module 23, in the enabled state, performs carry processing on the calculation result from the mantissa operation module 22. For example, when the mantissa of the first operand and the second operand are added to produce a carry to the integer part, the carry is performed: the exponent is incremented by 1, and the mantissa is shifted one position to the right. If the value of the calculation result from the mantissa operation module 22 decreases, the normalization module 23, in the enabled state, performs normalization processing on the calculation result from the mantissa operation module 22: based on the position of the first digit 1 in the mantissa of the calculation result from the mantissa operation module 22, the mantissa is shifted to the left to the highest bit, which is 1. The number of positions shifted corresponds to the number of positions the exponent is decremented. For example, if the first 1 in the mantissa appears in the 5th position, the mantissa is shifted 5 positions to the left, and the exponent is decremented by 5.

[0085] Normalization module 23 is connected to the normalization control terminal of the adder / subtractor. Normalization module 23 is configured to enable normalization processing in response to the normalization control terminal receiving the M-th level control signal, and to disable normalization processing in response to the normalization control terminal receiving the m-th level control signal.

[0086] According to embodiments of this application, the adder / subtractor includes an alignment module, a mantissa operation module, and a normalization module, which performs addition and / or subtraction operations on the first and second operands of a single-precision floating-point number. The normalization module is connected to the normalization control terminal of the adder / subtractor and configures control signals so that the normalization modules in multiple adders / subtractors at different levels of the calculation circuit are disabled under the control of the m-th level control signal and enabled under the control of the M-th level control signal, and normalize the calculation result, outputting the normalized calculation result. This achieves flexible control over whether to perform normalization operations.

[0087] like Figure 7 As shown, the alignment module 21 of the first type of adder / subtractor 11_1 may include a mantissa completion module 211 and a mantissa shifting module 212. The mantissa completion module 211 is connected to the normalization control terminal of the adder / subtractor. The mantissa shifting module 212 is connected to the mantissa completion module 211.

[0088] According to an embodiment of this application, the mantissa completion module 211 is configured to, in response to the normalization control terminal receiving a first-level control signal, take the mantissa of a predetermined number of bits from the first operand and the second operand to obtain a first mantissa and a second mantissa, and pad the first and second mantissas with a first value of a predetermined number of bits; in response to the normalization control terminal receiving any one of the second-level control signals to the Mth-level control signals, take the mantissa of a predetermined number of bits from the first preprocessing operand and the second preprocessing operand to obtain a first mantissa and a second mantissa, and pad the last digit of the first and second mantissas with a second value of a predetermined number of bits.

[0089] According to an embodiment of this application, the mantissa shifting module 212 is used to calculate the exponent difference between the first mantissa and the second mantissa output by the mantissa completion module 211 as the shift number, and shift one of the first mantissa and the second mantissa by the shift number to obtain the first mantissa and the second mantissa with the same exponent.

[0090] like Figure 7 As shown, in one example, the first type of adder / subtractor 11_1 in the first-level calculation circuit receives a first-level control signal STEP with a value of 0. Under the control of the first-level control signal STEP with a value of 0, the mantissa completion module 211 in the first-level calculation circuit fills the first and second mantissas with a first value of a predetermined number of bits, and the normalization module 23 in the first-level calculation circuit is disabled.

[0091] like Figure 7 As shown, in another example, the first type of adder / subtractor 11_1 in the second-level calculation circuit receives a second-level control signal STEP with a value of 1. Under the control of the second-level control signal STEP with a value of 1, the mantissa completion module 211 in the second-level calculation circuit fills the last digit of the first mantissa and the second mantissa with a second value of a predetermined number, and the normalization module 23 in the second-level calculation circuit is disabled.

[0092] like Figure 7 As shown, in another example, the first type of adder / subtractor 11_1 in the M-level calculation circuit receives the M-level control signal STEP with a value of 2. Under the control of the M-level control signal STEP with a value of 2, the mantissa completion module 211 in the M-level calculation circuit fills the last digit of the first mantissa and the second mantissa with a second value of a predetermined number, and the normalization module 23 in the M-level calculation circuit is enabled.

[0093] According to embodiments of this application, the exponent adjustment module, through a mantissa completion module and a mantissa shift module, adjusts the exponent bits of the first and second operands to have the same exponent, thereby obtaining a first and second mantissa that can be directly added or subtracted. The mantissa completion module is also controlled by the normalization control terminal, thus allowing normalized and denormalized operands to be mantissa-completed in different ways to accommodate the omission of normalization operations and ensure the accuracy of the calculation results.

[0094] According to embodiments of this application, such as Figure 7As shown, the exponent-alignment module 21 may further include a sign preprocessing module 213 and an operand preprocessing module 214. The sign preprocessing module 213 is connected to the second input of the adder / subtractor. The operand preprocessing module 214 is connected to the first input of the adder / subtractor and the sign preprocessing module. The sign preprocessing module 213, connected to the second input of the adder / subtractor, is configured to receive addition / subtraction instructions, perform an XOR operation between the sign bit of the second operand and the addition / subtraction instruction, and obtain a second operand with addition / subtraction operators. The operand preprocessing module is used to arrange the first operand and the second operand with addition / subtraction operators according to exponent size and then provide them to the mantissa completion module.

[0095] According to embodiments of this application, such as Figure 7 As shown, the mantissa operation module 22 of the first type of adder / subtractor 11_1 may include a mantissa sign processing module 221 and a mantissa addition module 222. The mantissa sign processing module 221 is connected to the mantissa shifting module 212. The mantissa addition module 222 is connected to the mantissa sign processing module 221. The mantissa sign processing module 221 is used to receive the first mantissa and the second mantissa output by the mantissa shifting module 212, and pads the first mantissa with a first sign value of a predetermined number of bits according to the sign of the first operand, and pads the second mantissa with a second sign value of a predetermined number of bits according to the sign of the second operand. The mantissa addition module 222 is used to perform an addition operation on the first mantissa and the second mantissa output by the mantissa sign processing module 221, and outputs the result of the addition operation as the calculation result at the output terminal of the adder / subtractor.

[0096] According to embodiments of this application, configuring different control signals allows the adders and subtractors in the computing circuit to selectively perform normalization under the control of the corresponding control signals, simplifying the calculation process. The first type of adder / subtractor can control the execution of addition or subtraction operations based on addition / subtraction instructions, reusing a single set of computing circuitry. This reduces hardware resource consumption, lowers chip power consumption and area, provides flexible switching, and balances computational efficiency with design versatility.

[0097] According to embodiments of this application, such as Figure 7 As shown, the first type of adder / subtractor 11_1 also includes a delay module 24, a sign storage module 25, and an exponent storage module 26. The delay module 24 is used to synchronize the calculation result with a timing signal. The sign storage module 25 is used to store the sign bits of the first and second operands provided by the operand preprocessing module 214. The exponent storage module 26 is used to store the exponents of the first and second operands provided by the operand preprocessing module.

[0098] According to an embodiment of this application, the delay module 24 can receive the input data valid signal enin, and delay the input data valid signal enin according to the number of clock cycles consumed by the addition and subtraction operations, so that the calculation result dou is synchronized with the timing signal, and outputs the synchronous data valid signal enout. This can align the effective timing of the calculation result dou, ensuring that the calculation result dou and the synchronous data valid signal enout are output synchronously. Dou does not participate in numerical calculation, but only serves as a timing synchronization buffer. The sign bits of the positive and negative identifiers obtained after sign preprocessing of the first operand and the second operand and operand preprocessing can also be temporarily sent to the mantissa operation module 22 for determining the positive and negative values ​​of the values ​​during the mantissa addition stage and completing the subtraction inversion logic. The sign bits are stored independently throughout the process to avoid the loss of sign information in multi-level mantissa pipeline operations. The exponent storage module 26 can cache the exponent difference of the two input floating-point numbers. The exponent storage module 26 can also directly send the exponent difference to the normalization module 23 to provide an exponent reference for normalization and carry adjustment after mantissa summation.

[0099] The following detailed explanation of the operation process of the adder / subtractor in the embodiments of this application is provided through a specific example.

[0100] like Figure 7 As shown, the process of addition and subtraction of the first type of adder / subtractor 11_1 in the first-stage computing circuit to the first operand dina and the second operand dinb is as follows.

[0101] (1) Input the second operand dinb and the addition / subtraction instruction opt into the sign preprocessing module 213. Perform an XOR operation between the sign bit S of the second operand dinb and the addition / subtraction instruction opt to obtain the second operand dinb with addition / subtraction operator signs.

[0102] The addition / subtraction instruction opt=1.

[0103] The first operand, dina, is 1100_0001_1000_1001_0100_1000_0000_0000.

[0104] The second operand, dinb, is 1100_0010_1000_0011_1010_0100_1110_0001.

[0105] The operator signa = 1 for the first operand dina.

[0106] The operator signb for the second operand dimb is 0.

[0107] (2) The first operand dina and the second operand dinb with addition and subtraction operators are arranged in order of exponent size and provided to the mantissa completion module 211. The operand with the larger exponent between the first operand dina and the second operand dinb with addition and subtraction operators is defined as the larger operand din0, and the operand with the smaller exponent is defined as the smaller operand din1.

[0108] The exponent of the first operand is expa = 1000_0011.

[0109] The exponent of the second operand is expb = 1000_0101.

[0110] The exponent expb of the second operand is greater than the exponent expa of the first operand. The sign of the larger operand din0 is sign00=0; the sign of the smaller operand din1 is sign01=1.

[0111] The exponent of the larger operand din0 is exp00 = 1000_0101.

[0112] The exponent of the smaller operand din1 is exp01 = 1000_0011.

[0113] Larger operand din0 = 1100_0010_1000_0011_1010_0100_1110_0001.

[0114] The smaller operand din1 = 1100_0001_1000_1001_0100_1000_0000_0000.

[0115] (3) The larger operand din0 and the smaller operand din1 received by the mantissa completion module 211 in the first-level calculation circuit are standard floating-point numbers. Under the control of the first-level control signal, the mantissa completion module 211 in the first-level calculation circuit takes a predetermined number of bits, such as the lower 23 bits, of the mantissa of the larger operand din0 and the smaller operand din1 to obtain the first mantissa and the second mantissa. The first bit of the first mantissa and the second mantissa is padded with the first value of the predetermined number of bits, such as 1, to form a 24-bit mantissa. The first mantissa part00 and the second mantissa part01 are output.

[0116] The first last digit, part00, is 1000_0011_1010_0100_1110_0001.

[0117] The second last digit, part01, is 1000_1001_0100_1000_0000_0000.

[0118] (4) The exponent of the larger operand din0 is greater than that of the smaller operand din1. The difference between the exponents of the first mantissa part00 and the second mantissa part01 output by the mantissa completion module 211 is used as the shift bit: exp00 - exp01 = 2. The second mantissa part01 is shifted by 2 bits to obtain the first mantissa and the second mantissa with the same exponent. The value of the exponent bit is the exponent of the larger operand din0. The first mantissa and the second mantissa with the same exponent are defined as part00 and part11, respectively.

[0119] The first last digit, part00, is 1000_0011_1010_0100_1110_0001.

[0120] The second mantissa with the same exponent is part11 = 0010_0010_0101_0010_0000_0000.

[0121] (5) Perform signed addition and subtraction of mantissa. The mantissa sign processing module 221 receives the first mantissa part00 and the second mantissa part11 with the same exponent as the first mantissa. Based on the sign bit of the larger operand din0, the first sign value of the first mantissa part00 is padded with a predetermined number of bits. Based on the sign bit of the smaller operand din1, the second sign value of the second mantissa part11 with the same exponent as the first mantissa is padded with a predetermined number of bits.

[0122] The sign of the first mantissa, part00, is 0. The first sign value of part00 is padded with a predetermined number of bits, for example, two zeros. The sign of the second mantissa, part11, which has the same exponent as the first mantissa, is 1. The first sign value of part11, which has the same exponent as the first mantissa, is padded with a predetermined number of bits, for example, two 1s. Since complex numbers are stored in two's complement form in the processor, the mantissa with a sign bit of 1 is inverted and incremented by 1 to obtain the 26-bit first and second mantissas. These 26-bit first and second mantissas are defined as part20 and part21, respectively.

[0123] The first digit of the 26-bit number, part20, is 00_1000_0011_1010_0100_1110_0001.

[0124] The second digit of the 26-bit number, part21, is 11_1101_1101_1010_1110_0000_0000.

[0125] (6) Perform addition on the 26-bit first mantissa part20 and the 26-bit second mantissa part21 output by the mantissa symbol processing module 221, and output the result of the addition operation as the calculation result at the output of the adder / subtractor.

[0126] The sum of the last digits is part3 = 00_0110_0001_0101_0010_1110_0001.

[0127] The result of the addition operation is that the sign of the sum of the mantissas part3 is positive (the highest bit is 0). The highest bit of the sum of the mantissas part3 is saved to the sign of the sum of the mantissas part3, sign4, so the sign of the sum of the mantissas part3 is sign4=0.

[0128] (7) The 24th bit of the summation mantissa part3, part3

[24] =0, the exponent of the summation mantissa part3 is equal to the exponent of the larger operand din0, exp00=1000_0101, without carry. Define the lower 24 bits of the summation mantissa part3 as the 24-bit summation mantissa part4, the exponent of the 24-bit summation mantissa part4, exp4=1000_0101.

[0129] The 24-bit sum of the mantissas part4 is 0110_0001_0101_0010_1110_0001.

[0130] (8) Directly concatenate the sign4 of the summation mantissa part3, the exponent exp4 of the 24-bit summation mantissa part4, and the high 23 bits of the 24-bit summation mantissa part4 to obtain the calculation result of the first-level calculation circuit. This result is a non-standard floating-point format.

[0131] The result of the first level calculation is dou1=0100_0010_1011_0000_1010_1001_0111_0000.

[0132] like Figure 7 As shown, the process of addition and subtraction of the first operand dina and the second operand dinb by the M-level computing circuit is as follows.

[0133] (1) Taking addition as an example, the second operand dinb and the addition / subtraction instruction opt are input into the sign preprocessing module 213. The sign bit S of the second operand dinb is XORed with the addition / subtraction instruction opt to obtain the second operand dinb with addition / subtraction operator signs.

[0134] The addition / subtraction instruction opt=0.

[0135] The first operand, dina, is 0100_0010_0101_0101_0001_0111_1111_1110.

[0136] The second operand, dinb, is 1100_0011_0100_1100_1101_0111_1101_1111.

[0137] The operator signa = 0 for the first operand dina.

[0138] The operator signb = 1 for the second operand dimb.

[0139] (2) The first operand dina and the second operand dinb with addition and subtraction operators are arranged in order of exponent size and provided to the mantissa completion module 211. The operand with the larger exponent between the first operand dina and the second operand dinb with addition and subtraction operators is defined as the larger operand din0, and the operand with the smaller exponent is defined as the smaller operand din1.

[0140] The exponent of the first operand is expa = 1000_0100.

[0141] The exponent of the second operand is expb = 1000_0110.

[0142] The exponent expb of the second operand is greater than the exponent expa of the first operand. The sign of the larger operand din0 is sign00=1; the sign of the smaller operand din1 is sign01=0.

[0143] The exponent of the larger operand din0 is exp00 = 1000_0110.

[0144] The exponent of the smaller operand din1 is exp01 = 1000_0100.

[0145] Larger operand din0 = 1100_0011_0100_1100_1101_0111_1101_1111.

[0146] The smaller operand din1 = 0100_0010_0101_0101_0001_0111_1111_1110.

[0147] (3) The larger operand din0 and the smaller operand din1 received by the mantissa completion module 211 in the M-level calculation circuit are non-standard floating-point numbers that have not been normalized. Under the control of the M-level control signal, the mantissa completion module 211 in the M-level calculation circuit takes a predetermined number of bits, such as the lower 23 bits, of the mantissa of the larger operand din0 and the smaller operand din1 to obtain the first mantissa and the second mantissa. The last bit of the first mantissa and the second mantissa is padded with a second value of a predetermined number of bits, such as 0, to form a 24-bit mantissa. The first mantissa part00 and the second mantissa part01 are then output.

[0148] The first last digit, part00, is 1001_1001_1010_1111_1011_1110.

[0149] The second last digit, part01, is 1010_1010_0010_1111_1111_1100.

[0150] (4) The exponent of the larger operand din0 is greater than that of the smaller operand din1. The difference between the exponents of the first mantissa part00 and the second mantissa part01 output by the mantissa completion module 211 is used as the shift bit: exp00 - exp01 = 2. The second mantissa part01 is shifted by 2 bits to obtain the first mantissa and the second mantissa with the same exponent. The value of the exponent bit is the exponent of the larger operand din0. The first mantissa and the second mantissa with the same exponent are defined as part00 and part11, respectively.

[0151] The first last digit, part00, is 1001_1001_1010_1111_1011_1110.

[0152] The second mantissa with the same exponent is part11=0010_1010_1000_1011_1111_1111.

[0153] (5) Perform signed addition and subtraction of mantissa. The mantissa sign processing module 221 receives the first mantissa part00 and the second mantissa part11 with the same exponent as the first mantissa. Based on the sign bit of the larger operand din0, the first sign value of the first mantissa part00 is padded with a predetermined number of bits. Based on the sign bit of the smaller operand din1, the second sign value of the second mantissa part11 with the same exponent as the first mantissa is padded with a predetermined number of bits.

[0154] The sign of the first mantissa, part00, is set to 1. Part00 is inverted and incremented by 1, with a pre-ordered sign value (e.g., two 1s) added to the beginning. The sign of the second mantissa, part11, which has the same exponent as the first mantissa, is set to 0. The pre-ordered sign value (e.g., two 0s) is added to the beginning of part11, resulting in a 26-bit first and second mantissa. These 26-bit first and second mantissas are defined as part20 and part21, respectively.

[0155] The first digit of the 26-bit number, part20, is 11_0110_0110_0101_0000_0100_0010.

[0156] The second digit of the 26-bit number, part21, is 00_0010_1010_1000_1011_1111_1111.

[0157] (6) Perform addition on the 26-bit first mantissa part20 and the 26-bit second mantissa part21 output by the mantissa symbol processing module 221, and output the result of the addition operation as the calculation result at the output of the adder / subtractor.

[0158] The sum of the last digits, part30, is 11_1001_0000_1101_1100_0100_0001.

[0159] The result of the addition operation is the sum of the mantissa part3, which is negative (the highest bit is 1), representing a negative number. Its two's complement is then obtained (inverted and 1 is added).

[0160] The sum of the mantissas and their complements, part31, is: 00_0110_1111_0010_0011_1011_1111.

[0161] Save the highest bit of the sum of the mantissa part30 into the sign4 of the mantissa sum calculation result, so that the sign4 of the mantissa sum calculation result is 1.

[0162] (7) The 24th bit of the summation mantissa complement part31, part31

[24] =0, the exponent of the summation mantissa complement part31 is equal to the exponent of the larger operand din0, exp00=1000_0110, without carry, and normalization is performed when the calculation result produces leading zeros. The lower 24 bits of the summation mantissa complement part31 are defined as the 24-bit summation mantissa part4, and the exponent of the 24-bit summation mantissa part4 is exp4=1000_0110.

[0163] The 24-bit sum of the mantissas part40 is 0110_1111_0010_0011_1011_1111.

[0164] The 24-bit summation mantissa part40 is normalized as follows.

[0165] 1) Swap the high and low bits of the 24-bit sum mantissa part40, then take the complement of the sum mantissa part41 after the high and low bits swap, that is, take the inverse complement and add 1. Perform a bitwise AND operation on the sum mantissa part41 after the high and low bits swap and the sum mantissa part42 after taking the complement. The result of the bitwise AND operation is recorded as data.

[0166] The sum of the mantissas after swapping the high and low digits is part41 = 1111_1101_1100_0100_1111_0110.

[0167] The sum of the two's complement values, part42, is 0000_0010_0011_1011_0000_1010.

[0168] The result of the bitwise AND operation is data=0000_0000_0000_0000_0000_0010.

[0169] 2) Calculate the number of digits site=1 according to the following formulas (3) to (7) to determine the number of digits of the first digit 1 in the mantissa of the calculation result of the mantissa operation module 22.

[0170] site[0] = data[1] | data[3] | data[5] | data[7] | data[9] | data

[11] | data

[13] | data

[15] | data

[17] | data

[19] | data

[21] | data

[23] (3);

[0171] site[1] = data[2] | data[3] | data[6] | data[7] | data

[10] | data

[11] | data

[14] | data

[15] | data

[18] | data

[19] | data

[22] | data

[23] (4);

[0172] site[2] = data[4] | data[5] | data[6] | data[7] | data

[12] | data

[13] | data

[14] | data

[15] | data

[20] | data

[21] | data

[22] | data

[23] (5);

[0173] site[3] = data[8] | data[9] | data

[10] | data

[11] | data

[12] | data

[13] | data

[14] | data

[15] (6);

[0174] site[4] = data

[16] | data

[17] | data

[18] | data

[19] | data

[20] | data

[21] | data

[22] | data

[23] (7).

[0175] Where data[a] represents the a-th bit of the result data of the bitwise AND operation, a = 0, 1, ..., 23; site[b] represents the b-th bit of the number site, b = 0, 1, 2, 3, 4; "|" represents the bitwise OR operation.

[0176] (8) Shift the lower 24 bits of the 24-bit sum mantissa part40 left by 1 bit to obtain the standard floating-point mantissa part43. Simultaneously, subtract site from the exponent exp00=1000_0110 of the larger operand din0 to obtain the exponent exp4 of the standard floating-point mantissa part43.

[0177] The sign4 of the summation result of the mantissa, the exponent exp4 of the standard floating-point mantissa part43, and the lower 23 bits of the standard floating-point mantissa part43 are combined to form a 32-bit result, which is the calculation result of the M-level calculation circuit. This result is in standard floating-point format.

[0178] The calculation result for level M is douM=1100_0010_1101_1110_0100_0111_0111_1110.

[0179] Combination Figure 2 In the second and third level addition / subtraction units, there are multiple pairs of identical inputs for addition and subtraction operations. If a complete adder and subtractor are used to implement butterfly computation, each adder or subtractor will perform alignment processing. This results in redundant alignment processing for the same operands in different adders and subtractors, leading to resource waste.

[0180] The embodiments of this application also propose a method to fuse adders and subtractors with the same input, share the alignment processing logic, thereby eliminating redundant alignment processing and reducing the consumption of hardware resources.

[0181] Figure 8 This is a schematic diagram of the processing flow of the processing unit in an embodiment of this application.

[0182] like Figure 8As shown, in the first-level computing circuit, the two first-type adders / subtractors, each under the control of the first-level control signal, perform alignment and mantissa processing on the multiple operands received by each of the two first-type adders / subtractors in the first-level computing circuit. In the second-level computing circuit, the second-type adders / subtractors, under the control of the second-level control signal, perform alignment and mantissa processing on the multiple operands received by the second-type adders / subtractors in the second-level computing circuit, so as to simultaneously perform addition and subtraction operations on the two calculation results output by the first-level computing circuit. In the third-level computing circuit, the second-type adders / subtractors, under the control of the third-level control signal, perform alignment, mantissa processing, and normalization processing on the multiple operands received by the second-type adders / subtractors in the third-level computing circuit, so as to simultaneously perform addition and subtraction operations on the two calculation results output by the second-level computing circuit. The second type of adder / subtractor can perform normalization and adaptive alignment processing by configuring control signals. At the same time, by merging the alignment modules of two first type adders / subtractors together, it can perform addition and subtraction operations on the two operands input to the second type adder / subtractor. Therefore, the second type of adder / subtractor can also be called a fused adder / subtractor.

[0183] Figure 9 This is a schematic diagram of the adder / subtractor according to an embodiment of this application.

[0184] like Figure 9 As shown, the second type of adder / subtractor 11_2 includes an alignment module 21, a mantissa operation module 22, and a normalization module 23. The alignment module 21 of the second type of adder / subtractor 11_2 may include a mantissa completion module 211, a mantissa shifting module 212, a sign preprocessing module 213, and an operand preprocessing module 214. The sign preprocessing module 213 is connected to the second input terminal of the adder / subtractor. The operand preprocessing module 214 is connected to the first input terminal of the adder / subtractor and the sign preprocessing module 213.

[0185] According to an embodiment of this application, the sign preprocessing module 213 is used to output the second operand received from the second input terminal to the operand preprocessing module. The operand preprocessing module 214 is used to arrange the first operand and the second operand in order of exponent size and then provide them to the mantissa completion module.

[0186] According to an embodiment of this application, the mantissa operation module 22 of the second type of adder / subtractor 11_2 includes a mantissa sign processing module 221, a mantissa addition module 222, and a mantissa subtraction module 223. The mantissa sign processing module 221 is connected to the mantissa shifting module 212. The mantissa addition module 222 is connected to the mantissa sign processing module 221. The mantissa subtraction module 223 is connected to the mantissa sign processing module 221.

[0187] The mantissa sign processing module 221 receives the first mantissa and the second mantissa output by the mantissa shifting module 212, and pads the first mantissa with a predetermined number of first sign values ​​according to the sign of the first operand, and pads the second mantissa with a predetermined number of second sign values ​​according to the sign of the second operand. The mantissa addition module 222 performs addition on the first mantissa and the second mantissa output by the mantissa sign processing module 221, and outputs the result of the addition as the first calculation result. The mantissa subtraction module 223 performs subtraction on the first mantissa and the second mantissa output by the mantissa sign processing module 221, and outputs the result of the subtraction as the second calculation result.

[0188] According to the embodiments of this application, considering the characteristics of butterfly computation in Fast Fourier Transform, butterfly computation involves a large number of addition operations on the first operand and the second operand, as well as subtraction operations on the first operand and the second operand. The embodiments of this application also propose an adder / subtractor that is not controlled by addition and subtraction instructions, which can simultaneously perform addition and subtraction operations on the first operand and the second operand. The addition and subtraction operations share a set of alignment modules, eliminating redundant calculations in alignment operations, greatly simplifying the circuit structure, and improving computational efficiency.

[0189] The normalization module 23 of the second type of adder / subtractor 11_2 includes a first normalization module 231 and a second normalization module 232. The first normalization module 231 is connected to the mantissa addition module 222. The second normalization module 232 is connected to the mantissa subtraction module 223.

[0190] According to an embodiment of this application, a first normalization module 231 is used to perform carry-over or normalization processing on a first calculation result. A second normalization module 232 is used to perform carry-over or normalization processing on a second calculation result. Both the first normalization module 231 and the second normalization module 232 are connected to the normalization control terminal of the adder / subtractor, and are configured to enable normalization processing in response to the normalization control terminal receiving an M-th level control signal, and to disable normalization processing in response to the normalization control terminal receiving an m-th level control signal.

[0191] According to the embodiments of this application, for the second type of adder / subtractor that simultaneously outputs a first calculation result and a second calculation result, a first normalization module and a second normalization module are configured to perform carry-over or normalization processing on the first calculation result and the second calculation result respectively, which can improve the accuracy of the calculation result of the second type of adder / subtractor.

[0192] According to embodiments of this application, such as Figure 9As shown, the second type of adder / subtractor 11_2 also includes a delay module 24, a sign storage module 25, and an exponent storage module 26. The delay module 24 is used to synchronize the calculation result with the timing signal. The sign storage module 25 is used to store the sign bits of the first and second operands provided by the operand preprocessing module. The exponent storage module 26 is used to store the exponents of the first and second operands provided by the operand preprocessing module.

[0193] According to an embodiment of this application, the first-stage computing circuit includes at least one complex multiplier, and the complex multiplier includes two adders / subtractors. The i-th-stage computing circuit includes multiple complex adders / subtractors, and the complex adders / subtractors include two adders / subtractors, where i = 2, 3, ..., M.

[0194] According to embodiments of this application, the adder / subtractor in the first-stage computing circuit is a first-type adder / subtractor, and the adder / subtractor in the i-th-stage computing circuit is either a first-type adder / subtractor or a second-type adder / subtractor. The first-type adder / subtractor is configured to receive addition / subtraction instructions and perform addition or subtraction operations on the received first operand and second operand according to the addition / subtraction instructions; the second-type adder / subtractor is configured to perform addition and subtraction operations on the received first operand and second operand respectively.

[0195] According to an embodiment of this application, the complex multiplier in the first-stage computing circuit is configured with two first-type adders / subtractors. The second-stage computing circuit and subsequent computing circuits can employ either first-type or second-type adders / subtractors. The choice between first-type or second-type adders / subtractors can be flexibly made as needed.

[0196] The butterfly calculation in the second-level calculation circuit and subsequent calculation circuits includes addition of the first operand and the second operand, as well as subtraction of the first operand and the second operand. In the second-level calculation circuit and subsequent calculation circuits, the addition and subtraction operations of the first operand and the second operand are combined by a second type of adder / subtractor. The second type of adder / subtractor can simultaneously output the calculation result of the addition operation of the first operand and the calculation result of the subtraction operation of the second operand.

[0197] According to an embodiment of this application, the butterfly computation is a radix-4 butterfly computation, where M=3. The first-stage computation circuit includes three complex multipliers, each containing two first-type adders / subtractors. The i-th-stage computation circuit includes four complex adders / subtractors, each containing two first-type adders / subtractors; or, the i-th-stage computation circuit includes two complex adders / subtractors, each containing two second-type adders / subtractors.

[0198] Figure 10 This is a schematic diagram of the processing unit in an embodiment of this application.

[0199] like Figure 10 As shown, the processing unit in this embodiment includes a three-level computing circuit, which can be applied to... Figure 2 The diagram illustrates a radix-4 butterfly computation. The first-stage computation circuit includes three complex multipliers, each containing two simplified adders / subtractors. The second-stage computation circuit includes four complex adders / subtractors, each containing two simplified adders / subtractors. The third-stage computation circuit includes four complex adders / subtractors, each containing two simplified adders / subtractors.

[0200] The addition / subtraction instruction opt can take the value 0 or 1. When opt=0, the simplified adder / subtractor performs an addition operation; when opt=1, the simplified adder / subtractor performs a subtraction operation, thus enabling complex number addition or subtraction.

[0201] The first-stage computing circuit has three complex multipliers that receive the input data validity flag xvalid. Each complex multiplier includes two simplified adders and subtractors, performing addition and subtraction operations respectively. Combined with... Figure 2 The first complex multiplier of the first-stage computing circuit receives the calculation result of the third subsequence. The calculation results of the real part x2real, the imaginary part x2imag, and the third subsequence The real part w2real and the imaginary part w2imag of the twitch factor; the second complex multiplier of the first-stage computing circuit receives the computation results of the second set of subsequences. The real part x1real, the imaginary part x1imag, and the calculation results of the second subsequence The real part w1real and the imaginary part w1imag of the twitch factor; the third complex multiplier of the first-stage computing circuit receives the computation results of the fourth subsequence. The calculation results of the real part x3real, the imaginary part x3imag, and the fourth subsequence The real part w3real and the imaginary part w3imag of the twitch factor. The first-level computation circuit also includes a data buffer module, which receives the computation results of the first set of subsequences. The real part is x0real, the imaginary part is x0imag, and the data caching module includes a mantissa processing module.

[0202] The first and third complex number adders / subtractors of the second-level computing circuit include two simplified adders / subtractors that perform addition operations. The second and fourth complex number adders / subtractors of the second-level computing circuit include two simplified adders / subtractors that perform subtraction operations.

[0203] The third-level computational circuit includes four complex number adders and subtractors, each outputting a data validity flag (yvalid). Each complex number adder and subtractor includes two simplified adders and subtractors. The two simplified adders and subtractors performing addition operations in the first complex number adder and subtractor of the third-level computational circuit output the real part (y0real) and imaginary part (y0imag) of the first calculation result, respectively. The two simplified adders and subtractors performing addition operations in the second complex number adder and subtractor of the third-level computational circuit output the real part (y1real) and imaginary part (y1imag) of the second calculation result, respectively. The two simplified adders and subtractors performing subtraction operations in the third complex number adder and subtractor of the third-level computational circuit output the real part (y2real) and imaginary part (y2imag) of the third calculation result, respectively. The two simplified adders and subtractors performing subtraction operations in the fourth complex number adder and subtractor of the third-level computational circuit output the real part (y3real) and imaginary part (y3imag) of the fourth calculation result, respectively.

[0204] The control signal STEP can take values ​​of 0, 1, and 2. The first-level calculation circuit receives a first-level control signal STEP with a value of 0. Under the control of this value, the mantissa completion module in the first-level calculation circuit pads the first and second mantissas with a pre-ordered first value, and the normalization module in the first-level calculation circuit is disabled. The second-level calculation circuit receives a second-level control signal STEP with a value of 1. Under this value, the mantissa completion module in the second-level calculation circuit pads the last digit of the first and second mantissas with a pre-ordered second value, and the normalization module in the second-level calculation circuit is disabled. The third-level calculation circuit receives a third-level control signal STEP with a value of 2. Under this value, the mantissa completion module in the third-level calculation circuit pads the last digit of the first and second mantissas with a pre-ordered second value, and the normalization module in the third-level calculation circuit is enabled. By configuring the STEP control signal, the first two stages of the computing circuit are controlled not to execute normalization operation logic, while the last stage of the computing circuit is controlled to execute normalization operation.

[0205] Figure 11 This is a schematic diagram of the processing unit in an embodiment of this application.

[0206] like Figure 11 As shown, the processing unit in this embodiment includes a three-level computing circuit, which can be applied to... Figure 2 The diagram illustrates a radix-4 butterfly computation. The first-stage computation circuit includes three complex multipliers, each comprising two type-1 adders / subtractors. The second-stage computation circuit includes two complex adders / subtractors, each comprising two fused adders / subtractors. The third-stage computation circuit includes two complex adders / subtractors, each comprising two fused adders / subtractors.

[0207] The first complex number adder / subtractor of the second-level computing circuit includes two fused adder / subtractors that perform addition and subtraction operations. The second complex number adder / subtractor of the second-level computing circuit includes two fused adder / subtractors that perform addition and subtraction operations.

[0208] The third-level computational circuit includes four fused adders / subtractors, each outputting a valid data flag (yvalid). The first fused adder / subtractor of the first complex adder / subtractor of the third-level computational circuit outputs the real part (y0real) of the first computation result and the real part (y2real) of the third computation result. The second fused adder / subtractor of the first complex adder / subtractor of the third-level computational circuit outputs the imaginary part (y0imag) of the first computation result and the imaginary part (y2imag) of the third computation result. The first fused adder / subtractor of the second complex adder / subtractor of the third-level computational circuit outputs the real part (y1real) of the second computation result and the real part (y3real) of the fourth computation result. The second fused adder / subtractor of the second complex adder / subtractor of the third-level computational circuit outputs the imaginary part (y1imag) of the second computation result and the imaginary part (y3imag) of the fourth computation result.

[0209] According to an embodiment of this application, the butterfly computation is a radix-2 butterfly computation, where M=2. The first-stage computation circuit includes one complex multiplier, which includes two adders and subtractors of the first type. The i-th-stage computation circuit includes two complex adders and subtractors, which include two adders and subtractors of the first type; or, the i-th-stage computation circuit includes one complex adder and subtractor, which includes two adders and subtractors of the second type.

[0210] Figure 12 This is a schematic diagram of the processing unit in an embodiment of this application.

[0211] like Figure 12 As shown, the processing unit in this embodiment includes a two-stage computing circuit, which can be applied to... Figure 1 The diagram illustrates a radix-2 butterfly computation. The first-stage computation circuit includes one complex multiplier, which in turn includes two simplified adders / subtractors. The second-stage computation circuit includes two complex adders / subtractors, which in turn include two simplified adders / subtractors.

[0212] The complex multiplier in the first-stage computing circuit receives the input data validity flag xvalid. The complex multiplier in the first-stage computing circuit includes two simplified adders and subtractors, which perform addition and subtraction operations respectively. Combined with... Figure 2 The complex multiplier of the first-stage computing circuit receives the calculation results of the second set of subsequences. The real part x1real, the imaginary part x1imag, and the calculation results of the second subsequence The real part wreal and the imaginary part wmag of the twitch factor; the first-stage computation circuit also includes a data cache module, which receives the computation results of the first set of subsequences. The real part is x0real, the imaginary part is x0imag, and the data caching module includes a mantissa processing module.

[0213] The second-stage computational circuit includes two complex number adders / subtractors, each outputting a data validity flag (yvalid). Each complex number adder / subtractor comprises two simplified adders / subtractors. The two simplified adders / subtractors performing addition operations in the first complex number adder / subtractor of the second-stage computational circuit output the real part (y0real) and the imaginary part (y0imag) of the first calculation result, respectively. The two simplified adders / subtractors performing subtraction operations in the second complex number adder / subtractor of the second-stage computational circuit output the real part (y1real) and the imaginary part (y1imag) of the second calculation result, respectively.

[0214] The control signal STEP can take values ​​of 0 and 1. The first-level calculation circuit receives a first-level control signal STEP with a value of 0. Under the control of this value, the mantissa completion module in the first-level calculation circuit pads the first and second mantissas with a first value of a predetermined number of bits, and the normalization module in the first-level calculation circuit is disabled. The second-level calculation circuit receives a second-level control signal STEP with a value of 1. Under this value, the mantissa completion module in the second-level calculation circuit pads the last digit of the first and second mantissas with a second value of a predetermined number of bits, and the normalization module in the second-level calculation circuit is enabled. By configuring the control signal STEP, the first-level calculation circuit can be controlled to not perform the normalization operation logic, while the second-level calculation circuit can be controlled to perform the normalization operation.

[0215] Figure 13 This is a schematic diagram of the processing unit in an embodiment of this application.

[0216] like Figure 13 As shown, the processing unit in this embodiment includes a two-stage computing circuit, which can be applied to... Figure 1 The diagram illustrates a radix-2 butterfly computation. The first-stage computation circuit includes one complex multiplier, which in turn includes two simplified adders / subtractors. The second-stage computation circuit includes one complex adder / subtractor, which in turn includes two fused adders / subtractors.

[0217] The second-stage computing circuit includes two fused adders / subtractors, each outputting a data validity flag, yvalid. The first fused adder / subtractor of the second-stage computing circuit outputs the real part y0real of the first calculation result and the real part y1real of the second calculation result. The second fused adder / subtractor of the third-stage computing circuit outputs the imaginary part y0imag of the first calculation result and the imaginary part y1imag of the second calculation result.

[0218] The processing unit in this application embodiment can be applied to butterfly calculations of fast Fourier transforms with different bases, all of which can reduce the amount of computation, and can be flexibly selected between "more flexible" and "more hardware resource saving" as needed.

[0219] The simplified adder / subtractor of the processing unit in this application embodiment can also be applied to butterfly calculations of fast Fourier transforms in architectures such as 4-point radix-2, 8-point radix-2, and 16-point radix-4.

[0220] The butterfly calculation of the 4-point radix-2 Fast Fourier Transform represents the sampling of a digital discrete signal into 4 sampling points. Therefore, the signal length is... Divided into Level butterfly computing unit.

[0221] Figure 14 This is a schematic diagram of a 4-point radix-2 FFT butterfly calculation.

[0222] like Figure 14 As shown, the signal length is 4, and the entire network is divided into two levels of butterfly computation. Each level of butterfly computation includes two butterfly units with two inputs and two outputs. The time-domain sequence passes through the first-level butterfly computation and the second-level butterfly computation in sequence, and finally obtains the ordered computation results directly at the output. to .

[0223] In Level 1 butterfly computation, the twitch factor for odd-numbered branches Even-indexed subsequences in a time series , The two input terms of the first butterfly unit are combined with the rotation factor corresponding to that butterfly unit. After completing the addition and subtraction operations, we obtain the two output items of the first butterfly unit, which are the calculation results of the first subsequence. and Odd-indexed subsequences in a time series , As the two input terms of the second butterfly unit, combined with the rotation factor corresponding to that butterfly unit. After completing the addition and subtraction operations, the two output items of the second butterfly unit, which are the calculation results of the second subsequence, are obtained. and .

[0224] (8);

[0225] (9).

[0226] In the second-level butterfly calculation, , Following the radix-2 merging rule shown in formula (1), the odd-numbered branch data is multiplied by the corresponding rotation factor, and then added to and subtracted from the even-numbered branch results respectively, finally outputting the calculation result. , , , The odd-numbered branches of the first butterfly unit are combined with the twitch factor corresponding to the first butterfly unit. The addition and subtraction operations are performed with the even-numbered branches of the first butterfly unit to obtain the two output terms of the first butterfly unit, namely... , The odd-numbered branches of the second butterfly unit are combined with the twitch factor corresponding to the second butterfly unit. The addition and subtraction operations are performed with the even-numbered branches of the second butterfly unit to obtain the two output terms of the second butterfly unit, namely... , .

[0227] (10);

[0228] (11).

[0229] The butterfly calculation of the 8-point radix-2 Fast Fourier Transform represents the sampling of a digital discrete signal into 8 sampling points. Therefore, the signal length is... Divided into Level butterfly computing unit.

[0230] Figure 15 This is a schematic diagram of an 8-point radix-2 FFT butterfly calculation.

[0231] As shown in Figure 15, the signal length is 8, and the entire network is divided into three levels of butterfly computation. Each level of butterfly computation includes four two-input, two-output butterfly units. The original time-domain sequence is first rearranged by bit reversal, and then sequentially processed through the first-level butterfly computation, the second-level butterfly computation, and the third-level butterfly computation, finally obtaining the ordered computation result directly at the output. to .

[0232] Figure 15 The leftmost dashed box in the diagram indicates the bit-flipping module. Original input sequence to The subscripts are replaced using crosshairs, which means the sequence is reordered according to the binary bit reversal rule. The rearranged sequence is used as the input data for the first-level butterfly computation.

[0233] In the first-level butterfly calculation, the twitch factors corresponding to the four butterfly units include: The time-domain sequence after bit reversal is grouped by even and odd indices. The even-indexed subsequence is used as the input to the first two butterfly units, and the odd-indexed subsequence is used as the input to the last two butterfly units. Each butterfly unit is then combined with its corresponding rotation factor. Complete addition and subtraction operations, output four sets of primary intermediate calculation results, and complete the first level of butterfly calculation.

[0234] In the second-level butterfly calculation, the twist factors corresponding to the four butterfly units are respectively , Using the primary intermediate calculation results output from the first-level butterfly calculation as input, and following the radix-2 FFT merging rules, the odd-numbered branch data is multiplied by the corresponding rotation factor and then added or subtracted from the even-numbered branch data to complete the 4-point radix-2 butterfly merging operation, outputting four sets of secondary intermediate calculation results, thus realizing the step-by-step merging of the sequence dimension.

[0235] In the third-level butterfly calculation, the twiddle factors corresponding to the four butterfly units are respectively: , , , Using the secondary intermediate calculation results output from the second-level butterfly calculation as input, the data from each branch are matched with corresponding rotation factors to complete weighting, addition, and subtraction operations. The four sets of 4-point intermediate results are finally merged into a complete 8-point frequency domain sequence, and the ordered final frequency domain results are output sequentially. , , , , , , , .

[0236] The butterfly calculation of the 16-point radix-4 Fast Fourier Transform represents the sampling of a digital discrete signal into 16 sampling points. Therefore, the signal length is... Divided into Level butterfly computing unit.

[0237] Figure 16 This is a schematic diagram of a 16-point radix 4FFT butterfly calculation.

[0238] like Figure 16As shown, the signal length is 16. The entire network consists of two levels of butterfly computation, each level comprising four four-input, four-output butterfly units. The original time-domain sequence is first rearranged by bit reversal, and then grouped based on the modulo-4 decimation rule. Subsequently, it is processed sequentially through the first and second level butterfly computations, ultimately yielding an ordered computation result directly at the output. to .

[0239] Figure 16 The leftmost dashed box in the diagram indicates the bit-flipping module. Original input sequence to The subscript permutation is performed using crosshairs, which means reordering the sequence according to the 4-bit reversal rule. The rearranged sequence serves as the input data for the first-level butterfly computation.

[0240] The first group of subsequences includes those from the original time series. , , , The calculation result corresponding to the first subsequence .

[0241] The second group of subsequences includes those from the original time series. , , , The calculation result corresponding to the second subsequence .

[0242] The third subsequence includes the original time series sequence. , , , The calculation result corresponding to the third subsequence .

[0243] The fourth subsequence includes the original time series sequence. , , , The calculation result corresponding to the 4th subsequence .

[0244] In the first-level butterfly calculation, the twitch factor corresponding to the four butterfly units is: Each butterfly branch is marked with fixed weighting coefficients such as "-1" and "-j". Following the modulo-4 time-domain decimation rule, the calculation results of the four subsequences are sent to the four butterfly units respectively. Each butterfly unit performs weighting and merging operations on the four input signals in combination with the corresponding rotation factor, and outputs four sets of primary intermediate calculation results, completing the first-level butterfly calculation.

[0245] In the second-level butterfly calculation, the twiddle factors corresponding to the four butterfly units include: , , , , , , Each branch within the butterfly is marked with fixed weighting coefficients such as "-1" and "-j". Using the initial intermediate calculation results from the first-level butterfly calculation as input, and adapting to the merging rules of radix-4 butterfly operations, the data from each branch is weighted according to the corresponding rotation factors to achieve orthogonal merging of multiple data streams. By merging the four intermediate subsequences step-by-step through four sets of radix-4 butterfly units, the global 16-point frequency domain sequence is integrated, ultimately outputting a complete and ordered frequency domain calculation result. to .

[0246] As can be seen from the radix-2 and radix-4 butterfly computing networks described above, the core operations of FFT computation can be summarized into two basic types of computations: complex number addition and complex number multiplication. Each butterfly unit contains several addition and multiplication operations, requiring a corresponding number of adders and multipliers in the hardware implementation. Multipath parallel computing is an effective method to improve FFT computation speed, that is, to perform butterfly computation on data in parallel through multiple radix-4 butterfly computing networks. However, multipath parallel computing increases the hardware resource overhead of the FFT core, thereby increasing chip area costs. Therefore, reducing the resource consumption of the butterfly computing network is key to optimized design.

[0247] The processing unit in this embodiment can perform butterfly calculations of fast Fourier transform based on a relatively simple adder / subtractor, thereby reducing the hardware resource consumption of the butterfly calculation network.

[0248] Figure 17 This is a schematic diagram of the signal transformation and processing flow of Fast Fourier Transform.

[0249] Figure 18 This is a schematic diagram illustrating the decomposition principle of multidimensional FFT.

[0250] The core function of FFT is to convert the acquired discrete-time domain signal into a frequency domain signal, enabling various signal processing operations such as filtering, data compression, and signal multiplication in the frequency domain. After frequency domain processing, the inverse fast Fourier transform (inverse FFT) is used to restore the processed frequency domain signal back to the time domain signal. The complete signal transformation and processing flow is as follows: Figure 17As shown. Based on the computational dimension, FFT mainly includes one-dimensional FFT, two-dimensional FFT, and three-dimensional FFT, among which one-dimensional FFT is the most basic algorithm form, while two-dimensional and three-dimensional FFTs are collectively referred to as multi-dimensional FFT. Different dimensions of FFT are suitable for different application scenarios: one-dimensional FFT is widely used in various one-dimensional signal processing scenarios, two-dimensional FFT is mainly used in image processing, and three-dimensional FFT is mostly used for solving partial differential equations. The essence of multi-dimensional FFT operation is to decompose it into multiple one-dimensional FFT operations performed step by step. Its hierarchical decomposition principle is as follows: Figure 18 As shown, this also makes one-dimensional FFT the core foundation of the entire FFT algorithm system.

[0251] Figure 19 This is a schematic diagram of a first application scenario of the processing unit in an embodiment of this application.

[0252] like Figure 19 As shown, the processing unit in this embodiment can be applied to millimeter-wave radar signal processing. The analog signal is input to the radio frequency front end, converted into a digital signal by a high-speed analog-to-digital converter, and then sent to the signal processing unit. The signal processing unit first performs interference suppression, preprocesses the time-domain signal, and then sequentially performs range FFT and Doppler FFT operations to obtain the two-dimensional range-Doppler frequency domain results corresponding to multiple antennas.

[0253] The processing unit in this embodiment can be applied to butterfly calculations of FFT in distance and Doppler dimensions, saving hardware resources for Fourier transform operations and achieving high-speed signal processing while reducing hardware resource consumption.

[0254] After completing two levels of FFT, beamforming is performed, followed by a filtering and screening process, specifically constant false alarm rate (CFAR) filtering, to identify point targets. Feature extraction is then performed, and angle calculations are completed for the filtered point targets. Based on the extracted distance, velocity, and angle information, target clustering is performed. In real-world scenarios, a target with the same large reflective area will generate multiple reflection points. The clustering algorithm is used to group points belonging to the same real target, thus completing target detection. Target tracking then proceeds. Thanks to the high signal processing speed of the entire process, the distance, velocity, and angle of the real target do not undergo significant abrupt changes. The system relies on this characteristic to predict and track the target trajectory, and combines vehicle motion information to determine whether the target will intrude into the vehicle's trajectory. Finally, control signals are generated and output to the sensors, issuing decision commands to the vehicle's actuators to complete vehicle control functions such as alarms, braking, acceleration / deceleration, and steering.

[0255] Figure 20 This is a schematic diagram of a second application scenario of the processing unit in an embodiment of this application.

[0256] like Figure 20As shown, the processing unit of this application embodiment can be applied to an image processing workflow. A two-dimensional forward FFT calculation is performed on the image data, first an X-dimensional FFT calculation, then a Y-dimensional FFT calculation, thereby converting the image data from the two-dimensional spatial domain to the frequency domain; subsequently, frequency domain filtering is performed on the frequency domain image data; after the filtering operation is completed, a two-dimensional inverse FFT calculation is performed, and finally the processed image data is output.

[0257] The processing unit in this application embodiment can be applied to butterfly calculations of two-dimensional forward FFT and two-dimensional inverse FFT, which can save hardware resources for Fourier transform operations and effectively reduce the consumption of hardware resources in the FFT operation process while realizing high-speed image data processing.

[0258] Figure 21 This is a schematic diagram of a third application scenario of the processing unit in an embodiment of this application.

[0259] like Figure 21 As shown, the processing unit in this embodiment can be applied to the Particle Mesh Ewald (PME) algorithm in molecular dynamics simulation. After inputting atomic charge, position, and other information, B-spline interpolation is performed to interpolate the randomly distributed charges in the molecular system to a three-dimensional mesh, generating an electrostatic potential energy matrix and a coefficient matrix. Subsequently, a three-dimensional forward FFT calculation is performed on the electrostatic potential energy matrix, sequentially performing X-dimensional FFT, Y-dimensional FFT, and Z-dimensional FFT operations to obtain the Fourier transform matrix. Matrix multiplication is then performed, multiplying the corresponding elements of the Fourier transform matrix and the coefficient matrix to obtain the initial inverse Fourier transform matrix. A three-dimensional inverse FFT calculation is then performed on the initial inverse Fourier transform matrix to output the electrostatic force matrix. Finally, the electrostatic force matrix is ​​processed again by B-spline interpolation to obtain the force situation of each atom in the original simulated molecular system.

[0260] The processing unit in this application embodiment can be applied to butterfly calculations of three-dimensional forward FFT and three-dimensional inverse FFT. It can save hardware resources for Fourier transform operations and effectively reduce the consumption of hardware resources in the FFT operation while achieving high-speed data processing.

[0261] At least one of the modules shown in the embodiments of this application can be implemented at least partially as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or any other reasonable means of integrating or packaging circuits, or implemented in hardware or firmware, or in any one of the three implementation methods of software, hardware, and firmware, or in a suitable combination of any of them.

[0262] Those skilled in the art will further recognize that the modules, units, and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0263] The processing unit provided in this application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A processing unit for performing butterfly calculations in Fast Fourier Transform, characterized in that, The processing unit includes: An M-level computing circuit, wherein multiple output terminals of the m-th level computing circuit are cross-connected with multiple input terminals of the (m+1)-th level computing circuit to form a butterfly network, where M is an integer greater than 1, and m = 1, 2, …, M-1; in, The computing circuit also has a control terminal, wherein the control terminal of the m-th level computing circuit is configured to receive the m-th level control signal, and the control terminal of the M-th level computing circuit is configured to receive the M-th level control signal; The computing circuit is configured as follows: In response to the control terminal of the computing circuit receiving the Mth level control signal, the computing circuit performs calculations on the multiple operands received at the multiple input terminals of the computing circuit, obtains multiple calculation results, normalizes the multiple calculation results, and outputs the normalized multiple calculation results at the multiple output terminals of the computing circuit. In response to the m-th level control signal received at the control terminal of the computing circuit, the circuit performs calculations on multiple operands received at multiple input terminals of the computing circuit, obtains multiple calculation results, and outputs the multiple calculation results at multiple output terminals of the computing circuit.

2. The processing unit according to claim 1, characterized in that, The calculation circuit includes multiple adders and subtractors. The first and second input terminals of the adders and subtractors serve as the input terminals of the calculation circuit, and the output terminals of the adders and subtractors serve as the output terminals of the calculation circuit. The adder / subtractor also has a normalization control terminal, which is connected to the control terminal of the calculation circuit. The adder / subtractor is configured as follows: In response to receiving the Mth level control signal at the normalization control terminal of the adder / subtractor, the adder / subtractor performs addition and / or subtraction operations on the operands received at the first and second input terminals of the adder / subtractor to obtain the calculation result, normalizes the calculation result, and outputs the normalized calculation result at the output terminal of the adder / subtractor. In response to receiving the m-th level control signal at the normalization control terminal of the adder / subtractor, the adder / subtractor performs addition and / or subtraction operations on the operands received at the first and second input terminals of the adder / subtractor to obtain the calculation result, and outputs the calculation result at the output terminal of the adder / subtractor.

3. The processing unit according to claim 2, characterized in that, Adders and subtractors include: The exponent alignment module connects the first and second input terminals of the adder / subtractor and is used to perform exponent alignment processing on the first operand received at the first input terminal and the second operand received at the second input terminal, so that the exponent of the first operand is equal to the exponent of the second operand. The mantissa operation module, connected to the alignment module, is used to perform addition and / or subtraction operations on the mantissas of the first and second operands output by the alignment module to obtain the calculation result; The normalization module, connected to the mantissa operation module, is used to perform carry or normalization processing on the calculation results of the mantissa operation module. Specifically, when the calculation result generates a carry to the integer part, the carry is performed, and when the calculation result generates leading zeros, the normalization processing is performed. The normalization module is connected to the normalization control terminal of the adder / subtractor. The normalization module is configured to enable normalization processing in response to the normalization control terminal receiving the M-th level control signal, and to disable normalization processing in response to the normalization control terminal receiving the m-th level control signal.

4. The processing unit according to claim 3, characterized in that, The alignment module includes: A mantissa padding module, connected to the normalization control terminal of the adder / subtractor, is configured to: in response to the normalization control terminal receiving a first-level control signal, extract the mantissa of a predetermined number of bits from the first and second operands to obtain a first mantissa and a second mantissa, and pad the first and second mantissas with a first value of a predetermined number of bits; in response to the normalization control terminal receiving any one of the second to Mth-level control signals, extract the mantissa of a predetermined number of bits from the first and second preprocessing operands to obtain a first mantissa and a second mantissa, and pad the last digit of the first and second mantissas with a second value of a predetermined number of bits; and The mantissa shifting module, connected to the mantissa completion module, is used to calculate the difference in exponents between the first mantissa and the second mantissa output by the mantissa completion module as the shift number, and shift one of the first mantissa and the second mantissa by the shift number to obtain the first mantissa and the second mantissa with the same exponent.

5. The processing unit according to claim 4, characterized in that, The alignment module also includes: The sign preprocessing module is connected to the second input terminal of the adder / subtractor and is configured to receive addition / subtraction instructions. It performs an XOR operation between the sign bit of the second operand and the addition / subtraction instruction to obtain the second operand with addition / subtraction operator signs. The operand preprocessing module connects the first input terminal of the adder / subtractor and the sign preprocessing module. It is used to arrange the first operand and the second operand with addition / subtraction operators in order of exponent size and then provide them to the mantissa completion module.

6. The processing unit according to claim 5, characterized in that, The mantissa calculation module includes: The mantissa sign processing module, connected to the mantissa shifting module, is used to receive the first mantissa and the second mantissa output by the mantissa shifting module, and to pad the first mantissa with a first sign value of a predetermined number of bits according to the sign of the first operand, and to pad the second mantissa with a second sign value of a predetermined number of bits according to the sign of the second operand. The mantissa addition module, connected to the mantissa sign processing module, is used to perform addition operations on the first and second mantissas output by the mantissa sign processing module, and outputs the result of the addition operation as the calculation result at the output terminal of the adder / subtractor.

7. The processing unit according to claim 4, characterized in that, The alignment module also includes: The sign preprocessing module is connected to the second input terminal of the adder / subtractor and is used to output the second operand received at the second input terminal to the operand preprocessing module. The operand preprocessing module connects the first input terminal of the adder / subtractor and the sign preprocessing module. It is used to arrange the first operand and the second operand according to the exponent size and then provide them to the mantissa completion module.

8. The processing unit according to claim 7, characterized in that, The mantissa calculation module includes: The mantissa sign processing module, connected to the mantissa shifting module, is used to receive the first mantissa and the second mantissa output by the mantissa shifting module, and to pad the first mantissa with a first sign value of a predetermined number of bits according to the sign of the first operand, and to pad the second mantissa with a second sign value of a predetermined number of bits according to the sign of the second operand. The mantissa addition module, connected to the mantissa sign processing module, is used to perform addition operations on the first and second mantissas output by the mantissa sign processing module, and output the result of the addition operation as the first calculation result. The mantissa subtraction module, connected to the mantissa sign processing module, is used to perform subtraction on the first and second mantissas output by the mantissa sign processing module, and output the result of the subtraction as the second calculation result.

9. The processing unit according to claim 8, characterized in that, The standardization module includes: The first normalization module, connected to the mantissa addition module, is used to perform carry-over or normalization processing on the first calculation result. The second normalization module, connected to the mantissa subtraction module, is used to perform carry-over or normalization processing on the second calculation result. The first normalization module and the second normalization module are both connected to the normalization control terminal of the adder / subtractor, and are configured to enable normalization processing in response to the normalization control terminal receiving the M-th level control signal, and to disable normalization processing in response to the normalization control terminal receiving the m-th level control signal.

10. The processing unit according to claim 5 or 7, characterized in that, Adders and subtractors also include: The delay module is used to synchronize the calculation results with the timing signal; The sign storage module is used to store the sign bits of the first and second operands provided by the operand preprocessing module; The exponent storage module is used to store the exponents of the first and second operands provided by the operand preprocessing module.

11. The processing unit according to claim 1, characterized in that, The first-level computing circuit includes at least one complex multiplier, and the complex multiplier includes two adders and subtractors; The i-th stage computing circuit includes multiple complex number adders and subtractors, and each complex number adder and subtractor includes two adders and subtractors, where i = 2, 3, …, M.

12. The processing unit according to claim 11, characterized in that, The adders and subtractors in the first-level computing circuit are of the first type, and the adders and subtractors in the i-th-level computing circuit are either of the first type or the second type. The first type of adder / subtractor is configured to receive addition / subtraction instructions and perform addition or subtraction operations on the received first operand and second operand according to the addition / subtraction instructions; the second type of adder / subtractor is configured to perform addition and subtraction operations on the received first operand and second operand respectively.

13. The processing unit according to claim 12, characterized in that, The butterfly calculation is a radix-4 butterfly calculation, where M=3; The first-level computing circuit includes three complex multipliers, each of which includes two type I adders and subtractors. The i-th stage computing circuit includes four complex number adders and subtractors, two of which are of the first type; or, the i-th stage computing circuit includes two complex number adders and subtractors, two of which are of the second type.

14. The processing unit according to claim 12, characterized in that, The butterfly calculation is a radix-2 butterfly calculation, where M=2; The first-level computing circuit includes one complex multiplier, which includes two adders and subtractors of type 1. The i-th stage computing circuit includes two complex number adders and subtractors, each of which is a first-type adder and subtractor; or, the i-th stage computing circuit includes one complex number adder and subtractor, which is a second-type adder and subtractor.

15. The processing unit according to claim 1, characterized in that, The multiple operands received by the multiple input terminals of the computing circuit are floating-point operands, which include a sign bit, an exponent bit, and a mantissa bit.