FPGA neural network accelerator for pipelined adc digital calibration
Patent Information
- Application Number
- CN202611003326.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-29
AI Technical Summary
但将神经网络应用于流水线ADC的实时在线校准仍面临诸多挑战:首先,神经网络计算量庞大且以乘累加操作为主,若直接在FPGA上部署,将大量占用DSP乘法器资源,或被迫将乘法运算映射至查找表(LUT),从而导致资源分配不均或资源消耗剧增;其次,网络参数及中间激活值的存储与频繁访问对带宽和块内存(BRAM)容量提出了较高要求,容易形成性能瓶颈并推高功耗;再者,神经网络各层对并行度的需求存在显著差异,使得采用固定拓扑结构的硬件阵列在不同计算阶段难以维持高效利用率,静态阵列往往造成资源闲置或吞吐能力受限;最后,在高采样率与严格的延迟约束条件下,还需在定点化精度、时序对齐与硬件映射之间寻求合理折衷,以防止量化误差削弱校准效果或引入额外处理延迟
[0011]本发明的二维并行乘累加阵列支持按功能周期动态选择活跃子PE阵列,能够在信息聚合、并行变换与残差合并、行级聚合等不同阶段灵活调整参与计算的行数与列数,从而在各计算阶段匹配神经网络层的并行度需求,有效减少空闲PE数量并提高资源利用率。
Smart Images

Figure CN122840141A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of programmable logic and digital signal processing, and more particularly to an FPGA neural network accelerator for pipelined ADC digital calibration. Background Technology
[0002] Pipeline ADCs, with their advantages of high bandwidth and high sampling rate, are widely used in communication, radar, and high-performance sampling systems. However, in actual conversion processes, these ADCs are often affected by factors such as gain mismatch, bias error, and nonlinear distortion. Therefore, digital calibration techniques are needed to improve the system's linearity and signal-to-noise ratio. In recent years, neural networks, with their superior nonlinear modeling capabilities, have been gradually introduced into the field of digital calibration to achieve higher calibration accuracy and stronger adaptability. However, applying neural networks to real-time online calibration of pipelined ADCs still faces many challenges: First, neural networks are computationally intensive and mainly involve multiply-accumulate operations. If deployed directly on an FPGA, they will consume a large amount of DSP multiplier resources or force multiplication operations to be mapped to lookup tables (LUTs), resulting in uneven resource allocation or a surge in resource consumption. Second, the storage and frequent access of network parameters and intermediate activation values place high demands on bandwidth and block RAM (BRAM) capacity, easily creating performance bottlenecks and increasing power consumption. Third, the parallelism requirements of different layers of a neural network vary significantly, making it difficult for hardware arrays with fixed topologies to maintain high utilization rates at different computation stages. Static arrays often result in idle resources or limited throughput. Finally, under high sampling rates and strict latency constraints, a reasonable trade-off must be sought between fixed-point accuracy, timing alignment, and hardware mapping to prevent quantization errors from weakening the calibration effect or introducing additional processing delays.
[0003] Therefore, there is an urgent need to design a dedicated hardware acceleration solution for pipelined ADC digital calibration, which can fully leverage the advantages of neural networks in calibration accuracy, and achieve high throughput and low latency online processing under the conditions of limited FPGA resources and meeting real-time requirements. Summary of the Invention
[0004] To address the aforementioned problems, this invention provides an FPGA neural network accelerator for pipelined ADC digital calibration, comprising a top-level interface module, a control and synchronization logic module, and a neural network calibration module. Upon receiving a start pulse, the top-level interface module transmits it to the control and synchronization logic module to trigger the neural network calibration module to perform digital calibration operations on the raw code of each node. The neural network calibration module includes:
[0005] The feature extraction unit is used to temporarily store the original encoding of multiple consecutive sampling times of the same node to form a short-time history sequence, and generate time-related features and spatial-related features based on the short-time history sequence. The time-related features and spatial-related features are combined into a fixed-width fixed-point vector to obtain the feature matrix.
[0006] The preprocessing and data flow control unit is used to send the elements of the normalized adjacency coefficient matrix and the elements of the feature matrix into the reconfigurable array unit according to the oblique in-time sequence under the coordination of the streaming control signal to perform matrix multiplication operations and output the support matrix after information aggregation.
[0007] The parallel transformation and residual merging unit is used to apply a linear transformation to the support matrix based on the reconfigurable array unit and generate residual branch results in parallel. The linear transformation results and residual branch results are then added element by element and the stage output is obtained after nonlinear activation.
[0008] Row-level aggregation unit is used to perform row-level aggregation operations on the stage output based on the reconfigurable array unit to obtain the final scalar output;
[0009] The reconfigurable array unit is used to configure the number of active rows and columns of the two-dimensional parallel multiply-accumulate array according to the dimension of the input matrix and the dimension of the weight matrix corresponding to the current computing task, and to control the timing of weight loading and data flow.
[0010] The beneficial effects of this invention are:
[0011] The two-dimensional parallel multiply-accumulate array of the present invention supports dynamic selection of active sub-PE arrays according to functional cycle. It can flexibly adjust the number of rows and columns involved in the calculation at different stages such as information aggregation, parallel transformation and residual merging, and row-level aggregation, thereby matching the parallelism requirements of neural network layers at each calculation stage, effectively reducing the number of idle PEs and improving resource utilization.
[0012] This architecture uses a diagonal injection method to synchronize multiple PEs on the diagonal, and a flow controller maintains the global time pointer and the preload of the spare buffer, thereby parallelizing a large number of inner product operations at the hardware level. This structure significantly improves the utilization of the DSP / multiply-accumulate unit, reduces overall latency, and increases system throughput.
[0013] To conserve hardware resources, this invention replaces some additional multiplication operations with row-level modulation and shift weighting in the row-level aggregation unit, and uses shifting and accumulation to achieve inter-row weight adjustment. This method achieves controllable row-level weighting with extremely low logic and routing overhead, significantly reducing dependence on multiplier resources, and is particularly suitable for resource-constrained FPGA platforms.
[0014] This invention can be adapted to various pipelined ADCs with different sampling rates and supports parameterization of arrays of different sizes. It eliminates the need to adjust the analog front end, greatly simplifying the complexity of engineering integration.
[0015] In summary, the synergistic effect of the reconfigurable PE array, oblique injection flow control, ping-pong buffered parallel prefetching, and shift weighting constitutes the beneficial effects of this invention in terms of performance improvement, resource saving, and engineering practicality. Attached Figure Description
[0016] Figure 1 This is a block diagram of the overall architecture of an FPGA neural network accelerator for pipelined ADC digital calibration according to the present invention.
[0017] Figure 2 This is a schematic diagram of the two-dimensional parallel multiply-accumulate array structure of the present invention;
[0018] Figure 3 This is a schematic diagram of the reconfigurable subarray selection of the two-dimensional parallel multiply-accumulate array of the present invention;
[0019] Figure 4 This is a comparison diagram of accelerator resources in an embodiment of the present invention;
[0020] Figure 5 This is a comparison of the spectrum of the commercial pipeline ADC (sampling rate 2.4G) before and after low-frequency / high-frequency calibration. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] This invention utilizes an FPGA as the target implementation platform to provide an FPGA neural network accelerator for pipelined ADC digital calibration. The system is implemented within the FPGA using RTL-level hardware description and mainly includes a top-level interface module, a control and synchronization logic module, and a neural network calibration module. Upon receiving a start pulse, the top-level interface module transmits it to the control and synchronization logic module, which then triggers the neural network calibration module to perform digital calibration operations on the raw code of each node.
[0023] In a preferred embodiment, the top-level interface module is provided with a parallel input bus, a system clock input terminal, an asynchronous reset input terminal, a start signal input terminal, a software clear signal input terminal, a busy flag output terminal, and a completion pulse output terminal.
[0024] A parallel input bus is used to receive the raw code from each node of the pipelined ADC output. To clarify the technical solution of this invention, the relevant terms are defined as follows:
[0025] A node refers to a sampling channel in a pipelined ADC. Each node generates a raw code at the effective edge of each sampling clock cycle; this raw code is the digital quantization value at that sampling moment. The system may contain one or more such nodes, and the FPGA neural network accelerator proposed in this invention performs digital calibration processing on at least one of these nodes.
[0026] The system clock input terminal is used to receive the system clock signal; the system clock signal provides a data sampling clock for the parallel input bus.
[0027] The asynchronous reset input terminal is used to receive the asynchronous reset signal;
[0028] The start signal input terminal is used to receive the start pulse (i.e., start signal) issued by the upper-level controller.
[0029] The software reset signal input terminal is used to receive the software reset pulse (i.e., software reset signal) issued by the upper-level controller.
[0030] The busy flag output terminal is used to output a busy flag signal to the upper-level controller to indicate whether the FPGA neural network accelerator is currently in operation;
[0031] The completion pulse output terminal is used to output a single-cycle completion pulse signal to the upper-level controller at the end of a complete digital calibration operation.
[0032] In this process, the original code for each node is output via the parallel input bus at the effective edge of each sampling clock and synchronously latched by the input register inside the FPGA neural network accelerator at the effective edge of the system clock. The asynchronous reset signal has higher priority than the system clock signal and forces the input register to remain cleared during the assertion period. After responding to the start signal, the control and synchronization logic module sets the busy flag signal to active and sets it to inactive after completing the calibration data processing of a preset length, and outputs a completion pulse signal. After responding to the software clear signal, the control and synchronization logic module immediately suspends the current processing and returns to the idle state.
[0033] In a preferred embodiment, the control and synchronization logic module is the global control core of the neural network calibration module. The control and synchronization logic module may include a main state machine, a clock counter, an enable signal generator, synchronization handshake logic, and a busy / complete flag generator. The main state machine manages the global operating state, including at least six states: idle, start, information aggregation, parallel transformation and residual merging, row-level aggregation, and complete. The clock counter generates clock cycle counts for the corresponding stage based on the current state (e.g., N cycles for the information aggregation stage, M cycles for the parallel transformation and residual merging stage, and L cycles for the row-level aggregation stage). The enable signal generator outputs clock enable, clear, and load control signals to the feature extraction unit, preprocessing and data flow control unit, reconfigurable array unit, row-level aggregation unit, and parallel transformation and residual merging unit, respectively, based on the current state and the clock count. The synchronization handshake logic processes the valid / ready handshake protocol at the data input / output ports of each module to prevent data overflow. The busy / complete flag generator generates a completion pulse when the global state enters the complete state and feeds it back to the top-level interface module.
[0034] Figure 1 This is a block diagram of the overall architecture of an FPGA neural network accelerator for pipelined ADC digital calibration according to the present invention.
[0035] In a preferred embodiment, such as Figure 1 As shown, the neural network calibration module includes a feature extraction unit, a reconfigurable array unit, a preprocessing and data flow control unit, a parallel transformation and residual merging unit, and a row-level aggregation unit.
[0036] The feature extraction unit is connected to the preprocessing and data flow control unit. It includes a short-term history register, which is used to temporarily store the original encoding of multiple consecutive sampling times of the same node to form a short-term history sequence. Based on the short-term history sequence, time-related features and spatial-related features are generated, and the time-related features and spatial-related features are combined into a fixed-width fixed-point vector to obtain the feature matrix.
[0037] The temporal correlation feature is used to characterize the dynamic changes of the same node at different times. Specifically, it can be composed of the original encoding of the same node at the current sampling time and the original encoding at the previous sampling time. It can also be extended to the original encoding of the same node at multiple consecutive sampling times stored in the short-term history register, thus reflecting the changing trend of the node in the time dimension. The spatial correlation feature is used to characterize the correlation between adjacent nodes or adjacent levels at the same sampling time. Specifically, it can be composed of the original encoding of the i-th level of the pipeline ADC at the same sampling time and the original encoding of the adjacent levels (i-1 and / or i+1 levels), thus reflecting the local consistency of the MDAC output of the sub-level ADCs within the pipeline ADC in the spatial neighborhood. Subsequently, the feature extraction unit concatenates the temporal correlation feature and the spatial correlation feature in a preset order to obtain the feature matrix.
[0038] The preprocessing and data flow control unit, connected to the parallel transformation and residual merging unit and the reconfigurable array unit, is used to feed the normalized adjacency coefficient matrix and the characteristic matrix into the reconfigurable array unit in a skew-in timing sequence under the coordination of streaming control signals to perform matrix multiplication operations, and output the aggregated support matrix. The parallel transformation and residual merging unit, connected to the row-level aggregation unit and the reconfigurable array unit, is used to apply a linear transformation to the support matrix and generate residual branch results in parallel. The linear transformation result and the residual branch result are then added element-wise, and after nonlinear activation, a stage output is obtained.
[0039] The row-level aggregation unit, connected to the reconfigurable array unit, is used to perform row-level aggregation operations on the stage output to obtain the final scalar output.
[0040] The reconfigurable array unit is used to configure the number of active rows and columns of the two-dimensional parallel multiply-accumulate array according to the dimension of the input matrix and the dimension of the weight matrix corresponding to the current computing task, and to control the timing of weight loading and data flow.
[0041] In a preferred embodiment, the normalized adjacency coefficients used by the preprocessing and data flow control unit are fixed-width, fixed-point representations obtained by sequentially performing mathematical normalization and fixed-point scaling on the original adjacency coefficients. The specific process for obtaining the normalized weight coefficients includes:
[0042] Step A: Determine the original adjacency coefficients based on the internal circuit connection topology definition of the pipeline ADC, and perform mathematical normalization on the original adjacency coefficients so that the normalized original adjacency coefficients fall within the preset value range.
[0043] Step B involves performing fixed-point scaling on the normalized original adjacency coefficients, i.e., multiplying them by a preset scaling factor, and rounding or truncating the product to obtain fixed-point adjacency coefficients in a fixed-width, fixed-point format.
[0044] Step C involves using the fixed-point adjacency coefficients as coefficients for matrix multiplication operations in the preprocessing and data flow control unit. To match the input bit width of the two-dimensional parallel multiply-accumulate array, the fixed-point adjacency coefficients are further sign-extended or zero-extended to obtain normalized adjacency coefficients, thereby avoiding bit width truncation errors during multiply-accumulate. Subsequently, the normalized adjacency coefficients are stored in the weight storage unit inside the FPGA neural network accelerator.
[0045] In a preferred embodiment, the parallel transformation and residual merging unit includes:
[0046] Local memory is used to store or load weight coefficients and biases to support vector-wise or element-wise linear transformations of the support matrix.
[0047] A linear transformation path connects the reconfigurable array unit and the preprocessing and data flow control unit, and is implemented based on a two-dimensional parallel multiply-accumulate array of the reconfigurable array unit. The linear transformation path receives the support matrix output from the preprocessing and data flow control unit and the weight matrix output from the weight memory, calls the reconfigurable array unit to perform parallel matrix multiplication on the support matrix and its corresponding weight matrix, and superimposes the results with corresponding bias terms to obtain the linear transformation result.
[0048] The residual path, independent of the linear transformation path, connects the reconfigurable array unit and the feature extraction unit, and is implemented based on a two-dimensional parallel multiply-accumulate array of the reconfigurable array unit. The residual path is used to receive the feature matrix output by the feature extraction unit and call the reconfigurable array unit to perform a projection transformation on the feature matrix to obtain a residual branch result that is aligned with the linear transformation result in both dimension and temporal order.
[0049] An element-wise additive unit array, connecting the linear transformation path and the residual path, is used to add the linear transformation result and the residual branch result element-wise, and then obtain the stage output after nonlinear activation. The stage output is organized into multiple row vectors, with different row vectors corresponding to different row weights.
[0050] The element-wise addition unit array is used to add the linear transformation result and the residual branch result element by element. When the bit width, integer bit configuration, sign attribute or numerical range of the two inputs are inconsistent with the target operation bit width of the element-wise addition unit array, or when there is a risk of overflow, bit width expansion and / or sign expansion are performed on the corresponding inputs to avoid truncation error.
[0051] In a preferred embodiment, the row-level aggregation unit includes:
[0052] The part and generation unit connects the reconfigurable array unit and the parallel transformation and residual merging unit. It is used to receive the stage output of the parallel transformation and residual merging unit and call the reconfigurable array unit to perform multiply-accumulate operation on the stage output to generate the first intermediate accumulated value and the second intermediate accumulated value corresponding to each row vector in the stage output.
[0053] The inline modulation unit, connecting the part and unit and the parallel transform and residual merging unit, is used to multiply the original encoded scalar of each row vector in the output of the generation stage by its first intermediate accumulated value, and add the product result to its second intermediate accumulated value to obtain the row-level output;
[0054] The aggregation unit, connected to the inline modulation unit, is used to left-shift and weight each row output according to a predetermined displacement, sum all the left-shift weighted results, and then perform an arithmetic right shift and bit width truncation to obtain the final scalar output.
[0055] In this way, the row-level aggregation process is directly built on a reconfigurable two-dimensional parallel multiply-accumulate array, thereby realizing differentiated weighted aggregation of different output row vectors.
[0056] In a preferred embodiment, the reconfigurable array unit includes a reconfigurable controller, a two-dimensional parallel multiply-accumulate array, a flow controller, and a weight configurator, wherein:
[0057] The reconfigurable controller is used to generate configuration signals based on the matrix dimension, computational precision, and weight parameters corresponding to the current computation task; the configuration signals include at least array size configuration signals, data bit width configuration signals, and weight loading control signals.
[0058] The matrix dimension corresponding to the current computation task refers to the dimension of the input vector matrix and the weight matrix multiplied with it. It is used to determine the number of rows and columns involved in the computation in the two-dimensional parallel multiply-accumulate array, thereby generating the array size configuration signal.
[0059] The computational precision refers to the fixed-point representation of the input vector, weighting coefficients, and multiplication-accumulation results, including the total bit width, integer bit width, and right shift truncation rules, which are used to generate the data bit width configuration signal for configuring the data path bit width.
[0060] Weight parameters refer to information used to identify the weight coefficients corresponding to the current layer and the group or base address biased in the weight storage unit, which is used to generate weight loading control signals to control weight loading.
[0061] The weight configurator, connected to the reconfigurable controller, is used to read the weight coefficients and biases of the corresponding group from the weight storage unit according to the weight loading control signal, and output them to the two-dimensional parallel multiply-accumulate array.
[0062] The flow controller, connected to the reconfigurable controller, is used to configure signals based on array size and data bit width, and inject the input vector and weight coefficients output by the weight configurator into the two-dimensional parallel multiply-accumulate array periodically according to the slant-in timing sequence.
[0063] The two-dimensional parallel multiply-accumulate array is connected to a reconfigurable controller, a weight configurator, and a flow controller. It consists of multiple processing elements (PEs) arranged in rows and columns, and the number of active rows and columns is dynamically set according to the array size configuration signal. Each processing element is configured to receive a weight coefficient and an input vector, perform multiply-accumulate operations, and pass the weight coefficient or input vector along the array direction to adjacent processing elements. Thus, under the scheduling of the reconfigurable controller, matrix operations of different sizes, bit widths, and weight configurations are completed, and the combined path delay is reduced.
[0064] Specifically, the two-dimensional parallel multiply-accumulate array is a clock-synchronized pipelined array. Each PE (Programmer) includes a configurable register stage on the data path, comprising at least an input register to store one of the following: an input register to store external row or column inputs, a pre-multiplication register, a post-multiplication register, and a post-addition register, to reduce combination path latency and increase the maximum operating frequency. Each PE also includes a multiplication register, a partial sum register, and an accumulation register, and in response to a clear_acc signal, the accumulation register is cleared at the start of each computation. After completing internal accumulation, the two-dimensional parallel multiply-accumulate array adds the corresponding column bias to the accumulation result and can perform optional right-shift scaling to match the bit width requirements of the fixed-point format, before outputting via clock register.
[0065] Furthermore, the reconfigurable array unit further includes:
[0066] Multiple input buffers are configured, each including a weight input buffer and a feature input buffer. The weight input buffer is used to temporarily store the weight coefficients read from the weight storage unit, and the feature input buffer is used to temporarily store the input vector. The multiple input buffers work alternately in a ping-pong manner, and under the coordination of the flow controller, the corresponding data is sent into the two-dimensional parallel multiply-accumulate array in an oblique in-order manner, so that the weight coefficients and input vectors are correctly paired in the internal processing unit of the two-dimensional parallel multiply-accumulate array and the matrix multiplication operation is completed. At any given time, only one input buffer is effective and drives the input of the two-dimensional parallel multiply-accumulate array, while the other input buffer groups prefetch the next batch of weight coefficients and input vectors in parallel.
[0067] Figure 2This is a schematic diagram of a two-dimensional parallel multiply-accumulate array structure within a reconfigurable array cell provided in some embodiments of the present invention. The two-dimensional parallel multiply-accumulate array is a matrix multiplier that operates in a systolic manner. It consists of multiple PEs arranged in rows and columns. Each PE performs multiply-accumulate operations and passes the weight coefficients or input vectors along the array direction to adjacent PEs to achieve parallel computation of matrix multiplication.
[0068] exist Figure 2 The system employs two alternating buffer queues: a first weight buffer `buf_a1`, a first feature buffer `buf_b1`, and a second weight buffer `buf_a2` and a second feature buffer `buf_b2`. At any given time, one set of input buffers (e.g., `buf_a1` and `buf_b1`) directly drives the input of the two-dimensional parallel multiply-accumulate array, while the other set of input buffers (`buf_a2` and `buf_b2`) prefetches the input vector for the next time slice from an external interface in parallel. By using alternating, effective buffer switching pulses, idle periods caused by waiting for data are avoided in the two-dimensional parallel multiply-accumulate array.
[0069] exist Figure 2 In the diagram, the solid arrow pointing from the reconfigurable controller to the two-dimensional parallel multiply-accumulate array represents the global time pointer and buffer switching decisions maintained by the flow controller. The flow controller injects the input vector and weighting coefficients into the two-dimensional parallel multiply-accumulate array periodically according to the global time pointer using a skew rule. This ensures that each PE performs multiplication after receiving the multiplier and multiplicand and accumulates the multiplier into its local accumulator register. Simultaneously, the flow controller is responsible for issuing backup buffer loading commands and switching pulses at the correct cycles, ensuring that when the two-dimensional parallel multiply-accumulate array completes the calculation for the current time slice, the data for the next time slice is already ready in the backup buffer, supporting loading and computation simultaneously.
[0070] exist Figure 2 In the middle, the slash t0…t n This indicates the timing of the oblique injection. At a certain time t... n On the diagonal, multiple PEs simultaneously receive valid multipliers and perform multiply-add operations in parallel. Slant injection ensures that data is advanced diagonally in the array to achieve streaming matrix multiplication.
[0071] In one specific embodiment, the current computation task can be any of the preprocessing and data flow control unit, linear transformation path, residual path, addition sub-unit, etc.
[0072] If the current computation task is the preprocessing and data flow control unit, meaning the preprocessing and data flow control unit calls the reconfigurable array to perform matrix multiplication, then the matrix dimension corresponding to the current computation task refers to the dimension of the feature matrix output by the feature extraction unit and the normalized adjacency coefficient matrix multiplied with it. The weight configurator reads the corresponding weight coefficients as normalized adjacency coefficients from the weight storage unit. Figure 3 This diagram illustrates the subarray selection of a reconfigurable array unit according to some embodiments of the present invention. It shows how the reconfigurable array unit adjusts the size of the matrix blocks and the active subarrays within the two-dimensional parallel multiply-accumulate array under different functional modules or computational stages. The diagram uses three parallel diagrams (left, middle, and right) to correspond to three types of computational functions: the left side represents the information aggregation stage (requiring N clock cycles to complete the aggregation of one slice), corresponding to the preprocessing and data flow control unit; the middle side represents the parallel transformation and residual merging stage (requiring M clock cycles); and the right side represents the row-level aggregation stage (requiring L clock cycles).
[0073] In each case, the red box indicates the currently enabled subarray region; this subarray can be dynamically selected or partially masked by the reconfigurable controller at compile time or runtime, thereby changing the number of matrix rows, columns, and parallelism effectively involved in the computation.
[0074] The reconfigurable controller is used to adjust the matrix size, computational accuracy, and data path configuration of the array according to different computational tasks, and to send control signals to the weight configurator according to the needs of the current computational task to control the loading, switching, or updating of the corresponding weight parameters. This enables the two-dimensional parallel multiply-accumulate array to load the matching weight configuration under different functional modules or different computational stages and complete the corresponding matrix operations.
[0075] Please see Figure 4 This figure compares the FPGA resource consumption of three network accelerators at different sampling rates. The left side shows the Transformer+CNN scheme with a calibrated ADC sampling rate of 700 MHz; the middle side shows the accelerator described in this invention with a calibrated ADC sampling rate of 2.4 GHz; and the right side shows the MLP implementation with a calibrated ADC sampling rate of 625 MHz. As shown in the figure, this invention maintains a low LUT and FF usage (approximately 4,547 and 2,771 respectively) at the highest sampling rate, and a moderate DSP usage (approximately 69), thus achieving high throughput while maintaining good logic resource efficiency on resource-constrained programmable logic. In contrast, the implementation on the right trades off significantly more LUT / FF consumption (approximately 12,085 and 4,803 respectively) with very few DSPs (2), reflecting another implementation strategy of mapping operations such as multiplication to general-purpose logic. The above comparison reveals different design trade-offs: this invention, through architectural optimization, has significant advantages in balancing high-frequency performance and resource conservation.
[0076] Please see Figure 5This figure illustrates the spectral changes and performance improvements of a commercial pipeline ADC before and after digital calibration using the digital calibrator based on the FPGA network accelerator described in this invention, under two sets of test conditions. The horizontal axis represents the normalized frequency, and the vertical axis represents the amplitude; in the legend, the gray curve represents before calibration, and the blue curve represents after calibration. In the left figure, the SNDR before calibration is approximately 46.6 dB and the SFDR is approximately 51.1 dB, while after calibration, the SNDR is approximately 59.4 dB and the SFDR is approximately 77.8 dB. In the right figure, the SNDR before calibration is approximately 47.6 dB and the SFDR is approximately 55.5 dB, while after calibration, the SNDR is approximately 59.3 dB and the SFDR is approximately 76.9 dB. After calibration, the noise floor decreases, spurious peaks are suppressed, and the dominant frequency components are more concentrated, indicating that digital calibration significantly improves the signal-to-noise ratio and linear distortion suppression capability.
[0077] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An FPGA neural network accelerator for pipelined ADC digital calibration, characterized in that, It includes a top-level interface module, a control and synchronization logic module, and a neural network calibration module; after receiving a start pulse, the top-level interface module transmits it to the control and synchronization logic module to trigger the neural network calibration module to perform digital calibration operations on the original code of each node; The neural network calibration module includes: The feature extraction unit is used to temporarily store the original encoding of multiple consecutive sampling times of the same node to form a short-time history sequence, and generate time-related features and spatial-related features based on the short-time history sequence. The time-related features and spatial-related features are combined into a fixed-width fixed-point vector to obtain the feature matrix. The preprocessing and data flow control unit is used to send the elements of the normalized adjacency coefficient matrix and the elements of the feature matrix into the reconfigurable array unit according to the oblique in-time sequence under the coordination of the streaming control signal to perform matrix multiplication operations and output the support matrix after information aggregation. The parallel transformation and residual merging unit is used to apply a linear transformation to the support matrix based on the reconfigurable array unit and generate residual branch results in parallel. The linear transformation results and residual branch results are then added element by element and the stage output is obtained after nonlinear activation. Row-level aggregation unit is used to perform row-level aggregation operations on the stage output based on the reconfigurable array unit to obtain the final scalar output; The reconfigurable array unit is used to configure the number of active rows and columns of the two-dimensional parallel multiply-accumulate array according to the dimension of the input matrix and the dimension of the weight matrix corresponding to the current computing task, and to control the timing of weight loading and data flow.
2. The FPGA neural network accelerator for pipelined ADC digital calibration according to claim 1, characterized in that, The normalized adjacency coefficient is a fixed-width, fixed-point representation obtained by sequentially performing mathematical normalization and fixed-point scaling on the original adjacency coefficients determined based on the internal circuit connections of the pipeline ADC.
3. An FPGA neural network accelerator for pipelined ADC digital calibration according to claim 1, characterized in that, The reconfigurable array unit includes a reconfigurable controller, a two-dimensional parallel multiply-accumulate array, a flow controller, and a weight configurator, wherein: The reconfigurable controller is used to generate configuration signals based on the matrix dimension, computational precision, and weight parameters corresponding to the current computation task; the configuration signals include at least an array size configuration signal, a data bit width configuration signal, and a weight loading control signal. The weight configurator is used to read the corresponding weight coefficients and biases from the weight storage unit according to the weight loading control signal, and output them to the two-dimensional parallel multiply-accumulate array; The flow controller is used to configure the signal and the data bit width according to the array size, and inject the input vector and the weight coefficients output by the weight configurator into the two-dimensional parallel multiply-accumulate array periodically according to the oblique in timing sequence. The two-dimensional parallel multiply-accumulate array is connected to a reconfigurable controller, a weight configurator, and a flow controller. It consists of multiple processing units arranged in rows and columns, and the number of active rows and columns is dynamically set according to the array size configuration signal. Each processing unit is configured to receive a weight coefficient and an input vector, perform multiply-accumulate operations, and pass the weight coefficient or input vector along the array direction to the adjacent processing unit.
4. An FPGA neural network accelerator for pipelined ADC digital calibration according to claim 3, characterized in that, The reconfigurable array unit further includes: Multiple input buffers are configured, each including a weight input buffer and a feature input buffer. The weight input buffer is used to temporarily store weight coefficients, and the feature input buffer is used to temporarily store input vectors. The multiple input buffers work alternately in a ping-pong manner, and under the coordination of the flow controller, the corresponding data is sent into the two-dimensional parallel multiply-accumulate array in an oblique in-order manner, so that the weight coefficients and input vectors are correctly paired in the internal processing unit of the two-dimensional parallel multiply-accumulate array and matrix multiplication is completed. At any given time, only one input buffer is effective and drives the input of the two-dimensional parallel multiply-accumulate array, while the other input buffer groups prefetch the next batch of weight coefficients and input vectors in parallel.
5. An FPGA neural network accelerator for pipelined ADC digital calibration according to claim 1, characterized in that, The parallel transformation and residual merging unit includes: The linear transformation path connects reconfigurable array cells and is used to perform matrix multiplication on the support matrix and its corresponding weight matrix. The result of the operation is then superimposed with the corresponding bias to obtain the linear transformation result. The residual path connects the reconfigurable array units and is used to project and transform the feature matrix output by the feature extraction unit to obtain the residual branch result that is aligned with the linear transformation result in both dimension and temporal order. An element-wise additive unit array is used to add the linear transformation result and the residual branch result element-wise, and then obtain the stage output after nonlinear activation.
6. An FPGA neural network accelerator for pipelined ADC digital calibration according to claim 1, characterized in that, Row-level clustering units include: Part and generation unit, connected to reconfigurable array unit, used to generate the first intermediate accumulated value and the second intermediate accumulated value corresponding to each row vector in the output of the generation stage; The inline modulation unit is used to multiply the original encoded scalar of each row vector in the output of the generation stage by its first intermediate accumulated value, and add the product result to its second intermediate accumulated value to obtain the row-level output. The aggregation unit is used to left-shift and weight each row-level output according to a predetermined displacement, sum all the left-shift weighted results, and then perform an arithmetic right shift and bit width truncation to obtain the final scalar output.
7. An FPGA neural network accelerator for pipelined ADC digital calibration according to claim 1, characterized in that, The top-level interface module is configured with: Parallel input bus for receiving the raw code of each node of the pipelined ADC output; The system clock input terminal is used to receive the system clock signal; The asynchronous reset input terminal is used to receive the asynchronous reset signal; The start signal input terminal is used to receive the start pulse sent by the upper-level controller; The software reset signal input terminal is used to receive the software reset pulse issued by the upper-level controller; The busy flag output terminal is used to output a busy flag signal to the upper-level controller to indicate whether the FPGA neural network accelerator is currently in operation; The completion pulse output terminal is used to output a single-cycle completion pulse signal to the upper-level controller at the end of a complete digital calibration operation.