Full-localization ultra-high-speed multichannel ADC (Analog to Digital Converter) acquisition and processing platform
By utilizing domestically produced high-stability crystal oscillators and FPGA dynamic phase adjustment technology, combined with calibration pulse injection and fractional delay filters, sub-sampling accuracy alignment and synchronous output of multi-channel signals at ultra-high sampling rates were achieved, solving the synchronization problem in multi-channel signal acquisition and improving the system's autonomy and security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-03
- Publication Date
- 2026-05-01
AI Technical Summary
Existing signal acquisition and processing solutions struggle to achieve synchronous capture of multi-channel signals at high sampling rates, leading to data distortion and misalignment between channels. Furthermore, they are heavily reliant on foreign core components, compromising autonomy and security.
The main clock signal is generated by a domestically produced high-stability crystal oscillator. The clock path deviation is compensated in real time by the dynamic phase adjustment circuit inside the FPGA. Combined with periodic calibration pulse injection and multi-stage pipeline comparators, the pulse edge position is accurately detected. Dynamic interpolation resampling is performed using a reconfigurable fractional delay filter to achieve subsampling accuracy alignment and hierarchical flow control output.
It achieves picosecond-level synchronization accuracy of the sampling reference clock for each channel, suppresses the phase difference between channels to within 0.05 sampling periods, solves the problems of data distortion and processing congestion under ultra-high sampling rates, and generates a synchronous data stream that can be used for radar signal processing.
Smart Images

Figure CN121966570A_ABST
Abstract
Description
A domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform Technical Field
[0001] This invention relates to the field of signal processing technology, specifically a fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform. Background Technology
[0002] In the field of modern information processing and communication, the importance of signal acquisition and processing technology is self-evident. It directly relates to the efficiency of data transmission, the stability of the system, and the execution capability of critical tasks. This field is not only a core support for industries such as defense, communications, and industrial automation, but also a crucial cornerstone for driving technological progress. Whether it's capturing radar signals or analyzing high-speed communication data, signal acquisition and processing plays an irreplaceable role, and its performance often determines the success or failure of the entire system. However, current mainstream signal acquisition and processing solutions often reveal some deep-seated shortcomings when dealing with complex requirements. These solutions are often limited by the flexibility and scalability of hardware architecture, making it difficult to maintain synchronization and stability under high-speed and multi-channel demands, especially when dealing with high-frequency signals, where performance bottlenecks frequently emerge. More importantly, many existing technologies rely heavily on foreign core components. Once the supply chain is restricted, autonomy and security are difficult to guarantee, which becomes a potential risk in critical application scenarios.
[0003] Focusing on the technical challenges, a core challenge in the field of signal acquisition and processing lies in how to achieve synchronous acquisition of multi-channel signals at ultra-high sampling rates. The higher the sampling rate, the more easily subtle deviations in the signal over time are amplified, leading to data distortion or misalignment between different channels. For example, when processing signals with frequencies as high as several gigahertz, even a slight misalignment in time among multiple acquisition channels can result in serious errors in the final synthesized signal, affecting the accuracy of subsequent analysis. More complexly, this high-speed synchronization requirement further intensifies the pressure on data interaction between hardware components. Traditional processing architectures often cannot handle such a high-intensity information flow, causing system slowdowns or even crashes. Summary of the Invention
[0004] The purpose of this invention is to provide a fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform, which enables sub-sampling accuracy alignment and hierarchical flow control output of multi-channel ADCs at ultra-high sampling rates.
[0005] The objective of this invention can be achieved through the following technical solution: This application provides a fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform, comprising: a high-stability clock distribution and calibration module, which uses a fully domestically produced high-stability crystal oscillator to generate a master clock signal, distributes the master clock signal to each acquisition channel through a dedicated clock driver chip, and uses the dynamic phase adjustment circuit inside the FPGA to compensate for clock path deviation in real time, thereby obtaining a synchronous sampling reference clock for each channel; a synchronous acquisition and pulse injection module, which drives the fully domestically produced ADC of each channel to work synchronously with the sampling reference clock, and periodically injects calibration pulses with known characteristics into the analog input terminal of each channel to obtain an original multi-channel high-speed sampling data sequence containing calibration pulses; a pulse detection and integer deviation module, which uses the high-speed parallel processing unit embedded in the FPGA to detect the rising or falling edge position of the calibration pulse of each channel in the original multi-channel high-speed sampling data sequence in real time, accurately extracts the pulse occurrence time through a multi-stage pipeline comparator, and calculates the fixed deviation of the sampling clock of each channel relative to the master clock as an integer multiple of the sampling period; and a coarse-grained translation and alignment module, which, according to the fixed deviation of the sampling period as an integer multiple of the sampling period, performs a configurable translation and alignment within the FPGA. The bit register dynamically shifts the data stream of each channel by an integer number of sampling periods to obtain a preliminary time-aligned multi-channel data sequence. The sub-sampling phase extraction module extracts waveform details of the calibration pulse from the preliminary time-aligned multi-channel data sequence and calculates the sub-sampling phase difference of each channel pulse relative to the reference channel using FFT interpolation or correlation peak fitting algorithms to obtain the remaining fractional sampling period deviation. The fractional delay interpolation alignment module uses a reconfigurable fractional delay filter to dynamically interpolate and resample the data of each channel according to the fractional sampling period deviation, obtaining a sub-sampling precision aligned multi-channel synchronous data sequence. The hierarchical storage and flow control module writes the sub-sampling precision aligned multi-channel synchronous data sequence into a domestic high-speed cache array according to a preset priority, and uses the FPGA's built-in intelligent flow control engine to dynamically adjust the cache strategy based on the instantaneous data rate and the load of subsequent processing modules, generating controlled parallel data blocks. The high-speed interface synchronous output module reads the controlled parallel data blocks from the domestic high-speed cache array in parallel with time consistency and transmits them to the subsequent processing modules through a fully domestically produced high-speed interface to obtain an ultra-high sampling rate synchronous data stream.
[0006] The beneficial effects of this invention are as follows: This invention, through the collaborative work of a high-stability clock distribution calibration module, a synchronous acquisition and pulse injection module, a pulse detection and integer deviation module, and a coarse-grained translation alignment module, uses a domestically produced high-stability crystal oscillator to generate the master clock signal and dynamically compensates for clock path deviations. Combined with periodically injected known characteristic calibration pulses, it utilizes the high-speed parallel processing unit embedded in the FPGA and a multi-stage pipeline comparator to accurately detect the pulse edge position, calculates the fixed deviation of the integer multiple sampling period for each channel, and performs coarse-grained translation alignment using a configurable shift register. This solves the problem of multi-channel sampling clock asynchrony and integer multiple period misalignment caused by clock path differences and hardware architecture limitations, achieving picosecond-level synchronization accuracy and preliminary time alignment of multi-channel data sequences for each channel's sampling reference clock. Through the fine processing of the sub-sampling phase extraction module and the fractional delay interpolation alignment module, it extracts calibration pulse waveform details from the preliminary time-aligned multi-channel data sequences, calculates the sub-sampling-level phase difference using FFT interpolation or correlation peak fitting algorithms, and uses a reconfigurable fractional delay filter to dynamically interpolate based on the fractional deviation. By sampling and iteratively optimizing through residual deviation detection and fine-tuning units until the deviation converges, the technical challenges of amplifying small sub-sampling-level deviations during UHF signal acquisition, leading to data distortion and inaccurate channel alignment, are solved. This achieves sub-sampling precision aligned multi-channel synchronous data sequences with channel phase differences suppressed to within 0.05 sampling periods. Through the organic combination of hierarchical storage, flow control modules, and high-speed interface synchronous output modules, the sub-sampling precision aligned multi-channel synchronous data sequences are written into a domestic high-speed cache array according to preset priorities. The FPGA's built-in intelligent flow control engine dynamically adjusts the cache strategy based on the instantaneous data rate and the load of subsequent processing modules, and reads data blocks in parallel according to timing consistency, transmitting them through a fully domestically produced high-speed interface. Channel alignment and timestamp correction are performed at the receiving end, solving the problems of processing congestion caused by explosive data growth at ultra-high sampling rates and additional time deviations introduced by differences in high-speed transmission paths. This generates an ultra-high sampling rate synchronous data stream that can be directly used for joint analysis tasks such as radar signal processing and interference cancellation, achieving smooth data flow scheduling and end-to-end accurate synchronous output. Attached Figure Description
[0007] To better understand and implement this application, the technical solution is described in detail below with reference to the accompanying drawings.
[0008] Figure 1 is a schematic diagram of the structure of a fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform provided in Embodiment 1 of this application; Figure 2 is a schematic diagram of the structure of the pulse detection and integer deviation module in the fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform provided in Embodiment 1 of this application; Figure 3 is a schematic diagram of the structure of the fractional delay interpolation alignment module in the fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform provided in Embodiment 1 of this application. Detailed Implementation
[0009] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, exemplary embodiments will be described in detail below, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0010] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0011] The following detailed description of the specific implementation methods, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided in detail.
[0012] Example 1, please refer to Figures 1-3. This example provides a fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform, including: a high-stability clock distribution and calibration module, which uses a fully domestically produced high-stability crystal oscillator to generate a master clock signal, distributes the master clock signal to each acquisition channel through a dedicated clock driver chip, and uses the FPGA's internal dynamic phase adjustment circuit to compensate for clock path deviation in real time, so as to obtain a synchronous sampling reference clock for each channel.
[0013] Furthermore, the high-stability clock distribution calibration module includes: a master clock generation and distribution unit, which generates a master clock signal using a domestically produced high-stability crystal oscillator, receives the master clock signal through a dedicated clock driver chip, and distributes the master clock signal to each channel according to the physical location and wiring length differences of each acquisition channel to obtain an initial clock path distribution; wherein, the domestically produced high-stability crystal oscillator is selected as model XX (e.g., domestic OCXO series), with a frequency stability better than ±0.1ppm and a phase noise lower than -150dBc / Hz@1kHz; the dedicated clock driver chip is a domestically produced multi-output clock buffer, such as a certain model from Guowei Electronics, with no less than the number of acquisition channels, and each output channel has independent output enable and drive strength configuration functions, which can dynamically compensate for the wiring load differences of different channels; the initial clock path distribution is obtained by calculating using static timing analysis tools in the PCB design stage in combination with the actual wiring length, and is fixed in the configuration ROM of the FPGA in the form of a lookup table as a baseline reference for subsequent phase deviation detection.
[0014] The phase deviation detection unit utilizes the FPGA's internal clock management module to extract the actual clock signals received by each acquisition channel, obtaining the actual arrival phase information for each channel. It then compares the actual arrival phase of each channel with the main clock reference phase to obtain the phase deviation value for each channel. The FPGA uses a Fudan Microelectronics JFM9RFVU3P series chip, whose internal clock management module includes multiple phase-locked loops (PLLs) and fine phase detectors. The PLLs synchronize the received clock signals from each channel to the FPGA's internal high-speed system clock domain. The phase detectors measure the time difference between the clock edge of each channel and the main clock reference edge using a counting method. The counting accuracy is determined by the delay of the FPGA's internal carry chain, reaching picosecond levels. The main clock reference phase is provided by a reference counter that increments from zero after system synchronization reset. The phase deviation value is stored as a signed integer, with positive values indicating delay and negative values indicating lead.
[0015] The dynamic phase adjustment unit determines whether the phase deviation value of each channel exceeds a preset threshold range. If it does, a programmable phase shift is applied to the clock path of the corresponding channel through the dynamic phase adjustment circuit inside the FPGA to obtain the adjusted channel clock phase. The dynamic phase adjustment circuit is implemented using the IODELAY primitive inside the FPGA. Each input clock pin is equipped with a programmable delay line, and the delay step is calibrated by the IDELAYCTRL reference clock, with a typical step of 78ps (corresponding to a 300MHz reference clock). The preset threshold range is set according to the maximum allowable clock skew required by the system, for example, ±50ps, and this value is stored in the FPGA's internal register. When the phase deviation of a certain channel exceeds the threshold, the state machine calculates the number of steps required for adjustment, and performs phase fine-tuning by dynamically configuring the delay value of IODELAY. The phase deviation detection result is updated immediately after adjustment.
[0016] The iterative synchronous convergence unit re-extracts the adjusted clock signals of each channel and updates the phase deviation value. It repeats the phase deviation comparison and dynamic phase adjustment until the phase deviation values of all channels converge to the preset range, and finally obtains a strictly synchronized sampling reference clock for each channel.
[0017] The iterative synchronization convergence unit is controlled by a finite state machine inside the FPGA. The state machine includes states such as idle, detection, adjustment, and verification. In each iteration, the phase deviation detection unit is first enabled to obtain the latest deviation value. If the deviation of any channel exceeds the threshold, the state is entered and the dynamic phase adjustment unit is called to correct it. Then, the state is returned to the detection state. The iteration count counter records the current loop count. When all channels are detected within the threshold twice in a row or the iteration count reaches the maximum value (e.g., 16 times), the state machine enters the completion state, latches the current adjustment parameters, and outputs a synchronization completion flag. Finally, the maximum static phase difference between the sampling reference clocks of each channel is suppressed within a preset threshold to ensure that subsequent multi-channel ADCs can achieve picosecond-level synchronous acquisition.
[0018] Specifically, by using dynamic phase adjustment and iterative convergence technology, the problem of asynchronous sampling clocks caused by differences in clock paths in multi-channel acquisition systems was solved, achieving picosecond-level synchronization accuracy of the sampling reference clocks for each channel.
[0019] The synchronous acquisition and pulse injection module drives the domestically produced ADCs of each channel to work synchronously with the sampling reference clock, and periodically injects calibration pulses with known characteristics into the analog input terminal of each channel to obtain the original multi-channel high-speed sampling data sequence containing the calibration pulses.
[0020] Furthermore, the synchronous acquisition and pulse injection module includes: a synchronous sampling and pulse injection unit, which drives the domestically produced ADCs of each channel to perform synchronous sampling operations with a sampling reference clock, and periodically injects calibration pulses with known characteristics into the analog input terminal to obtain the original multi-channel high-speed sampling data sequence containing the calibration pulses; wherein, the domestically produced ADC adopts the Fudan Micro 47DR series chip, which supports 14-bit resolution, a maximum single-channel sampling rate of 5Gps and 8-channel synchronous acquisition function. Its internally integrated phase-locked loop is locked to the sampling reference clock to ensure that all ADCs start sampling conversion on the same clock edge; the calibration pulse is generated by an arbitrary waveform generator integrated inside the FPGA. The FPGA adopts the Fudan Micro JFM9RFVU3P series. After the digital pulse waveform is converted into an analog signal by its internal digital-to-analog converter, it is injected into the analog input terminal of each channel through an RF switch; the known characteristics of the calibration pulse include pulse width, rise time, fall time, amplitude and pulse repetition interval. These characteristic parameters are set according to the linear frequency modulated pulse or Barker code pulse commonly used in radar countermeasure signal processing, and are pre-stored in the configuration register of the FPGA and loaded into the waveform generator during system initialization.
[0021] The pulse timing extraction and offset detection unit detects the periodic occurrence positions of calibration pulses in each channel from the original sampling sequence, generating a pulse timing sequence. It calculates the sampling point interval between adjacent pulses and identifies channels with sampling clock offsets by determining the interval constancy marker, thus obtaining multi-channel information including pulse timing and offset markers. The detection of the periodic occurrence positions of calibration pulses employs a combination of a digital energy detector and pulse shape matching. The energy detector determines the approximate arrival position range of the pulse by comparing the sampling point amplitude with a preset threshold. The pulse shape matching algorithm calculates the correlation coefficient between the sampling sequence and a pre-stored pulse template within the candidate range; the position corresponding to the maximum correlation coefficient is the precise arrival time of the pulse. The sampling point interval between adjacent pulses is calculated using a differential counter within the FPGA. The counter's clock frequency is an integer multiple of the sampling clock. If the difference between the interval value and the number of sampling points corresponding to the theoretical pulse repetition interval exceeds a preset tolerance (e.g., greater than one sampling period), it is determined that the channel has a sampling clock offset. The offset marker is appended to the channel status word in the form of a binary flag and stored together with the pulse timing sequence in a dual-port RAM for subsequent reading by subsequent units.
[0022] The pulse feature analysis and gain correction unit extracts the amplitude values and waveform samples of the calibration pulses for each channel based on multi-channel information, forming a pulse feature set. The amplitude values in this set are compared with a preset reference amplitude to determine the gain deviation of each channel and generate gain correction coefficients. These correction coefficients are then used to adjust the amplitude of the effective data segments in the original sequence, resulting in a gain-corrected multi-channel sampling sequence. The amplitude value extraction employs a peak detection circuit, taking N sampling points before and after the pulse moment to form a waveform window. The maximum value within the search window is taken as the pulse amplitude, and the sampling point position corresponding to this maximum value is recorded for subsequent time delay analysis. The waveform samples... This is a numerical sequence of all sampling points within the window, used for waveform similarity comparison between channels; the preset reference amplitude is the theoretical amplitude value of the calibration pulse injected during system calibration. This theoretical amplitude value is converted into a digital code value by the ADC nominal conversion gain and stored in the FPGA internal register; the gain deviation is the ratio of the measured amplitude to the reference amplitude, expressed in fixed-point decimal form, ranging from 0.5 to 2.0; the gain correction coefficient is calculated by the reciprocal of the deviation, and the effective data segments in the original sampling sequence, excluding the calibration pulse, are weighted and adjusted in real time using the parallel multiplier inside the FPGA to ensure the consistency of amplitude response of each channel.
[0023] The channel delay parameter calculation unit calculates the fixed delay difference of each channel relative to the reference channel based on the correspondence between the calibration pulse position and the sampling reference clock in the sampled sequence after gain correction, and generates multi-channel delay correction parameters for subsequent fine synchronization processing.
[0024] Specifically, based on the correspondence between the calibration pulse position and the sampling reference clock in the sampled sequence after gain correction, the fixed delay difference of each channel relative to the reference channel is calculated to generate multi-channel delay correction parameters for subsequent fine synchronization processing.
[0025] The reference channel is a pre-selected acquisition channel, typically the one with the central physical location or the shortest wiring length, and its calibration pulse position serves as a reference. The calibration pulse position is provided by the pulse timing sequence output by the pulse timing extraction unit. After gain correction, a waveform cross-correlation method is used for sub-sampling level precise positioning, and the fractional part of the peak position of the cross-correlation function is obtained through parabolic fitting. The fixed delay difference is the time difference between the pulse timing of each channel and the pulse timing of the reference channel, expressed in units of sampling periods. The integer part is determined by the difference in the number of sampling points, and the fractional part is determined by the cross-correlation fitting result. The multi-channel delay correction parameters are stored in a correction coefficient table inside the FPGA in floating-point format. Each channel corresponds to a 32-bit fixed-point number, with the high 16 bits representing the integer period delay and the low 16 bits representing the fractional period delay. This parameter table interacts with the subsequent fine synchronization processing module via the AXI bus.
[0026] Specifically, by using synchronous acquisition and calibration pulse injection technology, the problems of inconsistent initial sampling times and channel gain mismatch among multi-channel ADCs were solved, providing a raw data sequence containing amplitude and time references for subsequent accurate calibration between channels.
[0027] The pulse detection and integer deviation module uses the high-speed parallel processing unit embedded in the FPGA to detect the rising or falling edge position of the calibration pulse of each channel in the original multi-channel high-speed sampling data sequence in real time. It accurately extracts the pulse occurrence time through a multi-stage pipeline comparator and calculates the fixed deviation of the sampling clock of each channel relative to the main clock as an integer multiple of the sampling period.
[0028] Furthermore, the pulse detection and integer deviation module includes: a parallel pulse edge detection unit, which acquires the original multi-channel high-speed sampling data sequence and uses the high-speed parallel processing unit embedded in the FPGA to synchronously perform pulse edge detection on the data of each channel; by setting rising edge threshold conditions and falling edge threshold conditions, it records the sampling point indices that meet the conditions as candidate rising edge times and candidate falling edge times, respectively, to obtain the candidate edge time set for each channel; wherein, the rising edge threshold condition is that the amplitude of the current sampling point exceeds a preset high threshold and the amplitude of the previous point is lower than a preset low threshold, and the falling edge threshold condition is that the amplitude of the current sampling point is lower than the preset low threshold and the amplitude of the previous point exceeds a preset low threshold. High threshold; preset high and low thresholds are adaptively set according to the typical amplitude and noise floor of the calibration pulse. They are generated by the threshold learning module inside the FPGA after statistically analyzing the background noise of each channel during system initialization. The FPGA uses the Fudan Micro JFM9RFVU3P series chip, whose embedded high-speed parallel processing unit contains multiple digital signal processing slices, which can process eight channels of data in parallel at the same time. The candidate edge time is recorded in the form of sampling point count value, with additional channel identifier and edge type flag, and stored in distributed RAM. Each channel corresponds to an independent candidate edge queue, and the queue depth can be configured to 16 or 32 to adapt to different pulse repetition frequencies.
[0029] A multi-stage pipelined screening unit inputs the candidate edge time sets of each channel into a multi-stage pipelined comparator, compares the amplitude change rate of each candidate edge stage by stage, and selects the edge position with the steepest amplitude change as the precise occurrence time of the calibration pulse for that channel, thus obtaining the precise pulse time value for each channel. The multi-stage pipelined comparator consists of cascaded comparison stages. Each stage compares the amplitude change rate of two adjacent candidate edges. The amplitude change rate is obtained by calculating the sum of the absolute differences between two sampling points before and after the edge, i.e., change rate = |S(n+1)|. -S(n)|+|S(n+2)-S(n-1)|, where S(n) is the sampling point corresponding to the candidate edge; the comparison stage retains the candidate edge with the larger rate of change and eliminates the candidate edge with the smaller rate of change. After three-stage pipeline processing, each channel retains only one candidate edge with the largest rate of change as the precise pulse edge of that channel; the precise time value is the sampling point index corresponding to the candidate edge, and after subsampling interpolation correction, it is output in the form of a 32-bit fixed-point number, where the high 24 bits are the sampling point count and the low 8 bits are the subsampling offset.
[0030] The integer multiple deviation calculation unit performs a difference calculation between the precise pulse time of each channel and the master clock reference time to obtain the number of sampling points corresponding to the time difference. By determining that this number of points is an integer multiple of the sampling period, a fixed integer multiple sampling period deviation of each channel's sampling clock relative to the master clock is determined. The master clock reference time is a pre-set reference time, generated by an internal FPGA counter that increments from zero after system synchronization reset. The counter has a 48-bit width and can cover continuous acquisition time of up to several hours. The difference calculation is implemented using a pipelined subtractor to obtain the pulse time of each channel and the reference time. The difference between the sampling points is calculated. If the difference is divisible by the count corresponding to the sampling clock cycle, the quotient is directly output as an integer multiple of the deviation. If there is a remainder, the pulse time is fine-tuned and recalculated. The fine-tuning is achieved by rounding the lower 8 bits of the subsampling offset of the pulse time. That is, if the lower 8 bits are greater than 128, the sampling point count is incremented by 1; otherwise, it remains unchanged, ensuring that the obtained deviation is strictly an integer multiple of the sampling cycle. The integer multiple deviation is stored in 16-bit two's complement form, representing the number of cycles of delay or advance of the channel clock relative to the master clock, with a value range of -32768 to +32767.
[0031] The deviation value storage output unit stores the calculated fixed deviation values (in integer multiples) of each channel into the FPGA's internal register group, forming a deviation lookup table for the subsequent data alignment and correction module to read in real time.
[0032] The register group is implemented using dual-port RAM with a depth equal to the maximum number of channels (e.g., 16 channels). One port is used for writing and updating by the deviation calculation unit, and the other port is used for parallel reading by the subsequent coarse-grained translation and alignment module. The deviation lookup table is indexed by the channel address, and each storage unit contains a 16-bit integer multiple deviation value, a 1-bit valid flag, and a 4-bit cyclic redundancy check code to detect data transmission errors. When clock deviation drifts due to temperature or voltage changes during system operation, the deviation detection module can periodically recalculate and update the lookup table. The update cycle is configured by the system software, with a typical value of once every 100 milliseconds. The subsequent coarse-grained translation and alignment module reads the lookup table through the AXI bus to obtain the translation cycle number of each channel, ensuring the real-time performance and accuracy of data alignment.
[0033] Specifically, by using parallel pulse edge detection and multi-stage pipeline screening technology, the problem of accurate pulse edge detection under ultra-high-speed sampling was solved, and the accurate calculation and storage of fixed deviations of sampling periods of integer multiples for each channel were achieved.
[0034] The coarse-grained translation alignment module, based on the fixed deviation of the integer multiple sampling period, dynamically shifts the data stream of each channel by an integer number of sampling periods within the FPGA through a configurable shift register to obtain a preliminary time-aligned multi-channel data sequence.
[0035] Furthermore, the coarse-grained translation alignment module includes: an integer deviation parsing and mapping unit, which reads the fixed deviation of the sampling period in integer multiples of each channel from the register group, parses out the number of integer translation periods required for each channel; maps the number of translation periods to the depth parameter of the configurable shift register inside the FPGA, generates an independent shift depth configuration word for each channel, and synchronously writes it to the control register to provide a configuration basis for subsequent dynamic translation; wherein, the register group is the deviation lookup table formed in the aforementioned pulse detection and integer deviation module, adopts a dual-port RAM structure, and obtains the integer multiple deviation values of each channel in parallel through a dedicated read interface using the channel address as an index. The analysis process includes determining the sign of the deviation value. If the deviation value is positive, it indicates that the channel clock is lagging behind the master clock, and the data needs to be delayed by the corresponding number of cycles. If the deviation value is negative, it indicates that the data is ahead, and the data needs to be advanced by the corresponding number of cycles. However, due to the limitations of the causal system, in actual implementation, the same reference delay is applied to all channels before relative translation to ensure the continuity of the data stream. The shift depth configuration word is generated by binary encoding of the translation cycle number. Each channel corresponds to a set of depth parameters, which are synchronously written to the control register of each channel shift register through the AXI-Lite bus. The configuration process is completed within one sampling clock cycle.
[0036] The dynamic depth shift array unit dynamically adjusts the depth of each channel's shift register through FPGA logic based on the shift depth configuration word for each channel. The original multi-channel data stream is fed in parallel into the respective shift registers with their configured depths. Data is shifted stage by stage under the drive of the sampling clock. When the preset depth is reached after a certain number of shifts, data is extracted in parallel from the register output, achieving dynamic delay compensation for an integer number of sampling cycles, resulting in a coarse-grained shifted data stream for each channel. The dynamic depth shift array is implemented using SRL16E primitives or distributed RAM within the FPGA. Each channel corresponds to a cascaded chain of shift registers, and its depth is selected by the configuration word. The tap positions are dynamically adjusted; the FPGA uses Fudan Microelectronics JFM9RFVU3P series chips, which have rich configurable logic resources and can simultaneously support shift register arrays with up to 16 channels and a maximum shift depth of 1024 per channel; the data flow process in the shift register is as follows: at the rising edge of the sampling clock, the current input data is written to the first-level register, and the values of each level register are passed sequentially. After a preset depth of clock cycles, the original data appears at the output, thereby achieving integer cycle delay compensation; the shift operations of each channel are strictly synchronized to ensure that all channel data is shifted on the same clock edge.
[0037] The parallel buffer synchronous output unit writes the shifted data streams of each channel into the parallel buffer array according to the original timing sequence, eliminating the instantaneous output phase offset caused by the difference in the depth of the shift register; it applies a unified synchronous read enable signal to all buffer channels and reads the data of each channel at the same time to obtain a preliminary time-aligned multi-channel data sequence that is strictly arranged according to the original timing sequence and aligned with integer periods.
[0038] The parallel cache array is composed of BlockRAM within the FPGA, configured in a true dual-port mode. Each channel corresponds to an independent storage space, with the write port connected to the shift register output and the read port connected to the subsequent processing module. Writing according to the original timing sequence means that data from each channel is written to the cache sequentially according to the order of sampling times. The write address is uniformly generated by a global counter, ensuring that all channel data has the same time base in the cache. The synchronous read enable signal is generated by a timing control state machine. When all channels have completed a preset depth shift and sufficient data has accumulated in the cache, the state machine issues a unified read enable pulse. Each channel reads data in parallel from the same offset address in the same clock cycle, thereby eliminating the inconsistency in read times that may be introduced by differences in shift depth. In the final output multi-channel data sequence, data at the same sampling time is strictly aligned across all channels by integer multiples of the sampling period, laying the foundation for subsequent sub-sampling level fine synchronization.
[0039] Specifically, by using dynamic depth adjustment of configurable shift register arrays and parallel buffer synchronous output technology, the data misalignment problem caused by the difference in sampling periods of integer multiples of each channel was solved, and preliminary time synchronization with strict alignment of integer periods was achieved.
[0040] The subsampling phase extraction module extracts waveform details of calibration pulses from the initial time-aligned multi-channel data sequence, calculates the subsampling phase difference of each channel pulse relative to the reference channel using FFT interpolation or correlation peak fitting algorithms, and obtains the remaining fractional sampling period deviation.
[0041] Furthermore, the subsampling phase extraction module includes: a calibration pulse waveform extraction unit, which acquires a preliminary time-aligned multi-channel data sequence; for each channel data, it locates a time window based on the periodic occurrence position of the calibration pulse, extracts waveform sampling points within the window, and obtains the calibration pulse waveform for each channel; it selects one channel as a reference channel, fixes its calibration pulse waveform as the reference waveform, and uses the waveforms of the remaining channels as waveforms to be matched; wherein, the periodic occurrence position of the calibration pulse is provided by the pulse time sequence generated in the synchronous acquisition and pulse injection module, and each pulse time corresponds to a sampling point index; the time window is centered on this index, and the preceding... Each channel is then expanded by M sampling points. The value of M is set according to the calibration pulse width, typically twice the number of sampling points corresponding to the pulse rise time. For example, when the sampling rate is 2.4 Gps and the pulse rise time is 1 ns, M is 5. The waveform sampling points include the amplitude values of all sampling points within the window, forming a vector of length 2M+1. The selection of the reference channel is automatically determined based on the channel signal quality, usually selecting the channel with the highest signal-to-noise ratio or the smallest pulse amplitude fluctuation. If the data of this channel is abnormal, it is switched to the suboptimal channel. The calibration pulse waveforms of each channel are stored in the FPGA's internal BlockRAM in a fixed-point format for real-time reading in subsequent cross-correlation calculations.
[0042] The correlation peak fitting unit performs cross-correlation operations on the calibration pulse waveform of each channel to be matched and the reference waveform of the reference channel to obtain a cross-correlation function sequence; detects the peak position of the correlation function to obtain the preliminary time offset at the integer sampling point level; performs parabolic fitting on several points near the peak to accurately calculate the sub-sampling level peak offset, and obtains the sub-sampling level phase difference of each channel relative to the reference channel. If the phase difference exceeds a preset threshold, the channel data is marked as abnormal. The cross-correlation operation is implemented in parallel using digital signal processing slices within the FPGA, performing convolution operations on waveform vectors of length L to obtain a correlation function sequence of length 2L-1; the peak position detection compares the correlation function values point by point using a comparator, recording the index corresponding to the maximum value as the integer sampling point level offset; the parabolic fitting selects the peak and two points before and after it, for a total of five points, for quadratic curve fitting, with the fitting formula being y=ax. 2 +bx+c, the peak offset is calculated by -b / (2a), and the result is represented by a 32-bit fixed-point number, where the integer part occupies 16 bits and the fractional part occupies 16 bits, with a resolution of 1 / 65536 sampling period; the preset threshold is set according to the system synchronization accuracy requirements, for example, ±0.2 sampling period. If the phase difference of a certain channel exceeds this range, the abnormal flag in the status register of that channel is set to 1, and the abnormal event is recorded in the system log.
[0043] The FFT interpolation deviation correction unit performs FFT transformation on the calibration pulse waveform to obtain the frequency domain phase spectrum; it uses the phase spectrum difference to calculate the remaining fractional sampling period deviation between each channel and the reference channel to obtain the final deviation correction amount; and it outputs the correction amount to subsequent modules for fine time adjustment to ensure subsampling accuracy synchronization.
[0044] The FFT transformation employs a radix-2 time decimation algorithm, with the number of transformation points N being either 256 or 512, determined by the waveform window length. The transformation process is completed within the FFT IP core of the FPGA, which supports a pipelined architecture and can complete multi-channel waveform transformation within continuous sampling intervals. The frequency domain phase spectrum is obtained by calculating the complex phase angle at each frequency point. The phase spectrum difference is calculated as a weighted average of the phase differences between each channel and the reference channel at each frequency point, with the weights set according to the signal-to-noise ratio (SNR) at each frequency point; the higher the SNR, the greater the weight. The remaining fractional sampling period deviation is obtained by fitting the phase difference slope using the least squares method, with the calculation formula being Δτ=(1 / 2π)×(dφ / df), where dφ / df is the slope of the phase change with frequency. The final deviation correction is stored in a correction coefficient register in 32-bit floating-point format and transmitted in real-time to the fractional delay interpolation alignment module via the AXI-Stream interface as a control parameter for fine-tuning the time.
[0045] Specifically, by using cross-correlation peak fitting and FFT interpolation correction techniques, the problem of residual subsampling-level phase deviation after coarse-grained alignment was solved, and high-precision extraction of the remaining fractional sampling period deviation of each channel was achieved.
[0046] The fractional delay interpolation alignment module uses a reconfigurable fractional delay filter to dynamically interpolate and resample the data of each channel according to the fractional sampling period deviation, thereby obtaining a multi-channel synchronous data sequence with sub-sampling precision alignment.
[0047] Furthermore, the fractional delay interpolation alignment module includes: a deviation coefficient generation unit, which acquires the fractional sampling period deviation and original sampling period parameters for each channel, adopts a reconfigurable fractional delay filter structure for each channel, dynamically adjusts the filter tap coefficients according to the deviation, and generates fractional delay compensation coefficients adapted to the current deviation; wherein, the fractional sampling period deviation is provided by the final deviation correction amount output by the subsampling phase extraction module and stored in the correction coefficient register in 32-bit floating-point format; the reconfigurable fractional delay filter is implemented using a Farrow structure, and its coefficients are calculated in real time by polynomial interpolation based on the deviation, with the filter order selected as [missing information]. The filter can be configured dynamically according to system resources and accuracy requirements, with either 4th or 8th order. The FPGA uses Fudan Microelectronics' JFM9RFVU3P series chip, which has abundant DSP48E1 slices for parallel implementation of multiplication and accumulation operations. Each channel is allocated a set of DSP resources to ensure pipelined processing of coefficient generation and filtering operations. During the compensation coefficient generation process, the decimal deviation is first normalized to the range of [-0.5, 0.5]. Then, the coefficients of each order of the filter are quickly obtained by combining table lookup and linear interpolation. The coefficient update cycle is synchronized with the calibration pulse injection cycle, with a typical value of updating once every 100 microseconds to adapt to slow deviations caused by temperature or voltage.
[0048] The dynamic interpolation and resampling unit uses generated compensation coefficients to perform dynamic interpolation and resampling on the original data sequences of each channel. New sampling points are interpolated between integer sampling points through filter convolution operations to obtain a multi-channel data sequence with preliminary sub-sampling precision alignment. The original data sequence is a multi-channel data stream that has undergone coarse-grained translation alignment and is continuously input to the filter with the sampling clock as the beat. The interpolation and resampling process adopts a polyphase decomposition structure, integrating fractional delay filtering with downsampling / upsampling operations. At each original sampling point input, one or more interpolation points are output based on the currently required fractional delay. The filter convolution operation utilizes an FPGA. The internal multiply-accumulate array executes in parallel, with each channel configured with an independent filter engine, supporting an input data rate of up to 2.4Gps. The interpolation point calculation uses the formula y(n)=Σ_{k=0}^{N-1}h_k(μ)·x(nk), where h_k(μ) is the filter coefficient dependent on the fractional delay μ, and N is the filter order. After interpolation, the output data rate of each channel remains consistent with the original sampling rate, but each output point undergoes subsampling-level time correction, thereby achieving preliminary fine alignment between channels. During processing, pipelined registers are used to buffer intermediate results to ensure the timing consistency of the multi-channel data stream.
[0049] The residual deviation detection and fine-tuning unit synchronously compares the initially aligned multi-channel data sequences to detect residual minute time offsets between channels. If any are found, the residual deviation is calculated and fed back to the fractional delay filter. After adjusting the compensation coefficients, a second interpolation resampling is performed until the channel deviation converges to a preset threshold. The synchronous comparison employs a cross-correlation method, selecting one channel as a reference (usually consistent with the reference channel of the subsampling phase extraction module). The interpolated data from other channels are then subjected to a sliding correlation with the reference channel data, with a correlation window length of 128 sampling points. The peak position is calculated to obtain the residual time offset. The residual deviation is expressed as a floating-point number with an accuracy of 1 / 256 of the sampling period. If the residual deviation is found to be small, the residual deviation is calculated. If the absolute value of the residual deviation exceeds a preset threshold (e.g., 0.05 sampling periods), it is superimposed with the original decimal deviation to regenerate the compensation coefficient, triggering a second interpolation resampling. The fine-tuning process can be performed iteratively, with the residual deviation re-detected after each iteration until all channel deviations converge to within the threshold or the maximum number of iterations is reached (usually set to 3). To ensure real-time performance, residual deviation detection and coefficient updates are processed in parallel pipeline, meaning that while the current data block is being resampled, the residual deviation of the next data block is being calculated. The maximum time offset between channels in the final output multi-channel synchronous data sequence is suppressed to within 0.05 sampling periods, meeting the requirements of radar countermeasures and high-precision direction finding applications.
[0050] The synchronous verification output unit performs timestamp consistency verification on the final aligned multi-channel data sequence to ensure that all channel sampling points are strictly synchronized; the verified data sequence is packaged according to the original time sequence and outputs a multi-channel synchronous data sequence aligned with subsampling precision.
[0051] The timestamp consistency verification is achieved by comparing the timestamps carried by each channel's data blocks. Each data block is appended with a 48-bit timestamp generated by a global counter, which uses the sampling reference clock as its clock source and increments from zero after system synchronization reset. The verification unit reads the timestamps of each channel's data blocks in parallel and calculates the difference between the maximum and minimum values. If the difference is less than a preset synchronization tolerance (e.g., one sampling clock cycle), it is determined that all channel sampling points are strictly synchronized. If the difference exceeds the tolerance, an alarm is triggered and output is paused. Simultaneously, the reading time of subsequent data blocks is adjusted according to the magnitude of the difference until the synchronization condition is met again. After successful verification... The data from each channel is interleaved and packaged according to the original sampling time sequence to form a unified synchronous data frame. The frame structure includes a frame header, channel identifier, payload, and cyclic redundancy check field. The packaging process is completed in the high-speed serializer inside the FPGA. The output data is transmitted to the subsequent processing module in line-speed mode through a fully domestically produced high-speed interface (such as 10 Gigabit fiber, PCIe, or SRIO). Finally, a multi-channel synchronous data sequence with sub-sampling precision alignment is obtained. The data from each channel at the same sampling time in this sequence are strictly aligned in time, with a phase error of less than 0.05 sampling periods. It can be directly used for joint analysis tasks such as radar signal processing, direction of arrival estimation, and interference cancellation.
[0052] Specifically, by using a reconfigurable fractional delay filter and residual bias iterative fine-tuning technology, the problem of subsampling time mismatch caused by fractional sampling period deviation was solved, and subsampling accuracy alignment with inter-channel phase error of less than 0.05 sampling periods was achieved.
[0053] The hierarchical storage and flow control module writes the sub-sampling precision aligned multi-channel synchronous data sequence into a domestic high-speed cache array according to a preset priority, and uses the FPGA's built-in intelligent flow control engine to dynamically adjust the caching strategy based on the instantaneous data rate and the load of subsequent processing modules to generate controlled parallel data blocks.
[0054] Furthermore, the hierarchical storage and flow control module includes: a priority-based split storage unit, which acquires a multi-channel synchronous data sequence aligned to sub-sampling precision, sorts and groups the data of each channel according to a preset priority to obtain split data groups; writes the split data groups in parallel to a domestic high-speed cache array, dynamically allocates storage space according to priority, realizes hierarchical storage, and obtains a hierarchical storage cache data unit; wherein, the preset priority is predefined according to signal type and application scenario, and the typical priority in radar countermeasure signal processing is set as follows: radar pulse signal has the highest priority (level 0), communication signal is the second highest (level 1), and background noise and interference signal have the lowest priority (level 2); the sorting and grouping is implemented by a priority encoder inside the FPGA, which reads the corresponding priority level from the configuration register according to the channel identifier of each data block, and groups data of the same priority... Data blocks are grouped into the same virtual channel; the domestic high-speed cache array uses DDR4 / 3DRAM chips from Unisplendour or Changxin Memory, and parallel writing is achieved through a multi-port storage controller. Each priority level corresponds to an independent storage area, and the storage space is dynamically allocated according to a preset ratio. For example, 50% of the space is reserved for high-priority areas, 30% for medium-priority areas, and 20% for low-priority areas. When the storage space of a certain priority area is close to saturation, the controller automatically reclaims free space from the low-priority areas and reallocates it through an address remapping mechanism to ensure zero loss of high-priority data. The cache data units of the hierarchical storage are organized in a fixed-length frame format. Each frame contains a frame header (16 bytes), a priority flag (1 byte), a timestamp (8 bytes), a channel number (1 byte), and a payload (maximum 1024 bytes). A 32-bit CRC checksum is appended to the end of the frame.
[0055] The intelligent flow control scheduling unit utilizes the FPGA's built-in intelligent flow control engine to monitor the instantaneous data rate in real time, while simultaneously collecting the load status of subsequent processing modules. It dynamically calculates the scheduling order of each cached data unit and generates the data flow direction after scheduling. Based on the scheduling flow direction, it adaptively adjusts the caching strategy, dividing the cached data units into blocks to generate preliminary controlled parallel blocks. The intelligent flow control engine consists of a state machine, counters, and comparators within the FPGA, monitoring the write and read pointer positions of each priority storage area in real time, and calculating the water level (used space / total space) and data write rate of each area. The load status of the subsequent processing modules is collected through a dedicated interface, including the busy / idle flags of the digital signal processor (such as Phytium D2000 or Fudan Micro ARM), the congestion level of the PCIe link, and the cache status of the fiber optic interface. Occupancy status; the scheduling order calculation adopts a weighted round-robin algorithm, with a scheduling weight of 8 for high-priority channels, 4 for medium-priority channels, and 1 for low-priority channels, ensuring that high-priority data is output first, while avoiding low-priority data from being unable to receive service for a long time; the data flow after scheduling is stored in the queue management unit inside the FPGA in the form of a descriptor chain, and each descriptor contains source address, destination address, data length, and priority information; the cache strategy adjustment includes dynamically changing the burst transmission length according to the water level, for example, when the water level of the high-priority area exceeds 80%, its burst length is increased from 256 bytes to 512 bytes to speed up the data output speed; the initial controlled parallel block is a storage area with contiguous addresses, which is composed of multiple cache data units with the same or similar priorities, and the block size dynamically changes between 1KB and 16KB.
[0056] The parallel block generation output unit obtains the initial controlled parallel blocks. It uses a dynamic load adjustment mechanism to determine whether the data volume exceeds the load limit of subsequent modules. If it does, it performs a second split optimization to obtain optimized parallel data blocks. The optimized parallel data blocks are then formatted and integrated according to preset rules to generate the final controlled parallel data blocks.
[0057] The load limit of the subsequent modules is read from the system configuration register via the AXI-Lite interface. Typical values are: 4KB for PCIeDMA maximum transport block size, 256 bytes for SRIO maximum packet size, and 1500 bytes for 10 Gigabit fiber MTU. The determination is achieved by comparing the parallel block size with the load limit. If the parallel block size exceeds the limit, a split optimization is triggered. The split algorithm uses a greedy strategy to divide the parallel block into multiple sub-blocks not exceeding the limit, and adds a new frame header to each sub-block. The frame header contains the original block identifier and the sub-block sequence number, facilitating reassembly at the receiving end. During the secondary split optimization process... Maintaining the integrity of high-priority data means that high-priority sub-blocks are transmitted first, and cross-block splitting is not allowed; formatting and integration include adding a synchronization header (for receiver clock recovery), a channel mask (indicating which channel data is valid), and a timestamp reference (for time synchronization between multiple blocks) to each final parallel block; the final controlled parallel block is output through a high-speed serial interface that supports the AXI4-Stream protocol, and each block is marked with a TLAST signal to end; the output data stream is isolated from the clock domain of subsequent processing modules through an asynchronous FIFO to ensure the reliability of data transmission at different operating frequencies.
[0058] Specifically, by using priority-based hierarchical storage and intelligent flow control scheduling technology, the problem of rate matching and resource contention between ultra-high-speed multi-channel data streams and subsequent processing modules is solved, realizing hierarchical reliable storage and controlled parallel output of data.
[0059] The high-speed interface synchronous output module reads the controlled parallel data blocks from the domestic high-speed cache array in a time-consistent manner, and transmits them to the subsequent processing module through the fully domestic high-speed interface to obtain an ultra-high sampling rate synchronous data stream.
[0060] Furthermore, the high-speed interface synchronous output module includes: a timing-consistent read unit, which acquires all controlled parallel data blocks from the domestic high-speed cache array, globally sorts them according to the timestamps recorded in each data block, and generates an ordered data block sequence; applies a synchronous read enable signal to the sequence, and simultaneously reads data blocks in parallel from multiple cache channels to obtain a timing-consistent parallel read data stream; wherein, the domestic high-speed cache array uses multi-channel DDR4 SDRAM (such as Hefei Changxin CXDQ3A8AM), each channel corresponds to a priority storage area, the storage controller is embedded in Fudan Micro JFM9RFVU3P FPGA, and supports an AXI4 interface; each data block is appended with 48 bits during writing. The global timestamp is generated by a counter driven by the sampling reference clock, with an accuracy of 1 / 2.4GHz ≈ 0.416ns. Global sorting is accomplished by a hardware-implemented merge sorting network, which compares the timestamps of the head data blocks of each channel buffer in parallel, and selects the data block with the smallest timestamp for output each time, forming an ordered sequence that strictly increases in time. The synchronous read enable signal is generated by the timing control state machine. When the timestamp difference of several consecutive data blocks in the ordered sequence is less than a preset jitter threshold (e.g., 10ns), the state machine issues a parallel read enable and reads data blocks from all buffers involved in the channel at the same time, ensuring that data blocks at the same read time in the output data stream have similar timestamps, thereby achieving timing consistency.
[0061] The parallel interface transmission unit connects the time-consistent parallel read data stream to a fully domestically produced high-speed interface, which includes one or more of PCIe, SRIO, and 10 Gigabit fiber optic interfaces. The data stream is then packetized and encapsulated at the link layer in real time using the interface protocol stack, and transmitted to the subsequent processing module at line speed to obtain a continuously arriving raw transmission data stream. The fully domestically produced high-speed interface is dynamically selected based on system configuration: the PCIe interface uses a domestically produced PCIe controller core (such as the Fudan Microelectronics PCIe hard core), supports Gen3x8, and has a theoretical bandwidth of approximately 64Gbps; the SRIO interface uses a domestically produced SRIOIP core, supports 4x mode, and has a single-channel rate of 5Gbps; the 10 Gigabit fiber optic interface uses a domestically produced 10G Ethernet MAC+PCS / PMA interface, using SFP+... Optical module transmission; the interface protocol stack is implemented in FPGA logic, segmenting the parallel read data stream, adding protocol headers (PCIeTLP header, SRIO packet header or Ethernet MAC frame header), and calculating checksums (CRC32); line-rate mode is implemented through ping-pong buffer and DMA engine, data is directly pushed from the buffer to the interface send FIFO without CPU intervention, ensuring that the transmission rate matches the data generation rate; during transmission, each data packet carries the source channel number, data block sequence number and timestamp information, which facilitates the receiver to restore the original timing; the output data stream is sent out through a high-speed serial transceiver (GTV, rate up to 32.75Gbps), and subsequent processing modules (such as Phytium D2000 server or Loongson 3A5000 computer) receive it through the corresponding interface to obtain the original transmitted data stream.
[0062] The channel alignment correction unit unpacks the original transmitted data stream in the subsequent processing module, restores the data of each channel, and performs channel alignment operations according to the original channel number. It uses timestamp interpolation to detect and correct minor time deviations introduced by differences in the transmission path, obtaining a time-corrected data stream with strict channel alignment. The subsequent processing module uses a Phytium D2000 / 8 (ARMv8 architecture) or Loongson 3C5000 (LoongArch architecture) as the main processor, runs a real-time operating system, and receives data packets via PCIe or SRIO drivers. The unpacking process is completed by software, parsing the protocol header to extract the channel number, data block sequence number, timestamp, and payload, and aligning the data of the same channel in sequence. The data is reassembled into a continuous data stream. Since data from different channels may be transmitted via different physical paths (such as different PCIe channels or different optical fibers), the introduced path delay differences can reach the microsecond level, so timestamp correction is required. The timestamp interpolation method is as follows: using the timestamp of the reference channel (usually channel 0) as a benchmark, the timestamps of other channel data are linearly interpolated. The precise time corresponding to each sampling point is calculated using the timestamps of two adjacent data blocks. Then, the data of each channel is adjusted to a unified time grid through resampling (such as cubic spline interpolation). In the corrected data stream, the data of each channel at the same sampling time are strictly corresponding, with a time deviation of less than 0.1 sampling period. The output is a channel-aligned time-corrected data stream.
[0063] The synchronous data stream generation unit performs a synchronous merging operation on the time-corrected multi-channel data stream, re-interleaving and arranging the data from each channel according to the original sampling time to generate a single ultra-high sampling rate synchronous data stream for subsequent joint analysis tasks such as radar signal processing, interference cancellation, or direction of arrival estimation.
[0064] The synchronous merging operation is performed in the memory of the subsequent processing module. A multi-channel circular buffer is established, with each buffer storing the corrected data of one channel. The merging thread is triggered at a fixed time interval (e.g., 1μs) to read data points from each buffer at the same time and interleave them into a continuous byte stream according to the channel number order (0,1,2,…). The data stream rate after interleaving is equal to the number of channels multiplied by the single-channel sampling rate. For example, 8 channels @ 2.4Gps will result in an output rate of 19.2G samples / second, corresponding to a data rate of approximately 307.2Gbps. To meet the input requirements of subsequent processing algorithms, the data stream can be further encapsulated into a specific format, such as complex IQ data pairs (16-bit I + 16-bit Q) or floating-point format, and sent directly to the GPU (such as the domestic Tianshu Zhixin) or NPU for parallel processing via shared memory or RDMA. The final generated ultra-high sampling rate synchronous data stream maintains the sub-sampling accuracy synchronization relationship between channels and can be directly used for complex algorithms such as high-precision direction finding, digital beamforming, and interference cancellation.
[0065] Specifically, by using timing consistency reading and receiver timestamp interpolation correction technology, the timing misalignment problem caused by path differences when transmitting multi-channel data across high-speed interfaces is solved, ultimately generating an ultra-high sampling rate data stream with strict synchronization between channels.
[0066] Example 2 provides another method to implement a fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform. It employs an adaptive blind calibration technology based on time-interleaved sampling and data-driven methods, achieving channel synchronization through statistical analysis of signal characteristics. Specifically, it includes an optical clock distribution and passive synchronization module. A fully domestically produced low-phase-noise optical carrier generator (such as a domestically produced optical clock source) generates a master clock signal, which is broadcast to each acquisition channel via an optical fiber splitter. Each channel's photoelectric conversion module (a domestically produced PIN photodetector + transimpedance amplifier) converts the optical signal back into an electrical clock, and the clock data is then used to recover the clock signal. The complex circuit extracts the pure clock edge; due to the extremely low temperature sensitivity and inter-channel consistency of fiber optic transmission, the clock signal recovered by each channel naturally has sub-picosecond synchronization accuracy, without the need for dynamic phase adjustment; it solves the skew problem of long-distance transmission of electrical clocks and realizes passive high-precision synchronization of the entire system; the time-interleaved parallel acquisition and adaptive bias correction module uses multiple domestically produced low-resolution high-speed ADCs (such as a certain series of 8-bit ADCs from Fudan Microelectronics) to form an equivalent ultra-high-speed sampling channel in a time-interleaved manner: each ADC alternately samples under different phase drives of the sampling reference clock, synthesizing single-channel ultra-high sampling rate data. The module incorporates an adaptive bias correction circuit that monitors the mean and variance of the output code stream of each ADC in real time. It dynamically adjusts the bias voltage and gain coefficient of each ADC through a digital backend algorithm to eliminate mismatch errors introduced by process differences. The interleaved sampling data is aggregated to the FPGA through a high-speed SerDes interface to obtain the original multi-channel high-speed sampling data sequence. The statistical feature extraction and blind synchronization deviation estimation module uses the statistical computing unit embedded in the FPGA to perform real-time probability density analysis on the sampling data of each channel. In the absence of calibration pulses, the module assumes that the statistical characteristics (such as higher-order cumulants and power spectral density) of the received signals of each channel should be consistent in a statistically average sense.By calculating the cyclic autocorrelation function or fourth-order cumulant of each channel signal, time-shift-sensitive feature parameters are extracted. These feature parameters are compared with the mean features of all channels, and a gradient descent algorithm is used iteratively to solve for integer and fractional multiples of the sampling period deviation that minimize the feature differences. This approach is suitable for signal scenarios with statistically stationary characteristics, such as communication signals and radar pulses. The digital resampling and fractional delay a posteriori compensation module uses a fully digital resampling structure for time alignment based on the channel deviation output from the blind synchronization deviation estimation module. For integer multiple deviations, shifting is performed using a shift register within the FPGA; for fractional multiple deviations, a base-based shifting algorithm is used. Compensation is performed using a variable fractional delay filter with Lagrange interpolation. Unlike Example 1, the filter coefficients in this example are updated online adaptively based on the signal statistical characteristics, rather than relying on fixed calibration pulses. The resampled data stream enters a circular buffer, and phase continuity detection ensures no phase jumps between channels. The sparse representation and compression storage module utilizes the sparse characteristics of the signal in the time-frequency domain to perform compressed sensing processing on the multi-channel data sequence aligned with subsampling precision. The module incorporates a domestically produced sparse basis matrix (such as a discrete cosine transform basis) and reconstruction algorithm to perform real-time sparse decomposition on each channel's data, retaining only significant coefficients and their position indices while discarding redundant noise components. The sparse coefficients are stored in a domestically produced high-speed cache array (such as a hybrid storage of Unisplendour DDR4 + domestically produced NAND Flash) according to signal energy priority, improving storage capacity utilization by 3-5 times. The storage controller dynamically reconstructs the required precision of the data block output based on requests from subsequent processing modules, achieving a balance between storage and transmission bandwidth. The protocol-independent adaptive interface output module uses a protocol-aware hard core embedded in a domestically produced FPGA to detect the interface type and transmission protocol of subsequent processing modules in real time, automatically adapting the data packet format and link layer parameters. The module has a built-in multi-protocol conversion engine that dynamically encapsulates parallel data blocks into protocol frames for the target interface and outputs them through a high-speed serial transceiver. At the same time, the interface module monitors the link status and automatically initiates a retransmission mechanism when packet loss or bit errors are detected, ensuring lossless transmission of the ultra-high sampling rate synchronous data stream.
[0067] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any brief modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform, characterized in that: include: The high-stability clock distribution and calibration module uses a domestically produced high-stability crystal oscillator to generate the master clock signal. The master clock signal is distributed to each acquisition channel through a dedicated clock driver chip, and the clock path deviation is compensated in real time by the dynamic phase adjustment circuit inside the FPGA to obtain a synchronous sampling reference clock for each channel. The synchronous acquisition and pulse injection module drives the domestically produced ADCs of each channel to work synchronously with the sampling reference clock, and periodically injects calibration pulses with known characteristics into the analog input of each channel to obtain the original multi-channel high-speed sampling data sequence containing the calibration pulses; the pulse detection and integer deviation module uses the high-speed parallel processing unit embedded in the FPGA to detect the rising or falling edge position of the calibration pulse of each channel in the original multi-channel high-speed sampling data sequence in real time, and accurately extracts the pulse occurrence time through a multi-stage pipeline comparator to calculate the fixed deviation of the sampling clock of each channel relative to the main clock as an integer multiple of the sampling period. The coarse-grained translation alignment module, based on the fixed deviation of the integer multiple sampling period, dynamically shifts the data stream of each channel by an integer number of sampling periods within the FPGA through a configurable shift register to obtain a preliminary time-aligned multi-channel data sequence; The subsampling phase extraction module extracts waveform details of calibration pulses from the initially time-aligned multi-channel data sequence, calculates the subsampling-level phase difference of each channel pulse relative to the reference channel using FFT interpolation or correlation peak fitting algorithms, and obtains the remaining fractional sampling period deviation. The fractional delay interpolation alignment module uses a reconfigurable fractional delay filter to dynamically interpolate and resample the data of each channel according to the fractional sampling period deviation, obtaining a subsampling precision aligned multi-channel synchronous data sequence. The hierarchical storage and flow control module writes the subsampling precision aligned multi-channel synchronous data sequence into a domestic high-speed cache array according to a preset priority, and uses the FPGA's built-in intelligent flow control engine to dynamically adjust the cache strategy according to the instantaneous data rate and the load of subsequent processing modules, generating controlled parallel data blocks.
2. The domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 1, characterized in that: The high-stability clock distribution and calibration module includes: a master clock generation and distribution unit, which generates a master clock signal using a domestically produced high-stability crystal oscillator, receives the master clock signal through a dedicated clock driver chip, and distributes the master clock signal to each channel according to the physical location and wiring length differences of each acquisition channel to obtain an initial clock path distribution; a phase deviation detection unit, which uses the FPGA's internal clock management module to extract the actual clock signal received by each acquisition channel to obtain the actual arrival phase information of each channel; compares the actual arrival phase of each channel with the master clock reference phase to obtain the phase deviation value of each channel; a dynamic phase adjustment unit, which determines whether the phase deviation value of each channel exceeds a preset threshold range; if it does, it applies a programmable phase shift to the clock path of the corresponding channel through the FPGA's internal dynamic phase adjustment circuit to obtain the adjusted channel clock phase; and an iterative synchronization convergence unit, which repeatedly executes phase deviation detection and dynamic phase adjustment until the phase deviation values of all channels converge to the preset range to obtain a strictly synchronized sampling reference clock for each channel.
3. The domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 1, characterized in that: The synchronous acquisition and pulse injection module includes: a synchronous sampling and pulse injection unit, which drives the domestically produced ADCs of each channel to perform synchronous sampling operations with a sampling reference clock, and periodically injects calibration pulses with known characteristics into the analog input terminal to obtain the original multi-channel high-speed sampling data sequence containing the calibration pulses; a pulse moment extraction and offset detection unit, which detects the periodic occurrence position of the calibration pulses of each channel from the original sampling sequence to generate a pulse moment sequence, calculates the sampling point interval between adjacent pulses and marks the channels with sampling clock offsets, and obtains multi-channel information containing pulse moments and offset marks; a pulse feature analysis and gain correction unit, which extracts the amplitude value and waveform sample of the calibration pulse of each channel based on the multi-channel information to form a pulse feature set, compares the amplitude value with a preset reference amplitude to determine the gain deviation of each channel and generates a gain correction coefficient, and uses the correction coefficient to adjust the amplitude of the effective data segment in the original sequence to obtain the gain-corrected multi-channel sampling sequence; and a channel delay parameter calculation unit, which calculates the fixed delay difference of each channel relative to the reference channel according to the correspondence between the calibration pulse position and the sampling reference clock in the gain-corrected sampling sequence, and generates multi-channel delay correction parameters.
4. The domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 1, characterized in that: The pulse detection and integer deviation module includes: a parallel pulse edge detection unit, which acquires the original multi-channel high-speed sampling data sequence, uses the FPGA embedded high-speed parallel processing unit to synchronously detect the pulse edges of each channel's data, and records the candidate rising edge time and candidate falling edge time by setting rising edge threshold conditions and falling edge threshold conditions respectively, thus obtaining a set of candidate edge times for each channel; a multi-stage pipeline screening unit, which inputs the set of candidate edge times for each channel into a multi-stage pipeline comparator, compares the amplitude change rate of each candidate edge level by level, and selects the edge position with the steepest amplitude change as the precise occurrence time of the calibration pulse for that channel, thus obtaining the precise pulse time value for each channel; an integer multiple deviation calculation unit, which calculates the difference between the precise pulse time of each channel and the master clock reference time to obtain the number of sampling points corresponding to the time difference, and determines the fixed deviation of the sampling clock of each channel relative to the master clock as an integer multiple sampling period by judging that the number of points is an integer multiple sampling period; and a deviation value storage and output unit, which stores the calculated integer multiple fixed deviation value of each channel into the FPGA internal register group to form a deviation lookup table.
5. The domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 1, characterized in that: The coarse-grained translation alignment module includes: an integer deviation parsing and mapping unit, which reads the fixed deviation of each channel's sampling period as an integer multiple from the register group, parses the number of integer translation periods required for each channel, and maps it to the depth parameters of the configurable shift registers inside the FPGA to generate independent shift depth configuration words for each channel; a dynamic depth shift array unit, which dynamically adjusts the depth of each channel's shift register through FPGA logic according to the shift depth configuration words of each channel, sends the original multi-channel data stream in parallel into the shift registers of their respective configured depths, shifts and extracts data step by step under the driving of the sampling clock, realizes dynamic delay compensation for an integer number of sampling periods, and obtains the coarse-grained translation data stream for each channel; and a parallel buffer synchronization output unit, which writes the translation data stream of each channel into the parallel buffer array according to the original timing, applies a unified synchronization read enable signal to all buffer channels to read the data of each channel simultaneously, and obtains a preliminary time-aligned multi-channel data sequence with integer period alignment.
6. The domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 1, characterized in that: The subsampling phase extraction module includes: a calibration pulse waveform extraction unit, which acquires a multi-channel data sequence with preliminary time alignment, extracts waveform sampling points for each channel data according to the time window of the periodic occurrence position of the calibration pulse to obtain the calibration pulse waveform of each channel, and selects one channel as a reference channel to fix its calibration pulse waveform as the reference waveform; a correlation peak fitting unit, which performs cross-correlation operation on the calibration pulse waveform of each channel to be matched and the reference waveform of the reference channel to obtain a cross-correlation function sequence, detects the peak position of the correlation function to obtain the preliminary time offset at the integer sampling point level, performs parabolic fitting on the points near the peak to accurately calculate the subsampling level peak offset, and obtains the subsampling level phase difference of each channel relative to the reference channel; and an FFT interpolation deviation correction unit, which performs FFT transformation on the calibration pulse waveform to obtain the frequency domain phase spectrum, uses the phase spectrum difference to calculate the remaining fractional sampling period deviation between each channel and the reference channel, and obtains the final deviation correction amount.
7. The domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 1, characterized in that: The fractional delay interpolation alignment module includes: a deviation coefficient generation unit, which acquires the fractional sampling period deviation and original sampling period parameters for each channel, adopts a reconfigurable fractional delay filter structure for each channel, and dynamically adjusts the filter tap coefficients according to the deviation to generate fractional delay compensation coefficients adapted to the current deviation; a dynamic interpolation resampling unit, which uses the generated compensation coefficients to perform dynamic interpolation resampling processing on the original data sequence of each channel, and interpolates new sampling points between integer sampling points through filter convolution operations to obtain a multi-channel data sequence with preliminary sub-sampling precision alignment; a residual deviation detection and fine-tuning unit, which performs synchronous comparison on the preliminary aligned multi-channel data sequence to detect residual small time offsets between channels, and if any are found, calculates the residual deviation and feeds it back to the fractional delay filter to adjust the compensation coefficients before performing secondary interpolation resampling until the channel deviations converge to a preset threshold; and a synchronous verification output unit, which performs timestamp consistency verification on the finally aligned multi-channel data sequence to ensure that all channel sampling points are strictly synchronized, and packages the verified data sequence according to the original time sequence to output a multi-channel synchronous data sequence with sub-sampling precision alignment.
8. The domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 1, characterized in that: The hierarchical storage and flow control module includes: a priority-based split storage unit, which acquires multi-channel synchronous data sequences aligned to sub-sampling precision, sorts and groups the data of each channel according to a preset priority to obtain split data groups, writes them in parallel to a domestic high-speed cache array, and dynamically allocates storage space according to priority to achieve hierarchical storage, thus obtaining hierarchical cache data units; an intelligent flow control scheduling unit, which uses the FPGA's built-in intelligent flow control engine to monitor the instantaneous data rate in real time and collect the load status of subsequent processing modules, dynamically calculates the scheduling order of each cache data unit to generate the data flow direction after scheduling, and adaptively adjusts the cache strategy according to the scheduling flow direction to perform block processing on the cache data units to generate preliminary controlled parallel blocks; and a parallel block generation output unit, which acquires the preliminary controlled parallel blocks and combines the load dynamic adjustment mechanism to determine whether the data volume exceeds the load limit of subsequent modules. If it exceeds the limit, it performs secondary splitting optimization to obtain optimized parallel data blocks, and formats and integrates them according to preset rules to generate the final controlled parallel data blocks.
9. The domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 1, characterized in that: Also includes: The high-speed interface synchronous output module reads the controlled parallel data blocks from the domestic high-speed cache array in a time-consistent manner, and transmits them to the subsequent processing module through the fully domestic high-speed interface to obtain an ultra-high sampling rate synchronous data stream.
10. A fully domestically produced ultra-high-speed multi-channel ADC acquisition and processing platform according to claim 9, characterized in that: The high-speed interface synchronous output module includes: a timing-consistent reading unit, which obtains all controlled parallel data blocks from the domestic high-speed cache array, globally sorts them according to the timestamps recorded in each data block to generate an ordered data block sequence, applies a synchronous read enable signal to read data blocks in parallel from multiple cache channels to obtain a timing-consistent parallel read data stream; a parallel interface transmission unit, which connects the timing-consistent parallel read data stream to the fully domestically produced high-speed interface, performs real-time packet assembly and link layer encapsulation of the data stream through the interface protocol stack, and transmits it to the subsequent processing module in line-speed mode to obtain a continuously arriving original transmission data stream; a channel alignment correction unit, which unpacks the original transmission data stream in the subsequent processing module to restore the data of each channel and performs channel alignment operation according to the original channel number, uses timestamp interpolation method to detect and correct small time deviations introduced by transmission path differences, and obtains a time-corrected data stream with strict alignment between channels; and a synchronous data stream generation unit, which performs a synchronous merging operation on the time-corrected multi-channel data stream, and re-interleaves and arranges the data of each channel according to the original sampling time to generate a single ultra-high sampling rate synchronous data stream.
Citation Information
Patent Citations
Calibrating interleaved ADCs mismatch using signal injection
CN105052039A
High-precision multichannel synchronous high-speed data acquisition and processing system and method
CN113162626A
Clock signal dynamic alignment method and phase aligner
CN114465619A
Improved assembly line ADC digital calibration method based on dustpan line
CN118018019A
Self-adaptive multi-dimensional adjustment metering box based on Internet of Things
CN120449090A
Cited By
A method and apparatus for multi-channel data binding and clock compensation based on soft logic
CN122316971A
A method and system for processing multi-channel data synchronously in a TR chip
CN122348809A