FAST core array distributed collaborative observation and data fusion method and system based on RFSOC

Through the distributed collaborative observation and data fusion method based on RFSOC, the synchronization control, resource allocation and heterogeneous computing problems of traditional radio astronomical equipment are solved, efficient signal processing and observation efficiency are improved, and multi-beam synthesis and cross-regional observation are supported.

CN120743565AActive Publication Date: 2025-10-03NAT ASTRONOMICAL OBSERVATORIES CHINESE ACAD OF SCI +1

Patent Information

Application Number
CN202511241870.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-03
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Traditional radio astronomy observation equipment has problems such as low distributed collaboration efficiency, rigid resource allocation, fragmented heterogeneous computing, and poor cross-system compatibility, which leads to delayed response of fast radio burst signals and low observation efficiency, and cannot meet the needs of multi-band transient source monitoring.

Method used

It adopts a distributed collaborative observation and data fusion method based on RFSOC, and realizes nanosecond-level synchronization control, intelligent resource scheduling, heterogeneous computing acceleration, cross-domain data fusion and elastic interface adaptation through seven core functional modules, including signal acquisition and preprocessing, intelligent resource scheduling, heterogeneous computing acceleration, cross-domain data fusion, collaborative observation planning and data output, using technologies such as echo state neural network, improved ant colony algorithm, graph neural network and elastic interface module.

Benefits of technology

It achieves nanosecond-level synchronous control and dynamic collaboration of heterogeneous resources, improves signal superposition efficiency to 92% and resource utilization to 75%, supports multi-beam synthesis, cross-regional joint observation and real-time processing of massive data, and breaks through the performance bottleneck of traditional architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743565A_ABST
    Figure CN120743565A_ABST
Patent Text Reader

Abstract

The invention discloses an RFSOC-based FAST core array distributed collaborative observation and data fusion method and system, and the method comprises the steps: S1, system initialization: a master node RFSOC generates a global clock, achieves the phase synchronization of multiple board cards through an SYSREF differential signal, and completes the calibration of a three-stage clock tree; s2, signal acquisition and preprocessing: directly sampling a 3-8GHz radio frequency signal through an ADC (Analog to Digital Converter), and performing digital down-conversion to obtain a baseband signal; s3, intelligent resource scheduling: identifying a signal type based on an ESN (Echo State Neural Network), and dynamically allocating FPGA logic resources through an improved ant colony algorithm; s4, heterogeneous calculation acceleration: executing a 128-channel digital beam forming pipeline on the FPGA; s5, cross-domain data fusion; s6, collaborative observation planning; and S7, outputting data. The method is suitable for multi-beam synthesis, cross-region joint observation and mass data real-time processing scenes, and provides key technical support for solving the frontier scientific problems such as rapid radio storm origin and black hole activity monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radio astronomy observation technology, and in particular to a FAST core array distributed collaborative observation and data fusion method and system based on RFSOC. Background Art

[0002] In the field of radio astronomy, the FAST core array digital terminal system serves as the core equipment for high-precision radio signal acquisition, processing, and coordinated observation. Its performance directly determines the ability to capture cosmic transient phenomena (such as fast radio bursts and single pulses from pulsars). However, traditional architectures, such as the currently operational ROACH2 platform, face multiple technical bottlenecks: low distributed coordination efficiency (microsecond-level synchronization accuracy results in less than 70% efficiency in signal superposition for long-term observations), rigid resource allocation (fixed-function unit design results in FPGA resource utilization consistently below 40%), heterogeneous computing fragmentation (data exchange between the ADC analog front-end and the FPGA / ARM processing unit experiences a 120ns latency), and poor cross-system compatibility (supporting only single-telescope independent observations and incompatible with the SKA intermediate-frequency array's SDPv3 protocol and the VLBI Global Network's Mark 5B format). These problems are particularly prominent in the real-time tracking of fast radio bursts (FRBs). The traditional architecture has a response delay of more than 1.2 seconds to burst signals, resulting in only one effective sampling within the FRB millisecond period, which seriously restricts the analysis of the signal's fine time domain structure. In multi-band transient source monitoring, the fixed-band processing unit cannot dynamically allocate resources, resulting in less than 32 channels for observations in high-frequency bands above 2 GHz, and the probability of missed detection exceeds 25%.

[0003] The core problem lies in the fact that traditional designs fail to fully leverage the mixed-signal processing capabilities of FPGA chips. Their integrated 12-bit, 1.5GSPS ADC / DAC, 28nm FPGA logic (including 500k logic cells), and Cortex-A53 quad-core ARM processors could form an integrated architecture encompassing analog front-end, digital processing, and intelligent control. However, existing solutions utilize these chips solely as independent functional modules, failing to achieve nanosecond-level synchronization control and dynamic coordination of heterogeneous resources. For example, the ROACH2 platform uses external GPS timing, resulting in a synchronization error of ±500ns. However, the RFSOC's built-in phase-locked loop (PLL) combined with an onboard crystal oscillator (frequency stability of ±0.01ppm) theoretically improves synchronization accuracy to ±100ps. Furthermore, the traditional architecture's fixed pipeline design cannot adapt to the dynamic task switching requirements of "single pulse search, beamforming, and spectrum analysis" in radio observations, resulting in observation mode switching times exceeding 30 seconds. However, the RFSOC-based reconfiguration technology can complete hardware logic reconfiguration within 2ms. Summary of the Invention

[0004] To address the challenges of the existing technology, the present invention aims to provide a distributed collaborative observation and data fusion method for the FAST core array based on RFSOC. Through seven core functional modules, this method enables nanosecond-level synchronization control, intelligent resource scheduling, heterogeneous computing acceleration, cross-domain data fusion, collaborative observation planning, and flexible interface adaptation. The method is suitable for scenarios involving multi-beam synthesis, cross-regional joint observation, and real-time processing of massive amounts of data. Another object of the present invention is to provide a distributed collaborative observation and data fusion system for the FAST core array based on RFSOC that implements the above method.

[0005] To achieve the above objectives, the present invention provides a distributed collaborative observation and data fusion method for the FAST core array based on RFSOC, comprising the following steps: S1. System initialization: The master node RFSOC generates a global clock, synchronizes multiple boards using the SYSREF differential signal, and completes three-level clock tree calibration. S2. Signal acquisition and preprocessing: Directly sample the 3-8 GHz RF signal using an ADC and digitally down-convert it to baseband. S3. Intelligent Resource Scheduling: Using the Echo State Neural Network (ESN) to identify signal types, an improved ant colony algorithm dynamically allocates FPGA logic resources. S4. Heterogeneous Computing Acceleration: Executing a 128-channel digital beamforming pipeline on an FPGA. S5. Cross-domain data fusion: Delay compensation, gain equalization, and frequency-domain layered fusion of multi-board data. S6. Collaborative Observation Planning: Allocating Multi-Telescope Observation Resources Based on Graph Neural Networks (GATs); S7. Data output: The data is converted into SKA SDPv3, VLBI Mark5B, or a custom protocol format via a flexible interface module for data output.

[0006] Furthermore, the phase synchronization includes: a dynamic phase calibration algorithm to monitor the crystal oscillator frequency drift in real time and compensate for the phase accuracy up to λ / 16@2GHz; The event trigger signal is generated by the ARM core and synchronized to all processing nodes through the AXI bus, completing the three-level clock tree calibration of global clock → regional clock → logic unit clock, with a synchronization error of <500ps.

[0007] Furthermore, the intelligent resource scheduling includes: The ESN neural network contains a sparse randomly connected pool of 1000 neurons, outputs a resource allocation weight matrix, and the objective function is: ; in, To calculate the delay, is the resource utilization rate, is the power consumption; α, and γ are the corresponding weight coefficients, and their values ​​are dynamically and adaptively adjusted according to the requirements of the observation task; The incremental learning mechanism updates the ESN weight every time 200 new samples are added.

[0008] Furthermore, the heterogeneous computing acceleration includes: The ARM core dynamically adjusts the FPGA frequency through DVFS technology and monitors the FPGA temperature, logic unit occupancy, and DDR4 bandwidth in real time. Beamforming uses a double-precision floating-point arithmetic unit, with a phase error of ≤λ / 32@2GHz.

[0009] Furthermore, the cross-domain data fusion includes: Cross-band spectrum splicing uses multi-resolution wavelet transform; Radio frequency interference suppression uses a blind source separation algorithm, with an interference suppression ratio of 40dB.

[0010] Furthermore, the collaborative observation plan includes: Construct a multi-layer heterogeneous resource graph G = (V, E), where node attributes include device parameters, state parameters, and observation target parameters; The GAT model uses an 8-head attention mechanism and outputs the task assignment matrix A∈R N×M , N is the number of nodes, M is the number of tasks, and the constraints of single-node task number ≤ 4, frequency band isolation, and transmission delay < 50μs are met.

[0011] Furthermore, in S6, multi-telescope observation resources are allocated based on a graph neural network (GAT), and tasks are replanned within 100ms in emergencies using a PPO reinforcement learning algorithm. Emergency response includes: When an FRB signal with an SNR greater than 7dB is detected, 16 high-sensitivity channels are prioritized, with a signal capture delay of less than 200ms. The radio and optical telescopes are triggered synchronously in the time domain, and the alignment error between exposure time and integration time is <10ms.

[0012] Furthermore, the data output includes: Add a CRC-32 check module at the end of the FPGA pipeline to improve the bit error rate monitoring accuracy. , data reliability>99.999%; Supports clock synchronization and data stream interaction of the SKA medium frequency array and VLA telescope.

[0013] Furthermore, the system initialization phase also includes: Load the lightweight GAT model (the original GAT model has approximately 1.5 million parameters), pruned to 9.8MB (compressed 15 times), and the ESN model into the FPGA; The federated learning architecture aggregates model parameters updated locally at each node, including the ESN's pool connection weights and output layer weights, as well as the GAT's attention weights and graph convolution layer weights. These parameters are updated through a distributed mechanism to optimize signal recognition and resource allocation efficiency, while meeting the technical requirements of nanosecond synchronization and dynamic collaboration of heterogeneous resources.

[0014] The FAST core array distributed collaborative observation and data fusion system based on RFSOC includes the following modules: Nanosecond synchronization control module: The clock management unit integrated into the RFSOC generates a 100MHz±0.1ppm global clock and achieves multi-board synchronization through the SYSREF differential signal and a three-level clock buffer architecture. Intelligent resource scheduling module: The ESN neural network and improved ant colony algorithm deployed on the ARM core build a 3×5 feature-resource association matrix to dynamically allocate FPGA logic resources.

[0015] Heterogeneous computing acceleration module: A 128-channel digital beamforming pipeline on the FPGA, interacting with the ARM core through a ping-pong cache architecture; Cross-domain data fusion module: realizes spatiotemporal calibration and multi-level fusion of FAST and external telescope data; Collaborative Observation Planning Module: A task allocation engine based on the GAT model and a PPO reinforcement learning emergency response engine, supporting real-time re-planning of emergency events; Flexible interface adapter module: including 16-channel 3-8GHz RF input at the front end, dual 10G SFP+ interfaces and PCIe4.0 interfaces at the back end; External interface interaction module: supports PTP clock synchronization, SDP protocol data transmission and standard API interface.

[0016] The beneficial effects of the present invention are as follows: Through seven modular architecture designs (high-precision synchronization module, dynamic resource scheduling module, heterogeneous collaborative computing module, etc.), this invention has for the first time realized nanosecond-level distributed synchronization control (synchronization error <200ps), dynamic reconstruction of heterogeneous resources (function switching delay <5ms) and multi-protocol elastic adaptation (supporting six international standard protocols such as SKA and VLBI), increasing the signal superposition efficiency of long-term phase observation to more than 92%, and the resource utilization rate to 75%. It also supports real-time collaborative observations of FAST with the SKA medium-frequency array and the VLBI global network, providing key technical support for solving cutting-edge scientific problems such as the origin of fast radio bursts and monitoring black hole activities. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is the overall system architecture diagram; Figure 2 This is the timing diagram of the nanosecond-level synchronization control module; Figure 3 This is the hardware architecture diagram of the heterogeneous computing acceleration module; Figure 4 The performance comparison between the present invention and the traditional solution is shown in the actual measurement. Figure 5 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solution of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0020] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0021] The following combination Figure 1-Figure 5 The specific embodiments of the present invention are described in detail. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0022] The present invention is based on the FAST core array distributed collaborative observation and data fusion method and system of RFSOC, specifically a 500-meter Aperture Spherical Radio Telescope (FAST) core array distributed collaborative observation system and method based on radio frequency system on chip (RFSOC). Through seven core functional modules, it realizes nanosecond-level synchronization control, intelligent resource scheduling, heterogeneous computing acceleration, cross-domain data fusion, collaborative observation planning and flexible interface adaptation, and is suitable for multi-beam synthesis, cross-regional joint observation and real-time processing of massive data scenarios.

[0023] The distributed collaborative observation and data fusion method of FAST core array based on RFSOC, such as Figure 5 As shown, the following steps are included: S1. System initialization: The master node RFSOC generates a global clock, synchronizes multiple boards using the SYSREF differential signal, and completes three-level clock tree calibration. S2. Signal acquisition and preprocessing: Directly sample the 3-8 GHz RF signal using an ADC and digitally down-convert it to baseband. S3. Intelligent Resource Scheduling: Using the Echo State Neural Network (ESN) to identify signal types, an improved ant colony algorithm dynamically allocates FPGA logic resources. S4. Heterogeneous Computing Acceleration: Executing a 128-channel digital beamforming pipeline on an FPGA. S5. Cross-domain data fusion: Delay compensation, gain equalization, and frequency-domain layered fusion of multi-board data. S6. Collaborative Observation Planning: Allocating Multi-Telescope Observation Resources Based on Graph Neural Networks (GATs); S7. Data output: The data is converted into SKA SDPv3, VLBI Mark5B, or a custom protocol format via a flexible interface module for data output.

[0024] The present invention specifically adopts the following module architecture design and processing flow: 1. Core functional module architecture design, such as Figure 1 As shown: Module 1: Nanosecond synchronization control module, such as Figure 2 As shown, a high-precision synchronization benchmark is constructed for the entire system to support multi-node coherent synthesis and cross-telescope time slot alignment.

[0025] A hardware-level synchronization benchmark is employed: The RFSOC's internal clock management unit (CMT) generates a 100MHz ±0.1ppm global reference clock. The SYSREF differential signal is used to achieve multi-board phase synchronization, with a synchronization error of <500ps (differential signals suppress common-mode noise, resulting in a transmission delay deviation of <200ps). A three-level clock buffer architecture (global clock → regional clock → logic unit clock) is combined with a dynamic phase calibration algorithm (crystal oscillator frequency drift monitoring period of 10ms, phase compensation accuracy of λ / 16 at 2GHz). Crystal oscillator drift compensation is based on an adaptive Kalman filter framework, integrated with the FPGA hardware CMT unit design. The specific formula is as follows: Global clock drift compensation: ; in: (k) is the estimated value of the global clock frequency at the kth iteration, (k) is the current crystal oscillator frequency measurement value (obtained through PLL feedback), and α is the adaptive gain coefficient (dynamically adjusted according to the 10ms monitoring period).

[0026] Regional transfer delay compensation: ; in, is the SYSREF differential signal transmission delay deviation, c is the speed of light, is the regional delay compensation weight (calculated based on the phase accuracy of λ / 16@2GHz, ), The maximum allowable propagation delay deviation of the SYSREF differential signal (set to ≤ 200 ps) is used to ensure that the regional delay compensation weight β meets the phase accuracy constraint of λ / 16 (when the actual delay deviation ≤ When , the phase compensation error is ≤λ / 16).

[0027] Logic unit phase error correction: ; in, (k) is the actual phase error of the logic cell (measured by the phase detector), and γ is the logic cell correction factor (capacitor charge and discharge time in TSMC 16nm process). is the phase error of the logic unit after correction at the kth iteration, and its value is obtained by the previous error The current correction value (based on the deviation between the target accuracy λ / 16 and the measured error) is dynamically updated, and the phase error of the logic unit is finally converged to within λ / 16.

[0028] The coordinated total formula of the three-level calibration, the phase error transfer relationship between global, regional and logic units is: ; in, (k) is the total phase error of the system at the kth iteration, reflecting the comprehensive phase deviation of the entire clock tree after the three-level calibration; (k) is the phase error of the global clock at the kth iteration, which is caused by factors such as crystal oscillator frequency drift and temperature drift, and is dynamically corrected using the global clock drift compensation formula; (k) is the phase error of regional transmission at the kth iteration, which is converted from the delay deviation of the SYSREF signal in regional-level transmission; is the phase error of the logic unit after correction. By iteratively updating α, , γ coefficient, to achieve dynamic calibration.

[0029] Event triggering and time slot alignment: The ARM core generates nanosecond-level trigger signals (with a timestamp resolution of 1ns, including metadata such as task ID and integration time), which are synchronized to all processing nodes via the AXI bus. This supports the observation time slot alignment accuracy of <10ns for telescopes such as FAST and SKA, meeting the phase stability requirements of a coherent integration time of >1 hour.

[0030] Innovative advantages: Eliminates dependence on external synchronization modules, improves phase synchronization accuracy by two orders of magnitude compared to traditional solutions, and supports signal superposition for long-term phase observation and measurement.

[0031] Module 2: Intelligent Resource Scheduling Module dynamically allocates computing resources based on signal characteristics, enabling real-time switching between multiple observation modes. The ARM core deploys an Echo State Neural Network (ESN) to extract signal features and dynamically allocate computing resources. The 1GSps baseband signal is converted to a time-frequency spectrum matrix using STFT and then fed into a sparse randomly connected pool of 1000 neurons (sparseness 0.1, spectral radius 0.95). Linear regression training is used to output the weight matrix, with the objective function: ; in, To calculate the delay, is the resource utilization rate, The power consumption is represented by α, β, and γ, respectively, and the corresponding weight coefficients are determined by satisfying α + β + γ = 1. The weight coefficients are dynamically adaptive, achieving an accuracy rate of ≥ 98.2% for the recognition of three types of signals: pulsars, continuous spectra, and spectral lines. An incremental learning mechanism is introduced (weight updates are triggered every 200 new samples), supporting online adaptive adjustment of the model.

[0032] Multi-dimensional resource scheduling model: An improved ant colony algorithm is combined with the dynamic feature vectors output by the ESN to construct a 3×5 feature-resource association matrix. This optimizes the pheromone update strategy, implements differentiated resource allocation for different signal types, and enables dynamic allocation of FPGA logic resources (FFT cores, multipliers, and BRAMs) (for example, pulsar signals are allocated a 16k-point FFT core + 8-bit fixed-point arithmetic unit).

[0033] Dynamic channel switching mechanism: Real-time switching between eight observation modes of a single card is achieved through the RFSOC internal AXI-Stream interface. A dynamic FIFO buffer (depth 1024) is configured to avoid data congestion, and data routing delay is less than 1μs.

[0034] Among them, the traditional ant colony algorithm is suitable for static path optimization, but has the following shortcomings in FPGA dynamic resource scheduling: it cannot adapt to real-time changes in signal characteristics (such as the burst characteristics of FRB), it is difficult to balance convergence speed and global optimization with a fixed pheromone volatility coefficient, and it lacks direct mapping with hardware resource characteristics (such as logic units and BRAM). The innovative design of the improved ant colony algorithm: Feature-resource association matrix: Construct a 3×5 dimensional matrix (3 types of signal features × 5 types of resource types) to replace the traditional path diagram. This is achieved by improving the transition probability formula. The traditional transition probability calculation formula is: ; The improved transition probability formula combining signal characteristics and hardware status is: ; in, is the suitability of signal type s to resource j in the feature-resource matrix (e.g., the suitability of pulsar to 16k FFT kernel is 0.9); is the hardware load impact factor ( is the coefficient, is the current resource occupancy rate); Z is the normalization factor.

[0035] Dynamic pheromone update strategy: Adaptively adjust the volatility coefficient based on ESN signal recognition results and hardware status; Multi-objective joint optimization: Integrate computing latency, resource utilization, and power consumption. The calculation formula for the adaptive pheromone update strategy is: ; in, for The pheromone concentration from resource node i to resource node j at the iteration (i, j are resource type indices, such as i = FFT core, j = multiplier), is the pheromone increment from resource node i to j (the pheromone concentration added in this iteration), The pheromone volatility coefficient is dynamically adjusted over time (controls the decay rate of pheromones, the higher the value, the faster the decay). , dynamically adjusted with temperature, is the current FPGA temperature, is the volatility coefficient, the reference value =0.1, is the temperature influence coefficient, the reference value is =0.05; pheromone increment , where Q is a constant, is the path length of the kth ant (resource allocation cost), .

[0036] Innovative advantages: Resource utilization is increased by 60%, power consumption in typical scenarios is reduced by 40%, and dynamic reconstruction of hardware resources is supported for multi-target parallel observation.

[0037] Module 3: Heterogeneous computing acceleration module, such as Figure 3 As shown in the figure, a collaborative computing architecture of FPGA hardware acceleration and ARM intelligent scheduling is constructed to improve signal processing efficiency.

[0038] FPGA hardware-accelerated pipeline: Implements a 128-channel digital beamforming (DBF) pipeline with a processing rate of 2GSps and supports real-time operations in the complex domain with 16-bit precision (≥32 parallel multiplication and accumulation units). The pipeline includes an automatic gain control (AGC) module with a dynamic range of 100dB and supports direct sampling of 3-8GHz broadband signals (5GSps @ 14-bit ADC).

[0039] ARM intelligent scheduling and power consumption optimization: Real-time collection of FPGA temperature (accuracy ±1°C), logic unit occupancy (resolution 1%), and DDR4 bandwidth (sampling interval 5ms), dynamically adjusting the FPGA frequency (100-250MHz, in 10MHz steps) using DVFS technology. Power consumption is reduced to 50W at low loads (compared to ≥150W with traditional solutions) and maintained at 150W at full load.

[0040] Software and hardware collaborative interaction: Shared memory communication is achieved through the AXI-HP interface, supporting 8GB / s data throughput. The interaction delay between pre-processed data (such as beamforming baseband signals) and post-processing algorithms (such as general CLEAN imaging and pulsar period search algorithms) is less than 200ns.

[0041] Innovative advantages: End-to-end processing latency is reduced from 50μs to 15μs, and heterogeneous computing resource utilization is increased to 85%, breaking through the data interaction bottleneck of the traditional software and hardware separation architecture.

[0042] FPGA pipeline optimization: A double-precision floating-point arithmetic unit (supporting the IEEE 754-2019 standard) is used to dynamically update beamforming weights. Compared with traditional fixed-point arithmetic, the phase error is reduced to λ / 32@2GHz, and the synthesized beam sidelobe suppression ratio is improved by 10dB.

[0043] Heterogeneous computing collaboration mechanism: A ping-pong cache architecture (dual-port BRAM depth 64k) is designed to enable non-blocking data interaction between the FPGA pipeline and the ARM processor. Compared with traditional FIFO solutions, the throughput is increased to 16GB / s, and the interaction latency is stabilized within 100ns.

[0044] Module 4: Cross-domain Data Fusion Module implements spatiotemporal alignment and multi-level fusion of FAST and external telescope data, improving joint observation accuracy. This module utilizes a multi-source data spatiotemporal alignment design. For delay compensation, an FFT-based cross-correlation algorithm (calculation period <1μs) achieves compensation accuracy of 1 / 8 the sampling period (corresponding to a delay of <0.2ns at 5Gsps). For gain equalization, an adaptive filter with a dynamic range of 120dB (coefficient update frequency 10kHz) eliminates inter-channel gain differences (gain deviation <0.1dB after calibration). Regarding the layered fusion strategy, the time domain layer implements sliding window weighted fusion (with a dynamically adjustable window length from 100ns to 1μs and a Hanning window to reduce spectral leakage), while the frequency domain layer implements Wiener filtering for spectral line enhancement (improving noise suppression ratio by 25dB and real-time noise estimation and updating of filter coefficients). This design supports data fusion with telescopes from different systems, such as the SKA, enabling seamless integration of observation data across frequency bands (1GHz-8GHz) and across regions.

[0045] Cross-band fusion and expansion: Supports data fusion of FAST (1-3 GHz) and SKA medium-frequency array (3-8 GHz), realizes cross-band spectrum stitching through multi-resolution wavelet transform (5 decomposition layers), expands the joint imaging frequency coverage range to 1 GHz-8 GHz, and improves the resolution to 0.1 arc second (equivalent to a 10 km baseline interferometer).

[0046] Real-time interference suppression: The blind source separation algorithm (JADE algorithm) is introduced to separate RFI (radio frequency interference) in the time domain, achieving an interference suppression ratio of 40dB, significantly improving the detection capability of weak signals (such as pulsar single pulses).

[0047] Module 5: Collaborative Observation Planning Module, based on graph neural networks, achieves global optimization of multi-telescope resource allocation and supports real-time re-planning for emergencies. Observation resource graph modeling is performed, constructing a multi-layer heterogeneous resource graph G = (V, E). Node V includes 19 FAST feeds, 133 SKA antennas, and VLBI stations (with 32 attributes: device parameters, state parameters, and observation target parameters). This module implements a task allocation optimization engine and constructs a Graph Attention Network (GAT) with an 8-head attention mechanism, an embedding encoding dimension of 256, and a LeakyReLU activation function to capture node collaboration relationships (such as frequency band complementarity). The output is the task allocation matrix A∈R. N×M N is the number of nodes, and M is the number of tasks. The constraints of number of tasks per node ≤ 4, frequency band isolation, and transmission delay < 50 μs are met (integer linear programming is used to verify feasibility). This design implements a real-time replanning mechanism. When an emergency event (such as FRB, with a trigger threshold SNR > 7 dB) is detected, the PPO reinforcement learning algorithm (50-dimensional state space, task adjustment operation in the action space) is activated, completing resource reallocation and dynamically inserting priority tasks within 100 ms.

[0048] Emergency response performance test: In FRB simulation tests, the system's burst signal capture delay was 180ms, a 6.6-fold improvement compared to the ROACH2 platform (1.2s delay), ensuring at least three valid samples within the signal period.

[0049] Innovative advantages: The efficiency of multi-telescope joint observation has been increased by 50%, the resource conflict rate has been reduced from 15% to 2%, and time-domain and frequency-domain collaborative observations of radio / optical telescopes (such as LAMOST and JWST) have been realized for the first time.

[0050] Module 6: Flexible Interface Adapter Module implements modular hardware interface design and multi-protocol adaptation, supporting system expansion and cross-platform compatibility. It utilizes a modular hardware interface design, with 16 3-8 GHz RF inputs (SMA connectors with integrated anti-aliasing filters and a roll-off factor of 0.22) on the front end, a single-channel 5 GHz @ 14-bit ADC sampling, and dual 10 Gigabit SFP+ interfaces (supporting UDP / IP and SDP protocols, with a link aggregation bandwidth of 20 Gbps) and a 12 Gbps PCIe 4.0 interface on the back end, supporting cascading expansion up to 256 nodes (clock skew < 500 ps).

[0051] Protocol Adaptation Engine: The ARM core runs multi-protocol conversion middleware (format parsing / recombination / verification modules), supporting real-time conversion between FAST's custom format and the SKA standard format (latency < 5μs). It is compatible with the data interaction protocols of equipment such as the China VLBI Network and ALMA, and supports plug-and-play modular deployment.

[0052] Innovative advantages: System scalability is increased by 3 times, hardware compatibility covers 90% of mainstream radio equipment, and protocol conversion delay is reduced by 70% compared with traditional solutions. Figure 4 shown.

[0053] Module 7: External interface interaction module, which realizes standardized interaction between the system and external observation equipment and back-end processing platforms, and builds an open observation network.

[0054] Front-end signal access: Directly connected to the 19-unit feed network of the FAST core array, supporting 3-8 GHz full-band signal input, integrated feed noise suppression pre-processing module (noise figure <2 dB); compatible with feed polarization signal processing (left-handed / right-handed independent channels, polarization degree measurement accuracy ±1%).

[0055] Cross-telescope collaborative interface: supports clock synchronization (PTP protocol, synchronization accuracy <10ns) and data stream interaction (SDP protocol, transmission rate 10Gbps) with international telescopes such as the SKA medium-frequency array (133 units) and the VLA; provides a standard API interface for third-party devices to call, and supports remote configuration of observation parameters (latency <10ms).

[0056] Back-end data output: The 10G Ethernet port outputs the processed baseband signal (supports multicast mode, single-node bandwidth 4Gbps); the PCIe interface connects to the back-end correlation machine (such as the FAST pulsar search engine), with a data throughput of 12Gbps, meeting the requirements of real-time correlation processing (integration period <1ms).

[0057] Innovative advantages: Building a standardized interface system, supporting seamless access of multi-brand equipment, and promoting the openness and internationalization of radio astronomy observation networks.

[0058] 2. Processing Flow 1. System initialization phase: Synchronous network construction, master node RFSOC generates a global clock, slave nodes lock phases through SYSREF, and complete three-level clock tree calibration (<100ms); resource graph initialization, collecting device parameters (bandwidth, accuracy) and status (load rate, calibration status) of each node, and constructing the initial observation resource graph G0; model loading, FPGA loads the GAT task allocation model (9.8MB) and ESN signal recognition model (parameter size 3M), and the ARM core starts the resource monitoring service.

[0059] 2. Routine observation processing flow: For signal acquisition and preprocessing, the front-end ADC directly samples the 3-8 GHz signal (5 GSps), digitally down-converts it to baseband (I / Q two-way, 1 GSps), and inputs it into the FPGA to execute the DBF pipeline. The FPGA calculates the weighting coefficients of 128 channels in real time and outputs the synthesized high signal-to-noise ratio signal (SNR>15dB).

[0060] Intelligent scheduling and data fusion: The ARM core identifies the signal type and triggers the DDRP protocol to allocate resources (for example, a 4k-point FFT core + 4-bit quantization unit is allocated to a neutral hydrogen signal). The fusion module performs delay compensation, gain equalization, and three-layer fusion on multi-board data, and outputs feature data (such as spectral line profile and phase difference matrix).

[0061] Data output and storage are converted into a standard format through a flexible interface module and transmitted to the back-end related machine (latency < 5μs) or storage system (supporting RAID5, write speed 2GB / s) through a 10G network.

[0062] Data integrity check: Add CRC-32 check module at the end of FPGA pipeline to perform real-time check on 128-channel beamforming data, and the bit error rate monitoring accuracy reaches , ensuring that the reliability of data transmitted to the backend is >99.999%.

[0063] 3. Collaborative Observation Planning Process: Task reception and allocation: The master control node receives the observation target parameters (right ascension, declination, bandwidth), and the GAT model generates the task allocation matrix A to ensure that the single node load is ≤3 and there is no frequency band overlap; configuration instructions (including FFT kernel parameters, beam weights, and data routing paths) are issued to each RFSOC node.

[0064] In emergency response, when an FRB signal is detected, the PPO algorithm prioritizes scheduling 16 high-sensitivity channels (turning off low-priority tasks), with a rescheduling time of 95ms and a signal capture delay of <200ms; the resource map G is updated in real time, and emergency task processing logs are recorded (timestamp accuracy 1ns).

[0065] Radio-optical joint observations connect to the real-time data stream of the LAMOST optical telescope, achieve time-domain synchronous triggering (alignment error between optical exposure time and radio integration time <10ms) through a collaborative observation planning module, and support joint analysis of multi-band energy spectra of transient sources.

[0066] The present invention implements technical functions based on the following hardware platforms and software systems: 1. Hardware Platform Implementation 1. Core Board Design: Utilizes the Xilinx Zynq UltraScale+ RFSoC 4x2 device, integrating a 4-channel 14-bit ADC (5Gsps), a 2-channel 16-bit DAC (6.5Gsps), 3.4M system gate FPGA logic, and a quad-core Cortex-A53 ARM (1.5GHz). The board measures 160mm x 100mm, and the power supply design supports a wide input voltage (9-15V). It operates in a temperature range of -40°C to +85°C.

[0067] 2. System cascade solution: Synchronize over 100 boards using the SYSREF differential signal (maximum transmission distance 10m, signal attenuation <3dB). The 10G Ethernet port is configured with the TSN protocol (Time-Sensitive Network), with transmission delay jitter <10ns, supporting real-time aggregation of multi-board data.

[0068] 3. Actual synchronization performance: In a cascade test of 100 boards, the maximum synchronization error was 480ps and the root mean square error was 210ps, meeting the SKA's stringent requirements for distributed array synchronization accuracy (<500ps).

[0069] 4. Power consumption optimization verification: In pulsar search mode, the power consumption of a single board is as low as 45W (compared to 180W for the traditional ROACH2 platform). The temperature rise caused by heat accumulation during 8 hours of continuous observation is only 6°C, without the need for an additional liquid cooling system.

[0070] 2. Software System Implementation 1. Synchronization control software: Master node synchronization engine: CMT unit configuration, SYSREF signal generation, dynamic phase calibration algorithm (implemented in C language, running on ARM core 0); Slave node synchronization agent: PLL locking algorithm, phase error compensation (implemented in Verilog, running on FPGA logic).

[0071] 2. Resource Scheduling Software: CNN signal recognition model: Based on the TensorFlow Lite framework, quantized to 8-bit fixed-point operations, with inference latency of <1ms; Improved Ant Colony Algorithm: MATLAB modeling optimization and porting to ARM core, with a resource allocation cycle of 5ms.

[0072] 3. Collaborative Observation Planning Software: GAT model lightweight: the pruned model size is 9.8MB, FPGA hardware-accelerated inference (convolutional layer calculation speed increased by 20 times); PPO reinforcement learning engine: experience replay buffer capacity of 100,000, training update cycle of 100ms, running on ARM cores 1-3 (multi-core parallel).

[0073] 4. Distributed training support: The collaborative observation planning module supports the federated learning architecture. Each node can update the GAT model parameters locally and aggregate the global model through encrypted communication, protecting the privacy of observation data while improving the model's generalization ability.

[0074] Technical advantages of the present invention: Full-link high-precision synchronization: Nanosecond-level phase synchronization technology supports long-term phase observation and measurement, with a phase coherence synthesis efficiency exceeding 95%; Intelligent dynamic resource management: Deep learning-based signal recognition and ant colony algorithm scheduling improve resource utilization by 60% and reduce power consumption by 40%; Deep collaboration of heterogeneous computing: FPGA hardware acceleration and ARM intelligent scheduling combine to reduce end-to-end processing latency to 15μs, breaking through the performance bottleneck of traditional architectures. Cross-domain collaborative observation capability: Graph neural network-driven multi-telescope resource allocation improves joint observation efficiency by 50%, supporting real-time collaboration between radio and optical telescopes; Elastic scalability and compatibility: The modular interface design supports 256-node cascading, and the multi-protocol adaptive engine covers 90% of mainstream devices, building an open observation network.

[0075] Improved scientific discovery capabilities: Nanosecond-level synchronization and cross-domain fusion technology have improved FAST's positioning accuracy for FRBs from 10 square degrees to 0.1 square degrees, helping to achieve major scientific breakthroughs such as the "origin of fast radio bursts."

[0076] International cooperation and compatibility: The flexible interface adapter module supports the SKA standard protocol (SDPv3), VLBI related machine protocol (Mark 5B) and the custom format of the China VLBI network. It has passed the SKA organization's interoperability test (IOP) and became the first case of non-SKA native equipment accessing its global test bed.

[0077] Through the innovative design of seven modules, this invention solves the key technical difficulties of the FAST core array digital terminal and provides a reusable technical paradigm for the next generation of radio astronomy observation equipment.

[0078] Any process or method described in the flowchart of the present invention or in other ways herein can be understood as representing a module, segment or portion of code including one or more executable instructions for implementing specific logical functions or process steps, which can be implemented in any computer-readable medium for use by an instruction execution system, device or apparatus. The computer-readable medium can be any medium that stores, communicates, propagates or transmits a program for use by an execution system, device or apparatus, including read-only memory, magnetic disk or optical disk, etc.

[0079] Throughout this specification, reference to terms such as "embodiment" and "example" indicates that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, those skilled in the art may combine or integrate different embodiments or examples described in this specification, as well as features therein, without creating any inconsistency.

[0080] Although the above content has shown and described the embodiments of the present invention, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present invention. Ordinary technicians in this field can perform update operations such as changes, modifications, replacements and variations on the above embodiments within the scope of the present invention.

Claims

1. The FAST core array distributed collaborative observation and data fusion method based on RFSOC is characterized by: The following steps are involved: S1. System initialization: The master node RFSOC generates a global clock, synchronizes multiple boards using the SYSREF differential signal, and completes three-level clock tree calibration. S2. Signal acquisition and preprocessing: Directly sample the 3-8 GHz RF signal using an ADC and digitally down-convert it to baseband. S3. Intelligent Resource Scheduling: Using the Echo State Neural Network (ESN) to identify signal types, an improved ant colony algorithm dynamically allocates FPGA logic resources. S4. Heterogeneous Computing Acceleration: Executing a 128-channel digital beamforming pipeline on an FPGA. S5. Cross-domain data fusion: Delay compensation, gain equalization, and frequency-domain layered fusion of multi-board data. S6. Collaborative Observation Planning: Allocating Multi-Telescope Observation Resources Based on Graph Neural Networks (GATs); S7. Data output: The data is converted into SKA SDPv3, VLBI Mark5B, or a custom protocol format via a flexible interface module for data output.

2. The FAST core array distributed collaborative observation and data fusion method based on RFSOC according to claim 1 is characterized in that: The phase synchronization includes: Dynamic phase calibration algorithm monitors crystal oscillator frequency drift in real time, compensating phase accuracy up to λ / 16@2GHz; The event trigger signal is generated by the ARM core and synchronized to all processing nodes through the AXI bus, completing the three-level clock tree calibration of global clock → regional clock → logic unit clock, with a synchronization error of <500ps.

3. The FAST core array distributed collaborative observation and data fusion method based on RFSOC according to claim 1 is characterized in that: Intelligent resource scheduling includes: The ESN neural network contains a sparse randomly connected pool of 1000 neurons, outputs a resource allocation weight matrix, and the objective function is: ; in, To calculate the delay, is the resource utilization rate, is the power consumption, α, β, and γ are the corresponding weight coefficients, satisfying α+β+γ=1, and the weight coefficients are dynamically adaptive; The incremental learning mechanism updates the ESN weight every time 200 new samples are added.

4. The FAST core array distributed collaborative observation and data fusion method based on RFSOC according to claim 1 is characterized in that: Heterogeneous computing acceleration includes: The ARM core dynamically adjusts the FPGA frequency through DVFS technology and monitors the FPGA temperature, logic unit occupancy, and DDR4 bandwidth in real time. Beamforming uses a double-precision floating-point arithmetic unit, with a phase error of ≤λ / 32@2GHz.

5. The FAST core array distributed collaborative observation and data fusion method based on RFSOC according to claim 1 is characterized in that: Cross-domain data fusion includes: Cross-band spectrum splicing uses multi-resolution wavelet transform; Radio frequency interference suppression uses a blind source separation algorithm, with an interference suppression ratio of 40dB.

6. The FAST core array distributed collaborative observation and data fusion method based on RFSOC according to claim 1 is characterized in that: Collaborative observation planning includes: Construct a multi-layer heterogeneous resource graph G = (V, E), where node attributes include device parameters, state parameters, and observation target parameters; The GAT model uses an 8-head attention mechanism and outputs the task assignment matrix A∈R N×M , N is the number of nodes, M is the number of tasks, and the constraints of single-node task number ≤ 4, frequency band isolation, and transmission delay < 50μs are met.

7. The FAST core array distributed collaborative observation and data fusion method based on RFSOC according to claim 1 is characterized in that: In S6, multi-telescope observation resources are allocated based on a graph neural network (GAT), and tasks are replanned within 100ms in emergencies using a PPO reinforcement learning algorithm. Emergency response includes: When an FRB signal with an SNR greater than 7dB is detected, 16 high-sensitivity channels are prioritized, with a signal capture delay of less than 200ms. The radio and optical telescopes are triggered synchronously in the time domain, and the alignment error between exposure time and integration time is <10ms.

8. The FAST core array distributed collaborative observation and data fusion method based on RFSOC according to claim 1 is characterized in that: Data output includes: Add a CRC-32 check module at the end of the FPGA pipeline to improve the bit error rate monitoring accuracy. , data reliability>99.999%; Supports clock synchronization and data stream interaction of the SKA medium frequency array and VLA telescope.

9. The FAST core array distributed collaborative observation and data fusion method based on RFSOC according to any one of claims 1 to 8, characterized in that: The system initialization phase also includes: Load the lightweight GAT model and ESN model into FPGA; The model parameters updated locally on each node are aggregated through the federated learning architecture.

10. The FAST core array distributed collaborative observation and data fusion system based on RFSOC is characterized by: The system is used to implement the FAST core array distributed collaborative observation and data fusion method based on RFSOC according to any one of claims 1 to 8, and the system includes the following modules: Nanosecond synchronization control module: The clock management unit integrated into the RFSOC generates a 100MHz±0.1ppm global clock and achieves multi-board synchronization through the SYSREF differential signal and a three-level clock buffer architecture. Intelligent resource scheduling module: The ESN neural network and improved ant colony algorithm deployed on the ARM core build a 3×5 feature-resource association matrix to dynamically allocate FPGA logic resources; Heterogeneous computing acceleration module: A 128-channel digital beamforming pipeline on the FPGA, interacting with the ARM core through a ping-pong cache architecture; Cross-domain data fusion module: realizes spatiotemporal calibration and multi-level fusion of FAST and external telescope data; Collaborative Observation Planning Module: A task allocation engine based on the GAT model and a PPO reinforcement learning emergency response engine, supporting real-time re-planning of emergency events; Flexible interface adapter module: including 16-channel 3-8GHz RF input at the front end, dual 10G SFP+ interfaces and PCIe 4.0 interfaces at the back end; External interface interaction module: supports PTP clock synchronization, SDP protocol data transmission and standard API interface.

Citation Information

Patent Citations

  • Coherent multichannel transmit-receive system and method based on RFSoC

    CN116299259A

  • Full-digital large-scale phased array multi-beam forming hardware implementation network architecture

    CN116846420A

  • Fusion computing system and method based on quantum measurement and control board card

    CN118095460A

  • Electromagnetic spectrum real-time monitoring system and method based on radio frequency SoC chip

    CN119804978A

  • Digital rural intelligent management system based on multi-modal data fusion

    CN120495041A

Cited By

  • Delay neural network parameter ant colony optimization method and device based on pipeline analog-to-digital converter calibration

    CN121638301A

  • Ant colony optimization method and device for parameter of time delay neural network based on calibration of pipeline analog-digital converter

    CN121638301B

  • Pulse parameter intelligent detection method and system based on deep learning

    CN121721370A

  • High-precision modular event timing system

    CN121764299A

  • FPGA communication signal automatic calibration system

    CN122247578A