Heterogeneous Stream Processing Cluster Interconnect in Software-Defined Radio Systems
By designing the collaborative work of the RF front-end layer, heterogeneous computing pool layer, core interconnection layer, and cluster management layer, the problem of insufficient heterogeneous stream processing cluster interconnection performance in traditional software radio systems is solved, achieving efficient signal processing and system reliability, and is suitable for scenarios such as 5G/6G communication and deep space telemetry and control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DAYAO INFORMATION TECH (HUNAN) CO LTD
- Filing Date
- 2026-03-30
- Publication Date
- 2026-05-26
AI Technical Summary
In traditional software-defined radio systems, heterogeneous stream processing clusters suffer from insufficient cluster interconnection performance, high inter-cluster communication latency, low clock synchronization accuracy, poor scalability and compatibility, and cannot meet the high-performance signal processing requirements of 5G/6G communication and phased array radar.
Design a heterogeneous stream processing cluster interconnection device, including a radio frequency front-end layer, a heterogeneous computing power pool layer, a core interconnection layer and a cluster management layer. Employ a heterogeneous cluster internet gateway, a high-speed cluster data exchange matrix and a global cluster clock synchronization module to achieve direct interconnection and high-precision synchronization between hardware-level clusters, and support multi-node cluster collaborative work.
It significantly improves the cluster interconnection performance of heterogeneous stream processing clusters, enhances signal processing capabilities and system reliability, reduces energy consumption, and adapts to the signal processing needs of complex scenarios such as deep space telemetry and control.
Smart Images

Figure CN121940785B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wireless communication network technology and relates to a heterogeneous stream processing cluster interconnection device in a software radio system. Background Technology
[0002] Software-defined radio (SDR) refers to the use of different waveform software to form different radio signal processing systems on a general, standardized hardware and software architecture. Its core is component-based waveform design with decoupled hardware and software. Currently, general-purpose SDR systems, represented by USRP and GNU Radio, are mainly designed for low-speed signal processing using CPUs as computing units, and are mostly deployed in single-node or small-scale node configurations. In scenarios with extremely high signal processing performance requirements, such as 5G / 6G communication and phased-array radar, not only is high-speed signal stream processing capability required, but also multi-node cluster deployment is needed to achieve computing power aggregation and task division. Traditional single-computing-unit (such as using only CPU or FPGA) or small-scale non-clustered heterogeneous deployments can no longer meet these requirements. Therefore, the industry has gradually introduced heterogeneous units such as FPGAs, GPUs, DSPs, and APUs to form SDR processing clusters. However, the designed cluster interconnect architectures generally suffer from high communication latency, low clock synchronization accuracy and data forwarding efficiency between clusters, and poor cluster scalability and compatibility, exhibiting significant technical problems related to insufficient cluster interconnect performance. Summary of the Invention
[0003] To address the problems existing in the above-mentioned traditional methods, this invention proposes a heterogeneous stream processing cluster interconnection device in a software radio system, which can effectively improve the cluster interconnection performance of heterogeneous stream processing clusters.
[0004] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:
[0005] A heterogeneous stream processing cluster interconnection device is provided in a software radio system, including a radio frequency front-end layer, a heterogeneous computing power pool layer, a core interconnection layer and a cluster management layer. The radio frequency front-end layer includes multiple standardized radio frequency processing modules deployed in a cluster. Each standardized radio frequency processing module is used to downconvert the received radio frequency signal, perform high-speed AD conversion and preliminary filtering and then aggregate it into a baseband digital signal stream, and perform DA conversion and upconversion on the processing result and then transmit it in a cluster.
[0006] The core interconnection layer connects each standardized radio frequency processing module, the heterogeneous computing power pool layer, and the cluster management layer. It is used to forward the baseband digital signal stream in parallel without blocking to the heterogeneous computing power pool layer and to send the processing results output by the heterogeneous computing power pool layer back to the corresponding standardized radio frequency processing module to perform node synchronization within the cluster and phase calibration between clusters.
[0007] The heterogeneous computing power pool layer includes FPGA clusters, GPU clusters, DSP clusters, and APU clusters. The FPGA cluster is used to perform real-time parallel processing of radio frequency signals on the baseband digital signal stream. The GPU cluster is used to perform AI inference on the intermediate results output by the FPGA cluster. The DSP cluster is used to perform high-precision real-time signal processing on the classification results fed back after AI inference and output the processing results. The APU cluster is used to perform low-computing-power general computing tasks and energy-efficient scheduling of the computing power cluster.
[0008] The cluster management layer uses a distributed monitoring architecture to collect the operating parameters of each cluster node in real time, and performs hardware status monitoring and alarms for each standardized radio frequency processing module and each computing power cluster.
[0009] In one embodiment, the FPGA cluster uses a high-speed network interconnect topology and sets up a cluster central node to be responsible for data scheduling of each node in the cluster.
[0010] The GPU cluster adopts a fully interconnected topology within the cluster, and closed-loop communication is established between cluster nodes through high-speed dedicated links;
[0011] The DSP cluster uses a ring-shaped dedicated interface for interconnection, and each node connects to two adjacent nodes to form a redundant link.
[0012] The APU cluster uses a high-speed bus for self-interconnection and is configured with a cluster master node to directly manage cluster memory sharing and peripheral access.
[0013] In one embodiment, the core interconnect layer includes a heterogeneous cluster internet gateway, a high-speed cluster data exchange matrix, and a global cluster clock synchronization module;
[0014] The high-speed cluster data exchange matrix is composed of multi-port high-speed cross switch chips. The high-speed cluster data exchange matrix connects to each standardized radio frequency processing module, heterogeneous cluster Internet gateway and cluster management layer respectively. It supports parallel non-blocking forwarding of multiple cluster ports and reserves cluster expansion ports.
[0015] The heterogeneous cluster internet gateway integrates multiple types of cluster interface physical layer chips and switching units to complete direct hardware-level conversion and interconnection of various computing power clusters; the multi-cluster protocol dynamic multiplexing mode of the heterogeneous cluster internet gateway is configured through the internal registers of the cluster interface physical layer chip.
[0016] The global cluster clock synchronization module uses a high-precision clock source as the global master clock and configures an independent cluster-level clock phase-locked loop for each cluster. The global cluster clock synchronization module is used to first synchronize the clocks of all nodes in each cluster, and then calibrate the phase deviation between clusters by globally distributing the synchronization clock signal.
[0017] In one embodiment, the cluster management layer includes an industrial-grade cluster main control board, a multi-channel cluster status monitoring interface, and a unified cluster debugging interface;
[0018] The multi-channel cluster status monitoring interface adopts a distributed monitoring architecture to collect the operating parameters of each standardized radio frequency processing module, each node in each computing power cluster, heterogeneous cluster Internet gateway, high-speed cluster data exchange matrix and global cluster clock synchronization module in real time, and then aggregates and uploads them to the industrial-grade cluster main control board.
[0019] The industrial-grade cluster main control board comes pre-installed with an embedded cluster management system, which is used to manage the cluster's operating status according to operating parameters, trigger cluster-level hardware alarms, and send status signals to the APU cluster through the cluster status monitoring interface. The status signals are used to instruct the APU cluster to record the corresponding alarm status, and the alarm status is used for cluster-level fault location.
[0020] The unified cluster debugging interface is used to support the visual monitoring and centralized debugging of the hardware status of all clusters globally.
[0021] In one embodiment, all standardized RF processing modules are installed in a standard rack and deployed according to antenna array partitions. The standard rack has reserved heat dissipation channels between clusters and uses DC power supply for cluster power supply. The standardized RF processing modules support hot-swapping.
[0022] In one embodiment, each computing cluster of the heterogeneous computing power pool layer is independently deployed on a dedicated rack, and the nodes within the cluster are soldered to a standard backplane and fixedly installed on the dedicated rack through the standard backplane; the nodes within the cluster are interconnected through the wiring of the standard backplane.
[0023] In one embodiment, the core interconnect layer is integrated in the central area of a dedicated trunking switch rack, and the heterogeneous trunking internet gateway, high-speed trunking data exchange matrix, and global trunking clock synchronization module are connected via board-level high-speed signal traces.
[0024] In one embodiment, the ports of the high-speed cluster data exchange matrix are configured according to the cluster type; the partition configuration includes: ports 1 to 4 are used to connect to the cluster-level high-speed serial interface gateway of the radio frequency front-end layer, ports 5 to 8 are used to connect to the central node or master control node of each computing power cluster in the heterogeneous computing power pool layer, ports 9 to 10 are used to connect to the cluster management layer, and ports 11 to 16 are reserved for cluster expansion.
[0025] In one embodiment, the global cluster clock synchronization module supports GPS + BeiDou dual-mode time synchronization.
[0026] In one embodiment, the industrial-grade cluster main control board is equipped with LED alarm indicator lights and sound alarm modules dedicated to each computing power cluster.
[0027] One of the above technical solutions has the following advantages and beneficial effects:
[0028] The heterogeneous stream processing cluster interconnection device in the aforementioned software-defined radio system, through the design of a heterogeneous computing power cluster interconnection architecture in which the radio frequency front-end layer, heterogeneous computing power pool layer, core interconnection layer and cluster management layer work together, and with the core structural design of heterogeneous cluster internet gateway + high-speed cluster non-blocking switching + global cluster high-precision synchronization + layered cluster deployment, breaks through the various technical bottlenecks of existing heterogeneous computing power cluster integration from the hardware level, and achieves significant improvements in cluster collaboration performance, synchronization accuracy, transmission efficiency, compatibility, reliability and engineering feasibility, effectively improving the cluster interconnection performance of the heterogeneous stream processing cluster. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 This is a schematic diagram of the module architecture of a heterogeneous stream processing cluster interconnection device in a software radio system according to one embodiment;
[0031] Figure 2 This is a schematic diagram of the system interconnection of a heterogeneous stream processing cluster interconnection device in a software radio system according to one embodiment;
[0032] Figure 3 This is a schematic diagram of the system workflow of a heterogeneous stream processing cluster interconnection device in a software radio system according to one embodiment. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention.
[0034] It should be noted that, in this document, the reference to "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The presentation of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will understand that the embodiments described herein can be combined with other embodiments. The term "and / or" as used herein refers to any combination of one or more of the associated listed items, and all possible combinations, including such combinations.
[0035] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0036] In one embodiment, such as Figure 1 As shown, a heterogeneous stream processing cluster interconnection device in a software-defined radio system is provided, including a radio frequency (RF) front-end layer, a heterogeneous computing power pool layer, a core interconnection layer, and a cluster management layer. The RF front-end layer includes multiple standardized RF processing modules deployed in a cluster. Each standardized RF processing module is used to down-convert the received RF signal, perform high-speed AD conversion, and perform preliminary filtering before aggregating it into a baseband digital signal stream. It also performs DA conversion and up-conversion on the processing results before clustered transmission.
[0037] The core interconnect layer connects the standardized RF processing modules, the heterogeneous computing pool layer, and the cluster management layer. It forwards baseband digital signal streams in parallel without blocking to the heterogeneous computing pool layer and sends the processing results from the heterogeneous computing pool layer back to the corresponding standardized RF processing modules to perform intra-cluster node synchronization and inter-cluster phase calibration. The heterogeneous computing pool layer includes FPGA clusters, GPU clusters, DSP clusters, and APU clusters. The FPGA cluster performs real-time parallel RF signal processing on the baseband digital signal stream; the GPU cluster performs AI inference on the intermediate results output by the FPGA cluster; the DSP cluster performs high-precision real-time signal processing on the classification results fed back after AI inference and outputs the processing results; and the APU cluster performs low-computing-power general-purpose computing tasks and optimizes the energy efficiency scheduling of the computing cluster. The cluster management layer uses a distributed monitoring architecture to collect the operating parameters of each cluster node in real time, and performs hardware status monitoring and alarms for each standardized RF processing module and each computing cluster.
[0038] It is understandable that the heterogeneous stream processing cluster interconnection device adopts a core design concept of four-layer hardware layering + three-level cluster interconnection + global cluster synchronization. From top to bottom, it is designed with a radio frequency front-end layer, a heterogeneous computing power pool layer, a core interconnection layer, and a cluster management layer. The radio frequency front-end layer is vertically connected to the core interconnection layer through a cluster-level high-speed serial interface; for example... Figure 2As shown, the core interconnection layer includes core components such as a heterogeneous cluster internet gateway, a high-speed cluster data exchange matrix, and a global cluster clock synchronization module. These three components are directly interconnected through a board-level bus. At the same time, the heterogeneous cluster internet gateway connects downwards to the heterogeneous computing power pool layer. Each computing power cluster interfaces with the heterogeneous cluster internet gateway through a standardized cluster interface. The bottom layer is the cluster management layer, which connects to the core interconnection layer, the heterogeneous computing power pool layer, and the radio frequency front-end layer through distributed monitoring and debugging interfaces, respectively, to realize the status monitoring and control of the entire cluster system. The layers work together to form a complete software radio heterogeneous stream processing cluster system.
[0039] Each hardware layer works collaboratively through a standardized cluster interface to form a cluster interconnection system adapted to heterogeneous stream processing via software radio. The core hardware designs are as follows:
[0040] The hardware composition of the RF front-end layer (responsible for signal input / output clusters) is as follows: It adopts a clustered deployment mode and includes multiple standardized RF processing modules. Each standardized RF processing module includes multiple RF processing units. Each RF processing unit integrates commonly used RF front-end components in software radio, such as digital downconverters (DDC), digital upconverters (DUC), high-speed AD / DA converters, and low-noise amplifier circuits (adjustable gain). The link connection relationship of these components can be understood in the same way as the connection relationship in the existing RF front-end link structure in this field. The RF parameters of the standardized RF processing modules are adapted to the requirements of parallel acquisition and transmission of high-speed signal streams in software radio.
[0041] Interconnect design of the RF front-end layer: Each standardized RF processing module is configured with an independent cluster-level high-speed serial interface gateway, which supports high-speed differential transmission. It can be directly connected to the dedicated port of the high-speed cross switch chip in the high-speed cluster data exchange matrix of the core interconnect layer through high-speed cables, without the need for intermediate transfer circuits, so as to reduce cross-cluster signal attenuation and delay.
[0042] The core functions of the RF front-end layer are: each standardized RF processing module performs analog-to-digital conversion, frequency shifting and preliminary filtering of RF signals in parallel; after data aggregation within the cluster, the baseband digital signal stream is output to the heterogeneous computing power pool layer; at the same time, it receives the digital signals processed by each computing power cluster in the heterogeneous computing power pool layer and converts them into RF signals for clustered transmission, realizing the basic clustered function of signal transmission and reception of the software radio system.
[0043] The heterogeneous computing power pool layer (the core computing power processing cluster array of the device) adopts a clustered array + dedicated cluster interconnection design. Based on the computing power requirements and parallel characteristics of different software radio stream processing tasks, the heterogeneous computing power unit is divided into multiple functionally independent computing power clusters. Each computing power cluster is deployed independently and the hardware configuration is adapted to the dedicated task. The clusters achieve collaboration through the core interconnection layer.
[0044] The heterogeneous computing power pool layer includes FPGA clusters, GPU clusters, DSP clusters, and APU clusters. Figure 1 In this context, 'n' following each computing cluster name indicates the number of nodes in that cluster. The hardware configuration of the FPGA cluster can consist of multiple high-performance FPGA chips forming cluster nodes, with some nodes integrating processor cores. The cluster size can be expanded according to the task's computing power requirements for the FPGA chips (such as the number of antenna channels), as long as it can adapt to the parallel stream processing requirements. Furthermore, the interconnection method within the FPGA cluster can be: using a high-speed network (NoC) interconnection topology, with a central node responsible for data scheduling, supporting non-blocking communication and data aggregation among nodes within the cluster, ensuring efficient data interaction within the cluster. The core function of the FPGA cluster is to perform real-time parallel processing of radio frequency signals, such as, but not limited to, parallel processing of radio frequency signals in software radio core stream processing tasks like phased array radar beamforming, real-time signal filtering, and distributed processing of intermediate frequency signals, and outputting the processed intermediate results to other computing clusters.
[0045] The hardware configuration of a GPU cluster can consist of multiple high-performance GPU chips forming cluster nodes, directly connected via high-speed links, enabling massive parallel computing and AI inference capabilities. Furthermore, the interconnection method within the GPU cluster can be: a fully interconnected topology within the cluster, with closed-loop communication established between cluster nodes via high-speed dedicated links. For example, each cluster node can be directly connected to other cluster nodes via two high-speed dedicated links to form an internal closed-loop interconnection, ensuring high-bandwidth data interaction and load balancing, achieving high-speed data sharing within the cluster, and ensuring the parallel efficiency of AI inference tasks. The core function of the GPU cluster is to execute massive data AI inference tasks, such as, but not limited to, intelligent signal processing such as signal modulation recognition, interference type classification, and waveform feature extraction. The GPU cluster receives intermediate results from the FPGA cluster, performs AI inference tasks, and then feeds back the corresponding classification results. The number of cluster nodes in the GPU cluster can be configured according to the complexity of the AI inference task.
[0046] The hardware configuration of a DSP cluster can employ multiple high-performance DSP chips to form cluster nodes, supporting multi-core parallel computing and adapting to high-precision signal processing. Cluster nodes are deployed hierarchically according to processing accuracy requirements. Furthermore, the interconnection method within the DSP cluster can be: interconnection using dedicated ring interfaces, with each node connecting to adjacent nodes to form redundant links. For example, each node can connect to adjacent nodes through four dedicated interfaces to form a primary / backup dual-ring link, ensuring low-latency data transmission and fault redundancy. The core function of the DSP cluster is to perform high-precision real-time signal processing, such as, but not limited to, high-precision stream processing tasks like convolutional decoding, adaptive filtering, and signal encoding / decoding. The DSP cluster receives the classification results output by the GPU cluster and then performs high-precision real-time signal processing.
[0047] The hardware configuration of an APU cluster can utilize multiple APU chips integrating CPU and GPU to form cluster nodes, balancing general-purpose computing and lightweight graphics processing with excellent energy efficiency. Furthermore, the interconnection method within the APU cluster can be: high-speed bus self-interconnection within the cluster, with a cluster master node (with built-in memory controller) configured to directly manage cluster memory sharing and peripheral access, achieving load balancing for general-purpose computing tasks within the cluster. The core function of the APU cluster is to execute low-computing-power general-purpose computing tasks, such as, but not limited to, cluster status monitoring, cross-cluster data storage management, and waveform configuration loading and distribution, while also performing energy-efficient scheduling of the entire heterogeneous stream processing cluster interconnect device.
[0048] Among them, the APU cluster optimizes the energy efficiency scheduling of the computing power cluster of the entire heterogeneous stream processing cluster interconnection device. Specifically, it is implemented through the following mainstream scheduling methods: by offloading at the task level, lightweight general computing tasks and control and management tasks are converged to high-efficiency APUs for execution, so as to avoid high-performance computing units being idle and wasting power; based on real-time load, power consumption and temperature perception, global dynamic load balancing and power quota management are automatically realized; by optimizing the data flow path, the communication overhead between high-power nodes is reduced, and by closed-loop cluster status monitoring and abnormal scheduling, the system is ensured to achieve the lowest overall power consumption and the best energy efficiency ratio while meeting the functional and performance requirements.
[0049] The heterogeneous computing power pool layer, through clustered array deployment and dedicated interconnection design, enables complementary advantages and superposition of various computing power clusters, significantly improving the distributed processing capability of software radio systems for complex scene signals, and is suitable for cluster application scenarios with stringent signal processing requirements, such as deep space telemetry and control.
[0050] The core interconnect layer (i.e., the cluster interconnect hub) serves as the core for cross-cluster data transmission and synchronization in heterogeneous stream processing cluster interconnection devices, and is a key component in solving the interconnection and clock synchronization problems of software-defined radio heterogeneous stream processing clusters. For example... Figure 1 As shown, the hardware components of the core interconnection layer include a heterogeneous cluster internet gateway, a high-speed cluster data exchange matrix, and a global cluster clock synchronization module.
[0051] The heterogeneous cluster internet gateway is a cluster-level gateway built using mainstream chips that integrate high-speed interconnect protocol controllers. It integrates multiple types of cluster interface physical layer chips and switching units. The heterogeneous cluster internet gateway's multi-cluster protocol dynamic multiplexing mode can be configured through the internal registers of the cluster interface physical layer chips to support the dynamic multiplexing of multiple cluster interconnect protocols. This enables direct hardware-level conversion and interconnection of different heterogeneous cluster interfaces such as FPGA clusters, GPU clusters, DSP clusters, and APU clusters, without the need for software relay within the cluster nodes. The heterogeneous cluster internet gateway supports consistent access to the caches of each cluster device and shared memory between clusters, avoiding duplicate copying of cross-cluster data, reducing transmission latency, and adapting to the cross-cluster processing requirements of high-speed streaming data via software radio.
[0052] The high-speed trunking data exchange matrix can be composed of multi-port high-speed cross switch chips, supporting parallel non-blocking forwarding of multiple trunking ports. Each port is configured according to the trunking type, connecting to the RF front-end layer, heterogeneous trunking Internet gateway, and trunking management layer respectively, and reserving trunking expansion ports to ensure smooth link when multiple trunkings transmit high-speed streaming data at the same time, reduce cross-trunk forwarding delay fluctuations, and provide a stable transmission channel for software radio multi-trunk parallel streaming processing.
[0053] The global cluster clock synchronization module can use a high-precision clock source as the global master clock, and form a cluster synchronization network with a dedicated clock synchronization chip to distribute synchronization clock signals to each computing cluster and the RF front-end layer. The global cluster clock synchronization module configures an independent cluster-level clock phase-locked loop (PLL) and synchronization node for each cluster. After completing the clock synchronization of all nodes in each cluster, the phase deviation between clusters is calibrated by globally distributing the synchronization clock signal, ensuring the clock consistency between each computing cluster and the standardized RF processing module, and meeting the time accuracy requirements of distributed synchronization processing of multi-channel signals in software radio.
[0054] The heterogeneous cluster internet gateway enables direct hardware-level interconnection between clusters. Coupled with a high-speed cluster data exchange matrix, and through hardware-level protocol multiplexing and non-blocking forwarding design, it significantly reduces cross-cluster data transmission latency and reduces cross-cluster link congestion during multi-cluster parallel processing, greatly improving link utilization. This fully guarantees the timeliness of cross-cluster collaboration for core software radio tasks such as real-time RF signal processing and AI inference. The heterogeneous cluster internet gateway supports dynamic adaptation to multiple cluster protocols. When adding or upgrading computing clusters, there is no need to redesign the overall hardware circuit; only the corresponding switching unit in the gateway needs to be replaced or configured, significantly shortening the cluster adaptation cycle and adapting to the waveform iteration and computing power upgrade requirements of the software radio system. Furthermore, it integrates a global cluster clock synchronization module to achieve dual synchronization of intra-cluster node synchronization and inter-cluster phase calibration, keeping the clock deviation between multiple computing clusters and each standardized RF processing module within an extremely low range. This significantly improves the phase consistency of distributed processing of multi-channel RF signals and ensures signal processing accuracy.
[0055] The cluster management layer (responsible for cluster monitoring and control) can include an industrial-grade cluster main control board, multiple cluster status monitoring interfaces, and a unified cluster debugging interface. The cluster management layer adopts a distributed monitoring architecture, collecting real-time operating parameters such as power voltage, temperature, link status, and cluster load rate of each cluster node. Data is aggregated by monitoring nodes within the cluster and uploaded to the industrial-grade cluster main control board. When operating parameters exceed corresponding set thresholds, cluster-level hardware alarms are triggered (such as cluster-specific indicator lights + interrupt signals), and status signals are sent to the APU cluster via the cluster status monitoring interface for status recording. This allows for precise fault location (the fault can be traced back to its specific source cluster based on the correspondence between the status signal and the fault). Simultaneously, the provided unified cluster debugging interface supports global cluster hardware status query and configuration, ensuring the stable operation of the software-defined radio cluster system.
[0056] Through the cluster-level hardware status monitoring and alarm mechanism provided by the cluster management layer, monitoring nodes within the cluster collect real-time operating parameters of each cluster, effectively achieving cluster-level fault location and rapid response. This reduces the impact of a single cluster failure on the overall system operation and extends the stable operation time of the cluster system. Simultaneously, cluster alarm and thermal control monitoring functions are independently managed by the APU cluster, without occupying core computing cluster resources. Maintenance personnel can achieve parallel monitoring of multiple cluster statuses through a unified debugging interface, significantly simplifying the operation process, reducing manual maintenance costs and downtime losses, and improving the deployment reliability of the software-defined radio (SDR) cluster system. The provided unified cluster debugging interface also supports visualized monitoring and centralized debugging of the hardware status of all clusters globally.
[0057] The heterogeneous stream processing cluster interconnection device in the aforementioned software-defined radio system, through the design of a heterogeneous computing power cluster interconnection architecture in which the radio frequency front-end layer, heterogeneous computing power pool layer, core interconnection layer and cluster management layer work together, and with the core structural design of heterogeneous cluster internet gateway + high-speed cluster non-blocking switching + global cluster high-precision synchronization + layered cluster deployment, breaks through the various technical bottlenecks of existing heterogeneous computing power cluster integration from the hardware level, and achieves significant improvements in cluster collaboration performance, synchronization accuracy, transmission efficiency, compatibility, reliability and engineering feasibility, effectively improving the cluster interconnection performance of the heterogeneous stream processing cluster.
[0058] Compared to existing technologies, heterogeneous stream processing cluster interconnection devices significantly improve the overall system processing power and energy efficiency ratio by allocating cluster computing power on demand and controlling power consumption at the hardware level. This reduces the long-term energy consumption of cluster devices, making them particularly suitable for software radio cluster applications with low power consumption requirements, such as deep space telemetry and field deployment, thereby reducing long-term operating costs.
[0059] The system initialization process for the heterogeneous stream processing cluster interconnection device in the aforementioned software-defined radio system is as follows:
[0060] Cluster power-on startup: Power on in the order of cluster management layer, core interconnection layer, heterogeneous computing power pool layer to radio frequency front-end layer. The industrial-grade cluster main control board first completes self-test and queries the initial status of each cluster through the distributed monitoring network. If there is no abnormality, the power normal indicator light of each cluster is lit; if there is an abnormality, a fault alarm is issued for the abnormal cluster.
[0061] Cluster Protocol Negotiation: The heterogeneous cluster internet gateway establishes a connection with the central / master node of each computing power cluster through the basic interconnection protocol, automatically identifies the interface type and scale of each computing power cluster, loads the corresponding protocol conversion firmware, completes hardware-level cluster protocol adaptation, and lays the foundation for cross-cluster data transmission.
[0062] Cluster clock synchronization calibration: After the high-precision clock source of the global cluster clock synchronization module is started, it completes satellite timing lock (locking no less than 4 satellites), outputs a reference clock signal, and distributes it to the synchronization nodes of each cluster through a dedicated clock synchronization chip. First, the clock synchronization of all nodes in the cluster is achieved through the phase-locked loop chip configured in the cluster. Then, the industrial-grade cluster main control board distributes the synchronization clock signal globally through the clock synchronization protocol to calibrate the phase deviation between clusters and ensure that the clock consistency of all clusters meets the standard.
[0063] Cluster system ready: The industrial-grade cluster main control board sends initialization commands to each corresponding cluster through the high-speed cross switch chip of the high-speed cluster data exchange matrix, enabling each standardized RF processing module to complete AD / DA converter calibration and gain configuration, and enabling each computing power cluster to complete load initialization; after all clusters have fed back "ready" signals to the industrial-grade cluster main control board, the system enters standby mode, waiting for software radio waveform loading and cluster task issuance.
[0064] The system workflow of the heterogeneous stream processing cluster interconnection device in the above software-defined radio system is as follows: Figure 3 As shown, the RF front-end layer receives RF signals, performs AD conversion, down-conversion, and preliminary filtering, and then outputs the corresponding baseband digital signal. After receiving the baseband digital signal, the high-speed cluster data exchange matrix of the core interconnect layer forwards the baseband digital signal to the target computing power cluster according to the task requirements and the corresponding protocol. For example, the FPGA cluster performs real-time parallel processing on the baseband digital signal, such as beamforming and frequency offset compensation; the GPU cluster performs AI inference on the intermediate results output by the real-time parallel processing of the FPGA cluster, such as modulation recognition and interference classification; and the DSP cluster receives the classification results output by the GPU cluster and performs high-precision real-time signal processing, such as encoding / decoding and adaptive filtering, and then feeds the processing results back to the core. The high-speed cluster data exchange matrix of the core interconnect layer transmits the processing results back to the RF front-end layer. The RF front-end layer performs up-conversion and DA conversion on the processing results to convert the digital signals into RF signals for transmission. At the same time, the APU cluster records the processing data of the aforementioned signal processing tasks, such as computing power usage and processing status. Finally, the task records and status monitoring data are uploaded to the industrial-grade cluster main control board (which acts as a monitoring terminal) to achieve traceability of the task process and status monitoring. When the RF signal processing is completed and no new RF signals are received, the system goes into standby mode; otherwise, it continues to receive and process RF signals.
[0065] In one embodiment, all standardized RF processing modules in the aforementioned RF front-end layer are further installed in a standard rack, deployed according to antenna array partitions, with reserved heat dissipation channels between clusters (i.e., standardized RF processing modules) for equipment cooling and DC power supply for cluster power supply. Hot-swapping of cluster-level (i.e., standardized RF processing modules) is supported, facilitating maintenance and expansion, and better meeting the clustering requirements of software-defined radio systems for signal transmission and reception. The standard rack can be a commonly used RF front-end deployment rack in software-defined radio systems.
[0066] In one embodiment, the aforementioned heterogeneous computing power pool layer further adopts a clustered and compact layout. Each computing power cluster is independently deployed on a dedicated rack, and the nodes within the cluster are soldered to a standard backplane. The nodes are fixedly installed on the dedicated rack via the standard backplane, and interconnection within the cluster is achieved through the wiring on the standard backplane, thereby reducing external interference and improving the stability of signal transmission within the cluster. The standard backplane can be a PCB board, and the dedicated rack can be a host mounting rack / cabinet capable of accommodating various processor boards. Typically, the dedicated rack can also integrate the power ports and cooling system required by the computing power cluster.
[0067] In one embodiment, the core interconnect layer, acting as the central hub for cluster interconnection, can be integrated into the central area of a dedicated cluster switch rack. The heterogeneous cluster internet gateway, high-speed cluster data exchange matrix, and global cluster clock synchronization module of the core interconnect layer are connected via board-level high-speed signal traces, eliminating the need for external adapter circuits and reducing cross-cluster signal attenuation and delay. The dedicated cluster switch rack can be a variable-structure rack or customized according to the required component installation layout.
[0068] Among them, the heterogeneous cluster internet gateway can allocate cluster protocol resources according to task priority. For example, real-time processing tasks have the highest priority and are allocated the required cluster protocol resources first. The heterogeneous cluster internet gateway connects the master control nodes of each computing power cluster through multiple high-speed signal links to realize hardware-level protocol conversion and cluster interconnection.
[0069] Optionally, the ports of the high-speed cluster data exchange matrix are configured according to the cluster type. For example, ports 1 to 4 connect to the cluster-level high-speed serial interface gateway of the radio frequency front-end layer, ports 5 to 8 connect to the master control nodes of each computing power cluster in the heterogeneous computing power pool layer, ports 9 to 10 connect to the cluster management layer, and ports 11 to 16 are reserved for cluster expansion. The port rate and forwarding mode of the high-speed cluster data exchange matrix can be locked through existing configuration tools to ensure cross-cluster link bandwidth and low-latency transmission, and adapt to the high-speed streaming data requirements of software radio.
[0070] Optionally, the global cluster clock synchronization module supports GPS+BeiDou dual-mode timing and outputs a 10MHz reference clock signal. The global cluster clock synchronization module expands the clock output port through two dedicated clock synchronization chips. Each dedicated clock synchronization chip can output eight synchronization clock signals to the nodes in each computing cluster and the RF front-end layer that require clock synchronization, ensuring that the clock consistency of each cluster meets the standards efficiently.
[0071] In one embodiment, the industrial-grade cluster main control board of the cluster management layer further integrates a large-capacity memory and storage unit, and pre-installs an existing embedded cluster management system. In the distributed monitoring network deployed by the industrial-grade cluster main control board, each computing power cluster can be configured with two temperature sensors and one current sensor. The two temperature sensors are respectively installed on the chip surface of the central / main control node of the respective computing power cluster, and the current sensor is installed on the output port of the power module of the respective computing power cluster to monitor the power output current. Each sensor is connected to the monitoring node in the cluster through an I2C bus. The monitoring node aggregates the data collected by the sensors and uploads it to the industrial-grade cluster main control board.
[0072] Optionally, the cluster management layer can design a unified cluster debugging interface. The unified cluster debugging interface can be installed on the front of the rack where the industrial-grade cluster main control board is located. The industrial-grade cluster main control board can also be configured with dedicated LED alarm indicators for each computing power cluster (such as indicators corresponding to power failure, temperature over-limit and link interruption respectively). It can also integrate an audible alarm module (such as a buzzer) to realize rapid early warning and location of cluster faults.
[0073] In some implementations, the workflow of the heterogeneous stream processing cluster interconnection device in the above-mentioned software-defined radio system is further described using distributed processing of phased array radar radio frequency signals as an example:
[0074] Signal clustering input: The radio frequency signals received by the phased array radar antenna array are input to the radio frequency front-end layer according to the channel partition. Each standardized radio frequency processing module simultaneously performs analog-to-digital (AD) conversion, frequency shifting (down-conversion to baseband signal) and preliminary filtering. After aggregation by the nodes in the radio frequency front-end layer cluster, the baseband digital signal stream is output, completing the signal clustering input and preprocessing of the software radio system.
[0075] Cross-cluster data distribution: Each standardized RF processing module transmits the baseband digital signal stream to the high-speed cross switch of the high-speed cluster data exchange matrix in the core interconnect layer through the corresponding cluster-level high-speed serial interface gateway. According to the signal processing requirements (such as real-time beamforming), the cross switch forwards the data directly to the cluster center node of the FPGA cluster through a high-speed protocol, and the cluster center node distributes it to each node in the FPGA cluster for processing.
[0076] Cluster collaborative processing: The nodes of the FPGA cluster achieve parallel computing through NoC interconnection, completing real-time processing tasks such as radar beamforming and Doppler frequency offset compensation. The processed data is aggregated by the central node of the cluster and shared to the GPU cluster through a memory sharing protocol. The master node of the GPU cluster distributes data to each node and uses parallel computing capabilities to complete AI inference tasks such as signal modulation recognition and interference type classification. The classification results are aggregated and transmitted to the DSP cluster through a general interconnection protocol. The DSP cluster achieves multi-node collaboration through dual-ring interconnection, completing signal convolution decoding and adaptive filtering of the classification results, outputting high-precision processing results, and realizing collaborative scheduling of heterogeneous clusters.
[0077] Clustered feedback results: The processed digital signal is returned to the RF front-end layer via the high-speed cross switch of the high-speed cluster data exchange matrix in the core interconnection layer. The DUC chip of each standardized RF processing module performs up-conversion to RF signal and transmits it in a clustered manner through the phased array radar antenna array. During signal processing, the APU cluster records the processing data of each cluster in real time (such as computing power utilization and signal processing status) and uploads it to the monitoring terminal through the cluster management layer to realize traceability of the cluster collaboration process.
[0078] In some implementation methods, the following can also be carried out:
[0079] Cross-cluster transmission delay test: Standard pulse signals are input to each standardized RF processing module in the RF front-end layer through a signal generator. The timestamps of the signal input to the standardized RF processing module, the FPGA cluster receiving, the GPU cluster receiving, and the DSP cluster output are collected by an oscilloscope. The transmission delay of each cross-cluster link is calculated to verify whether the delay index meets the real-time processing requirements of software radio.
[0080] Inter-cluster clock synchronization accuracy test: The phase difference of the output clock of each computing cluster and each standardized RF processing module is measured using a spectrum analyzer to verify the synchronization accuracy of nodes within the cluster and the synchronization accuracy between clusters, ensuring that the phase consistency requirements of multi-channel signal distributed processing are met.
[0081] Cluster computing power performance test: Run a pre-built radar signal processing standard test set, and use the monitoring terminal to count the load rate of each computing cluster and the total computing power of the system to verify the system's distributed processing capability for complex signals, especially the ability to capture weak signals.
[0082] Cluster scalability test: Remove one node card from the GPU cluster, expand by adding two new GPU nodes to form a new GPU cluster, replace only the firmware of the corresponding switching unit of the heterogeneous cluster Internet gateway, and restart the system. The cluster collaborative signal processing function is normal, verifying the scalability of the architecture cluster.
[0083] Through the above tests and verifications, the heterogeneous stream processing cluster interconnection device in the above software radio system can be fully realized and verified. It meets the high-speed signal processing requirements of software radio cluster systems such as phased array radar, 5G / 6G communication and deep space telemetry and control, and has reproducibility, scalability and engineering feasibility.
[0084] Each module in the heterogeneous stream processing cluster interconnection device of the aforementioned software-defined radio system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of a device with data processing capabilities, or stored in software within the memory of the aforementioned device, so that the processor can invoke and execute the operations corresponding to each module. The aforementioned device can be, but is not limited to, various types of computer devices already existing in the art.
[0085] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0086] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of protection of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and all such modifications and improvements fall within the scope of protection of the present invention.
Claims
1. A heterogeneous stream processing cluster interconnection device in a software-defined radio system, characterized in that, It includes an RF front-end layer, a heterogeneous computing power pool layer, a core interconnection layer, and a cluster management layer. The RF front-end layer includes multiple standardized RF processing modules deployed in a cluster. Each standardized RF processing module is used to downconvert the received RF signal, perform high-speed AD conversion and preliminary filtering, and then aggregate it into a baseband digital signal stream. It also performs DA conversion and upconversion on the processing results before cluster transmission. The core interconnection layer connects each standardized radio frequency processing module, the heterogeneous computing power pool layer, and the cluster management layer. It is used to forward the baseband digital signal stream in parallel without blocking to the heterogeneous computing power pool layer and to send the processing results output by the heterogeneous computing power pool layer back to the corresponding standardized radio frequency processing module to perform node synchronization within the cluster and phase calibration between clusters. The heterogeneous computing power pool layer includes FPGA clusters, GPU clusters, DSP clusters, and APU clusters. The FPGA cluster is used to perform real-time parallel processing of radio frequency signals on the baseband digital signal stream. The GPU cluster is used to perform AI inference on the intermediate results output by the FPGA cluster. The DSP cluster is used to perform high-precision real-time signal processing on the classification results fed back after AI inference and output the processing results. The APU cluster is used to perform low-computing-power general computing tasks and energy-efficient scheduling of the computing power cluster. The cluster management layer adopts a distributed monitoring architecture to collect the operating parameters of each cluster node in real time, and performs hardware status monitoring and alarms for each standardized radio frequency processing module and each computing power cluster. The FPGA cluster uses a high-speed network interconnect topology and sets up a cluster central node to be responsible for data scheduling of each node in the cluster. The GPU cluster adopts a fully interconnected topology within the cluster, and closed-loop communication is established between cluster nodes through high-speed dedicated links; The DSP cluster uses a ring-shaped dedicated interface for interconnection, and each node connects to two adjacent nodes to form a redundant link. The APU cluster uses a high-speed bus for self-interconnection and is configured with a cluster master node to directly manage cluster memory sharing and peripheral access. The core interconnection layer includes a heterogeneous cluster internet gateway, a high-speed cluster data exchange matrix, and a global cluster clock synchronization module; The high-speed cluster data exchange matrix is composed of multi-port high-speed cross switch chips. The high-speed cluster data exchange matrix connects to each standardized radio frequency processing module, heterogeneous cluster Internet gateway and cluster management layer respectively. It supports parallel non-blocking forwarding of multiple cluster ports and reserves cluster expansion ports. The heterogeneous cluster internet gateway integrates multiple types of cluster interface physical layer chips and switching units to complete direct hardware-level conversion and interconnection of various computing power clusters; the multi-cluster protocol dynamic multiplexing mode of the heterogeneous cluster internet gateway is configured through the internal registers of the cluster interface physical layer chip. The global cluster clock synchronization module uses a high-precision clock source as the global master clock and configures an independent cluster-level clock phase-locked loop for each cluster. The global cluster clock synchronization module is used to first synchronize the clocks of all nodes in each cluster, and then calibrate the phase deviation between clusters by globally distributing the synchronization clock signal.
2. The heterogeneous stream processing cluster interconnection device in the software-defined radio system according to claim 1, characterized in that, The cluster management layer includes an industrial-grade cluster main control board, multiple cluster status monitoring interfaces, and a unified cluster debugging interface; The multi-channel cluster status monitoring interface adopts a distributed monitoring architecture to collect the operating parameters of each standardized radio frequency processing module, each node in each computing power cluster, heterogeneous cluster Internet gateway, high-speed cluster data exchange matrix and global cluster clock synchronization module in real time, and then aggregates and uploads them to the industrial-grade cluster main control board. The industrial-grade cluster main control board is pre-installed with an embedded cluster management system, which is used to manage the cluster's operating status according to operating parameters, trigger cluster-level hardware alarms, and send status signals to the APU cluster through the cluster status monitoring interface. Status signals are used to indicate the corresponding alarm status recorded by the APU cluster, and alarm status is used for cluster-level fault location. The unified cluster debugging interface is used to support the visual monitoring and centralized debugging of the hardware status of all clusters globally.
3. The heterogeneous stream processing cluster interconnection device in the software-defined radio system according to claim 1, characterized in that, All standardized RF processing modules are installed in standard racks and deployed according to antenna array zones. The standard racks have reserved heat dissipation channels between clusters and use DC power supplies for cluster power supply; the standardized RF processing modules support hot-swapping.
4. The heterogeneous stream processing cluster interconnection device in the software-defined radio system according to claim 1, characterized in that, Each computing cluster in the heterogeneous computing power pool layer is independently deployed on a dedicated rack, and the nodes within the cluster are soldered to a standard backplane and fixedly installed on the dedicated rack through the standard backplane; the nodes within the cluster are interconnected through the wiring of the standard backplane.
5. The heterogeneous stream processing cluster interconnection device in the software-defined radio system according to claim 1, characterized in that, The core interconnect layer is integrated in the central area of the dedicated trunking switch rack. The heterogeneous trunking internet gateway, high-speed trunking data exchange matrix, and global trunking clock synchronization module are connected through board-level high-speed signal traces.
6. The heterogeneous stream processing cluster interconnection device in the software-defined radio system according to claim 5, characterized in that, Each port of the high-speed cluster data exchange matrix is configured according to the cluster type. The partition configuration includes: ports 1 to 4 are used to connect to the cluster-level high-speed serial interface gateway of the radio frequency front-end layer; ports 5 to 8 are used to connect to the central node or master control node of each computing power cluster in the heterogeneous computing power pool layer; ports 9 to 10 are used to connect to the cluster management layer; and ports 11 to 16 are reserved for cluster expansion.
7. The heterogeneous stream processing cluster interconnection device in the software-defined radio system according to claim 5, characterized in that, The global cluster clock synchronization module supports GPS + BeiDou dual-mode time synchronization.
8. The heterogeneous stream processing cluster interconnection device in the software-defined radio system according to claim 2, characterized in that, The industrial-grade cluster main control board is equipped with dedicated LED alarm indicators and audible alarm modules for each computing power cluster.
Citation Information
Patent Citations
Wireless communication system and method based on reconfigurable FPGA modularization development
CN119316851A
Cluster channel machine based on RF Transceiver software radio
CN222981548U