Spaceborne Communication Equipment Fault Self-Monitoring and Hierarchical Single-Event Soft Error Handling System
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-08-11
AI Technical Summary
第一类是单粒子翻转,导致存储器单元逻辑状态翻转,表现为配置参数错乱、逻辑功能异常、数据传输错误、电路功能异常等
[0026] The aforementioned spaceborne communication equipment fault self-monitoring and hierarchical single-event soft error (SEE) handling system constructs six complementary fault monitoring points covering hardware status, service links, and FPGA dedicated interfaces. Through cross-validation to eliminate false alarms, the system significantly reduces the fault detection miss rate and can promptly detect various latent SEE errors. A dual response mechanism combining fault-triggered immediate refresh and timed hierarchical refresh is established, enabling rapid repair after a fault occurs, solving the problem of passive lag and waiting hours for repair in existing technologies. Three hierarchical repair modes are implemented: dynamic refresh, full refresh, and global reset. Most SEE errors can be resolved through dynamic refresh without service interruption, with only a very small number of severe faults requiring a global reset. The introduction of a satellite transit and service status awareness mechanism automatically avoids high-impact repair operations during peak service periods, significantly reducing the impact of repair operations on user service quality. The system automatically selects the optimal repair method based on the severity of the fault, avoiding unnecessary global resets and full refreshes, reducing system resource waste, and improving the overall operating efficiency of the spaceborne communication equipment.
Smart Images

Figure CN122553980A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of radiation hardening and common basic technologies for satellites, and in particular to a self-monitoring and graded single-event soft error handling system for spaceborne communication equipment faults. Background Technology
[0002] Low-Earth orbit (LEO) satellite internet constellations have entered the stage of large-scale commercialization. As the core equipment for data transmission between satellites and the ground, and between satellites, the onboard communication unit extensively uses commercial off-the-shelf signal processing FPGAs to implement baseband processing, protocol stack operation, interface control, and digital signal processing functions. Signal processing FPGAs have advantages such as high integration, high flexibility, short development cycle, and low cost. However, their configuration memory, block memory, distributed memory, and logic resources are highly sensitive to single-event effects caused by high-energy charged particles in space.
[0003] In the space environment, the semiconductor devices of FPGAs used for signal processing by high-energy proton and heavy-ion bombardment can trigger two types of core soft errors. The first type is single-event upsets (SWEs), which cause memory cell logic states to flip, manifesting as incorrect configuration parameters, abnormal logic functions, data transmission errors, and circuit malfunctions. The second type is single-event interrupts (SWEs), which cause soft failures in the FPGA's global control logic, configuration interface circuits, or clock circuits, manifesting as an overall FPGA crash, unresponsive interfaces, and complete service interruption.
[0004] Single-event soft errors have become the primary factor affecting the on-orbit reliability of spaceborne communication systems. Statistics show that over 70% of on-orbit failures in low-Earth orbit satellite communication systems are caused by single-event effects, which can severely paralyze the entire satellite's communication services.
[0005] Currently, single-event soft error handling technology for spaceborne communication devices suffers from the following unavoidable core defects.
[0006] First, the fault monitoring system is too simplistic and has a high failure rate. Existing technologies generally rely solely on watchdog timers or simple core register checks for fault monitoring, failing to cover latent faults such as PLL lockout, abnormal link signal-to-noise ratio, CAN bus interface interruption, and baseband processing logic errors. According to actual test data, the failure rate of existing single monitoring systems is as high as 35%, with a large number of soft errors going undetected in time, leading to the continuous spread of faults and prolonged service interruptions.
[0007] Second, the refresh mechanism is passive and lagging, resulting in untimely responses. Existing solutions mostly use fixed-period global resets to handle single-event soft errors that cannot be resolved by dynamic refreshes. These single-event reset cycles are typically 24 hours or longer. After a fault occurs, several hours must pass before repair is triggered, during which the communication link remains abnormal, leading to a poor user experience and requiring manual intervention for timely recovery. While some solutions introduce fault triggering mechanisms, they can only respond to the most severe FPGA crashes and cannot address early, latent faults in a timely manner.
[0008] Third, the recovery methods are limited, and the impact on services is significant. Existing technologies almost entirely rely on FPGA dynamic refresh for recovery, which forcibly interrupts all services during the reset process, with interruptions lasting three to ten seconds. This fails to meet the over 99.99% availability requirement for telecom-grade satellite communications. A global reset also causes all user connections to be lost and data to be lost, severely impacting the quality of commercial satellite internet services.
[0009] Fourth, the timing of resets is inappropriate and not adapted to satellite operational characteristics. Current timed reset operations do not differentiate between satellites within ground station coverage areas or consider current service load, frequently performing resets during peak hours, causing a large number of users to lose connection simultaneously and leading to user complaints. Low-Earth orbit satellites have a single ground station transit time of only about ten minutes, a period coinciding with peak service hours. Inappropriate timing of timed resets and full refreshes has a significant impact on service quality. Summary of the Invention
[0010] Therefore, it is necessary to provide a self-monitoring and hierarchical single-event soft error handling system for spaceborne communication equipment that can achieve multi-dimensional fault monitoring, hierarchical processing, and intelligent timing optimization, in response to the above-mentioned technical problems.
[0011] A self-monitoring and graded single-event soft error handling system for a spaceborne communication device adopts an A / B dual-machine hot backup hardware architecture. Both machine A and machine B contain a monitoring Flash-FPGA and a signal processing SRAM-type FPGA. The monitoring Flash-FPGA establishes a physical connection with the signal processing SRAM-type FPGA, the radio frequency module, and the peer communication interface through a dedicated interface.
[0012] The Flash-FPGA is internally equipped with a multi-dimensional fault monitoring module, a decision control module, a hierarchical handling and execution module, and a business status perception module. These modules are interconnected through an internal high-speed bus and form a collaborative working mechanism.
[0013] The multi-dimensional fault monitoring module is used to collect fault monitoring data from multiple dimensions in real time from the signal processing SRAM-type FPGA, RF module and peer communication interface through a dedicated interface, and send the single-event soft error monitoring results to the decision control module.
[0014] The decision control module is used to receive monitoring results, cross-validate data from multiple related sources in the monitoring results, generate graded handling instructions based on the error impact range and severity of single-event soft errors, and output the graded handling instructions to the graded handling execution module.
[0015] The graded processing execution module is used to receive graded processing instructions and perform graded repair operations on the signal processing SRAM-type FPGA through a dedicated interface.
[0016] The service status awareness module is used to acquire the satellite transit status and the service load status of the communication device, and feeds the perceived status data back to the decision control module. The decision control module optimizes the timing of the hierarchical handling execution module to perform repair operations based on the status data.
[0017] In one embodiment, the multi-dimensional fault monitoring module includes a phase-locked loop (PLL) status monitoring unit, a communication link interaction monitoring unit, a service channel locking monitoring unit, an RF link status comparison unit, a dual-machine uplink status comparison unit, and a configuration interface status monitoring unit. The PLL status monitoring unit monitors the locking indication status of the PLL inside the signal processing SRAM-type FPGA. The communication link interaction monitoring unit monitors the continuity of communication interaction between the signal processing SRAM-type FPGA and an external satellite computer. The service channel locking monitoring unit monitors the validity of the carrier lock, bit synchronization, and frame synchronization indication signals output by the signal processing SRAM-type FPGA. The RF link status comparison unit reads and compares the automatic gain control value of the RF module with the baseband demodulation signal-to-noise ratio. The dual-machine uplink status comparison unit obtains the uplink status data of both machine A and machine B through the peer communication interface and compares the status differences between machine A and machine B. The configuration interface status monitoring unit reads the status register and configuration register data of the signal processing SRAM-type FPGA before configuration refresh.
[0018] In one embodiment, the decision control module includes a fault cross-verification unit, a tiered handling decision unit, and a timed task management unit. The fault cross-verification unit, upon receiving an abnormal alarm from any monitoring unit, synchronously calls the monitoring results of other monitoring units physically and logically related to the abnormal alarm. If all related monitoring results meet preset abnormal conditions, a single-event fault is confirmed. The tiered handling decision unit, after confirming a single-event fault, classifies the fault into three handling levels and generates repair instructions corresponding to each level. Specifically, the first-level instruction instructs the tiered handling execution module to perform a local module reset and immediately trigger a dynamic refresh; the second-level instruction instructs a full refresh of the configuration memory; and the third-level instruction instructs a global reset and reload of the configuration bitstream. The timed task management unit maintains timed repair tasks with preset cycles and, based on satellite transit and service load status feedback from the service status awareness module, determines whether to adjust the triggering timing of the timed repair tasks.
[0019] In one embodiment, the hierarchical processing execution module is connected to the configuration memory of the signal processing SRAM-type FPGA via the SelectMap interface, and supports configuration memory dynamic refresh mode, configuration memory full refresh mode, and FPGA global reset mode. The configuration memory dynamic refresh mode is used to dynamically refresh all configuration memory bits in the signal processing SRAM-type FPGA except for distributed RAM, shift register chains, and BRAM. The configuration memory full refresh mode is used to incrementally refresh all configuration memory bits of the signal processing SRAM-type FPGA page by page. The FPGA global reset mode is used to send a reset signal to the reset pin of the signal processing SRAM-type FPGA and control the signal processing SRAM-type FPGA to reload the complete configuration bitstream.
[0020] In one embodiment, the service status awareness module includes a satellite transit status determination unit and a service load status determination unit. The satellite transit status determination unit calculates the relative elevation angle between the satellite and the ground station based on ephemeris data, and determines that the current location is in a satellite transit state when the relative elevation angle meets a preset coverage threshold. The service load status determination unit determines whether the communication device is in a preset peak service period state based on the uplink signal lock status, automatic gain control value, demodulation signal-to-noise ratio, and downlink transmitter on / off status of the communication device.
[0021] In one embodiment, the decision control module optimizes the timing of repair execution based on status data. When a pre-defined periodic repair task or a configuration memory full refresh task is triggered, and the service status awareness module reports that the communication device is in a peak service period, a delayed execution instruction is sent to the hierarchical handling execution module. During the delay, the status data reported by the service status awareness module is continuously monitored, and the pending repair operation is executed after the status data returns to a service idle state. If an immediate emergency repair request is triggered by the multi-dimensional fault monitoring module during the delay, the decision control module prioritizes outputting the corresponding immediate repair instruction to the hierarchical handling execution module.
[0022] In one embodiment, the collaborative working mechanism of the functional modules via an internal high-speed bus is as follows: the multi-dimensional fault monitoring module collects fault monitoring data of various dimensions according to a preset cycle and outputs it to the decision control module. The decision control module receives the monitoring data in real time and jumps to the verification stage when a suspected abnormal state is detected, comparing it with other monitoring data associated with the suspected abnormal state to confirm whether a real fault has occurred. After confirming the fault, the decision control module selects the corresponding graded handling instruction based on the scope and severity of the fault impact and transmits it to the graded handling execution module. After the graded handling execution module completes the graded repair operation, it feeds back the completion signal to the decision control module, which triggers the multi-dimensional fault monitoring module to retest the repaired data to verify the repair effectiveness. If the verification of the repair fails, the decision control module will automatically upgrade the repair level and output the corresponding repair instruction again. The service status awareness module feeds back the satellite transit status and service load status to the decision control module in real time to adjust the timing of the output of the graded handling instructions.
[0023] A self-monitoring and graded single-event soft error (SEE) handling device for spaceborne communication equipment includes a multi-dimensional fault monitoring module, a decision control module, a graded handling execution module, and a service status awareness module. The multi-dimensional fault monitoring module collects fault monitoring data from multiple dimensions in real time via a dedicated interface from the signal processing SRAM-type FPGA, the radio frequency module, and the peer communication interface, and outputs SEE monitoring results. The decision control module receives the monitoring results, cross-validates them to eliminate false alarms, and generates graded handling instructions based on the scope and severity of the SEE's impact. The graded handling execution module receives the graded handling instructions and performs graded repair operations on the signal processing SRAM-type FPGA via a dedicated interface. The service status awareness module acquires the satellite's transit status and the communication equipment's service load status, and feeds the perceived status data back to the decision control module. The decision control module optimizes the timing of the graded handling execution module's repair operations based on the status data.
[0024] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the functions of any of the above-mentioned systems.
[0025] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the functions of any of the above-mentioned systems.
[0026] The aforementioned spaceborne communication equipment fault self-monitoring and hierarchical single-event soft error (SEE) handling system constructs six complementary fault monitoring points covering hardware status, service links, and FPGA dedicated interfaces. Through cross-validation to eliminate false alarms, the system significantly reduces the fault detection miss rate and can promptly detect various latent SEE errors. A dual response mechanism combining fault-triggered immediate refresh and timed hierarchical refresh is established, enabling rapid repair after a fault occurs, solving the problem of passive lag and waiting hours for repair in existing technologies. Three hierarchical repair modes are implemented: dynamic refresh, full refresh, and global reset. Most SEE errors can be resolved through dynamic refresh without service interruption, with only a very small number of severe faults requiring a global reset. The introduction of a satellite transit and service status awareness mechanism automatically avoids high-impact repair operations during peak service periods, significantly reducing the impact of repair operations on user service quality. The system automatically selects the optimal repair method based on the severity of the fault, avoiding unnecessary global resets and full refreshes, reducing system resource waste, and improving the overall operating efficiency of the spaceborne communication equipment. Attached Figure Description
[0027] Figure 1 This is a framework diagram of a single-event soft error classification monitoring and processing system for a spaceborne communication device in one embodiment; Figure 2 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0029] This application provides a single-event soft error (SEE) classification monitoring and processing system for spaceborne communication equipment. This system is deployed in the onboard communication terminal of a low-Earth orbit satellite internet constellation and adopts an A / B dual-machine hot backup hardware architecture. Machines A and B have identical hardware configurations. During normal operation, machine A acts as the primary machine, and machine B as the backup machine. In the event of a primary machine failure, the system automatically switches to the backup machine. Each communication unit includes a monitoring Flash-FPGA, a signal processing SRAM-type FPGA, a power supply module, an RF module, an ADDA module, and a clock module. The monitoring Flash-FPGA uses a Flash-type FPGA, possessing extremely high single-event immunity, and is responsible for fault monitoring, decision control, and repair execution for the entire system. The signal processing SRAM-type FPGA uses commercial COTS devices and is responsible for baseband processing, protocol stack operation, interface control, and digital signal processing functions. Four core functional units are constructed between the monitoring Flash-FPGA and the signal processing SRAM-type FPGA: a multi-dimensional fault monitoring module, a decision control module, a classification processing execution module, and a service status awareness module. These units are interconnected through an internal serial bus to collaboratively complete the detection, decision-making, and repair of single-event soft errors.
[0030] In one embodiment, such as Figure 1 As shown, a single-event soft error (SEE) classification monitoring and processing system for a spaceborne communication device is provided. The system adopts an A / B dual-machine hot-backup hardware architecture. Both machine A and machine B include a monitoring Flash-FPGA and a signal processing SRAM-type FPGA. The monitoring Flash-FPGA establishes a physical connection with the signal processing SRAM-type FPGA, the radio frequency module, and the peer communication interface through a dedicated interface.
[0031] The Flash-FPGA is internally equipped with a multi-dimensional fault monitoring module, a decision control module, a hierarchical handling and execution module, and a business status perception module. These modules are interconnected through an internal high-speed bus and form a collaborative working mechanism.
[0032] The multi-dimensional fault monitoring module is used to collect fault monitoring data from multiple dimensions in real time from the signal processing SRAM-type FPGA, RF module, and peer communication interface through a dedicated interface, and sends the single-event soft error monitoring results to the decision control module. This module solves the problem that the failure rate of existing technologies relying on only a single monitoring method is as high as 35% or more. Its key lies in the construction of six complementary monitoring points covering hardware status, business links, and FPGA dedicated interfaces, which simultaneously collect fault characteristics from different physical levels and data links, so that no latent faults are missed.
[0033] The decision control module receives monitoring results, cross-validates data from multiple related sources within the monitoring results, and generates tiered handling instructions based on the scope and severity of the single-event soft error's impact. These instructions are then output to the tiered handling execution module. The core innovation of this module lies in establishing a cross-validation mechanism and a three-level fault classification mapping relationship, tightly coupling fault detection with the selection of repair strategies, thus avoiding the crude, one-size-fits-all approach of global reset found in traditional solutions.
[0034] The tiered handling execution module receives tiered handling instructions and performs tiered repair operations on the signal processing SRAM-type FPGA through a dedicated interface. This module connects to the configuration memory of the SRAM-type FPGA via the SelectMap interface and supports three tiered repair modes: dynamic refresh, full refresh, and global reset. It automatically matches the optimal repair method according to the severity of the fault, realizing progressive fault-tolerant processing from uninterrupted repair to full reconfiguration.
[0035] The service status awareness module acquires satellite transit status and communication equipment service load status, and feeds the perceived status data back to the decision control module. The decision control module then optimizes the timing of repair operations performed by the tiered handling execution module based on the status data. This module is the first to incorporate satellite orbital operation characteristics and communication service status into single-event soft error repair decision-making, enabling high-impact repair operations to automatically avoid peak service periods.
[0036] The aforementioned single-event soft error (SEE) classification monitoring and processing system for spaceborne communication devices achieves a complete processing chain from fault detection, cross-validation, classification decision-making, classification repair to adaptive optimization of timing by constructing a collaborative working system of multi-dimensional fault monitoring modules, decision control modules, classification handling execution modules, and service status awareness modules. The system comprehensively covers six types of complementary fault monitoring points, establishes a dual response mechanism combining fault-triggered immediate repair and timed classification repair, minimizes service interruptions through three classification repair modes—dynamic refresh, full refresh, and global reset—and automatically optimizes the timing of repair execution by utilizing satellite transit and service status awareness.
[0037] In one embodiment, the multi-dimensional fault monitoring module includes a phase-locked loop status monitoring unit, a communication link interaction monitoring unit, a service channel locking monitoring unit, a radio frequency link status comparison unit, a dual-machine uplink status comparison unit, and a configuration interface status monitoring unit.
[0038] The phase-locked loop (PLL) status monitoring unit is used to monitor the lock-in indicator status of the PLL inside the SRAM-type FPGA used for signal processing. Deployed within the monitoring Flash-FPGA, this unit reads the PLL lock-in status register value once per second via the SRAM-type FPGA's internal register interface, while simultaneously monitoring the PLL lock-in indicator pin level. If three consecutive reads show a lost-lock status (register value 0x00 and pin level low), a PLL lock-in failure is identified. This failure corresponds to a single-event upset causing a reversal of the PLL configuration parameters.
[0039] The communication link interaction monitoring unit is used to monitor the continuity of communication between the signal processing SRAM-type FPGA and the external spaceborne computer. This unit uses a CAN bus interface. Every second, the spaceborne computer sends a telemetry polling request frame to the signal processing SRAM-type FPGA. The frame format is 0xAA plus address code plus function code plus checksum, requesting the return of core operating parameters, including temperature, voltage, and operating status. If the signal processing SRAM-type FPGA does not receive a telemetry polling command from the spaceborne computer for five consecutive cycles, or fails to return a valid response frame to the spaceborne computer for five consecutive cycles, it is determined to be a CAN interface failure. This failure corresponds to a single event interruption (SEE) causing FPGA interface function interruption.
[0040] The service channel lock monitoring unit is used to monitor the validity of the carrier lock, bit synchronization, and frame synchronization indication signals output by the signal processing SRAM-type FPGA. This unit monitors the carrier lock, bit synchronization, and frame synchronization lock indication signals output by the signal processing SRAM-type FPGA in real time. These lock indication signals are three TTL level signals, containing the currently locked channel number (eight bits) and the signal strength (a sixteen-bit ADC sample value). If all service channels are not locked for two consecutive seconds and there is no valid signal strength output, a baseband processing fault is determined.
[0041] The RF link status comparison unit reads and compares the automatic gain control (AGC) value of the RF module with the baseband demodulation signal-to-noise ratio (SNR). The AGC value ranges from 0 to 2.5 volts, and the SNR ranges from 10 to 255. This unit reads the AGC and SNR values every second via the RF module's SPI interface. When the RF channel AGC value is within the medium-strong signal operating range, but the baseband signal processing SNR value is below 100 and remains below 100 for more than five seconds, a demodulation module fault is identified. This unit overcomes the limitation of traditional solutions where independent monitoring of AGC and SNR fails to detect latent anomalies in the demodulation module.
[0042] The dual-machine uplink status comparison unit is used to acquire the uplink status data of machines A and B respectively through the peer communication interface and compare the status differences between machines A and B. This unit exchanges uplink status data every 300 milliseconds through the communication interface between machines A and B. The data length is 32 bytes and includes RF channel receive AGC, signal processing SNR, and signal lock status. If the uplink status of the two machines is inconsistent and lasts for more than ten seconds, for example, if the SNR difference between the two machines exceeds ten decibels, it is determined that a single-event fault has occurred in the FPGA of one of the machines.
[0043] The configuration interface status monitoring unit is used to read the status register and configuration register data of the signal processing SRAM-type FPGA before configuration refresh. This unit reads the configuration status register and configuration register of the SRAM-type FPGA once before refreshing the configuration via the SelectMap interface. If the return value is inconsistent with the preset normal value, or if the interface times out without response after three consecutive reads, it is determined to be a single-event interrupt (SEE) fault of the FPGA dedicated interface. This fault corresponds to a single event causing the FPGA configuration circuit logic to fail.
[0044] Through the coordinated deployment of the above six monitoring units, this embodiment constructs a complementary fault monitoring system covering six categories: phase-locked loop (PLL) hardware status, CAN communication link, service channel locking, RF and baseband signal matching, dual-machine status cross-comparison, and FPGA configuration interface. From a monitoring perspective, the PLL status monitoring unit and configuration interface status monitoring unit cover the hardware circuit level; the communication link interaction monitoring unit covers the space-based communication link level; the service channel locking monitoring unit and RF link status comparison unit cover the service data processing link level; and the dual-machine uplink status comparison unit covers hidden anomalies that are difficult for a single machine to self-test through redundant comparison. The monitoring data from these six dimensions complement and corroborate each other, fundamentally solving the technical bottleneck of high false negative rates in existing single-monitoring systems.
[0045] In one embodiment, the decision control module includes a fault cross-verification unit, a hierarchical handling decision unit, and a timed task management unit. This module is deployed within the monitoring Flash-FPGA and implemented using a finite state machine (FSM). Its operating states include normal operation, fault detection, fault verification, repair execution, and repair verification.
[0046] The fault cross-verification unit is used to simultaneously retrieve the monitoring results of other monitoring units that have a physical logical association with the abnormal alarm when an abnormal alarm is received from any monitoring unit. A single-event fault is confirmed when all associated monitoring results meet preset abnormal conditions. For example, when the phase-locked loop (PLL) status monitoring unit alarms, the fault cross-verification unit simultaneously verifies the AGC / SNR status of the RF link status comparison unit and the service signal lock status of the service channel lock monitoring unit. If the AGC value is abnormal and the service signal is not locked, a PLL lockout fault is confirmed; if the AGC value is normal and the service signal is locked, it is determined to be a false alarm from the PLL status monitoring unit, and the false alarm is eliminated. This cross-verification mechanism effectively eliminates single-point misjudgments that may be caused by factors such as electromagnetic interference in the space environment.
[0047] The graded response decision unit is used to classify a single-event fault into three response levels after confirming the fault, and to generate repair instructions corresponding to the response levels.
[0048] The first-level instruction corresponds to scenarios where a single functional module fails without affecting the overall service chain. Examples include a demodulation module failure or a phase-locked loop (PLL) lockout failure. The first-level instruction instructs the tiered handling execution module to first send a reset signal to the faulty module, with a reset time of one hundred milliseconds, and simultaneously request immediate execution of a dynamic refresh of the configuration memory. The dynamic refresh only applies to configuration storage bits other than distributed RAM, shift register chain (SRL), and block RAM (BRAM). During the refresh process, other FPGA logic functions operate normally, and the service is completely uninterrupted. The total refresh time is approximately two thousand milliseconds.
[0049] The second-level instruction corresponds to scenarios where faults are still detected after dynamic refresh, or where multiple functional modules fail simultaneously without causing a complete FPGA crash. For example, a phase-locked loop (PLL) lockout combined with a CAN interface fault. The second-level instruction instructs the tiered handling execution module to immediately perform a full refresh operation on the SRAM-type FPGA configuration memory. The full refresh incrementally refreshes all configuration memory resources of the FPGA, including distributed RAM, shift register chain (SRL), and block RAM, page by page. It monitors the Flash-FPGA writing all configuration pages sequentially according to their addresses, with each page write time not exceeding one millisecond, and the total refresh time for the entire chip not exceeding three thousand milliseconds. During the refresh process, most FPGA logic functions are unaffected; only functional modules using distributed RAM, SRL, and BRAM resources will experience a brief interruption of no more than one millisecond, ensuring minimal service interruption.
[0050] The third-level instruction corresponds to faults that cannot be recovered after a full refresh, single-event interrupt faults of the FPGA dedicated interface, or scenarios where the entire FPGA crashes. The third-level instruction instructs the graded handling execution module to send a low-level reset signal to the reset pin of the signal processing SRAM-type FPGA. The reset duration is 200 milliseconds. Simultaneously, the complete configuration bitstream is reloaded via the SelectMap interface. The total reset and load time does not exceed five seconds, during which service is interrupted. The third-level handling is only executed when both dynamic refresh and full refresh fail.
[0051] The scheduled task management unit maintains scheduled repair tasks with preset cycles. Based on satellite transit and service load status feedback from the service status awareness module, it determines whether to adjust the triggering timing of scheduled repair tasks. Scheduled tasks include 12-hour periodic repair tasks and 8-minute periodic dynamic refresh tasks. The 8-minute periodic dynamic refresh tasks do not affect service operation and do not require delay. The 12-hour periodic repair tasks are executed by default during non-satellite transit periods; delay adjustments are triggered when the service status awareness module reports that the communication device is experiencing peak service.
[0052] The aforementioned decision control module employs a three-level fault classification mechanism to precisely correlate fault severity with repair strategies. From a technological perspective, traditional solutions typically use a single global reset strategy, initiating the same full-scale repair process regardless of whether the fault is a single-bit flip in a local register or an interruption of global control logic, resulting in numerous unnecessary service interruptions. This module, after eliminating false alarms through fault cross-validation, automatically handles faults in tiers based on their impact: over 90% of local functional module anomalies require only Level 1 handling for recovery, with zero service interruption; more severe anomalies occurring concurrently in multiple modules are escalated to Level 2 handling, resulting in only millisecond-level brief service jitter; a Level 3 global reset is triggered only in extreme cases such as when a full refresh fails or an FPGA crash. This tiered design achieves a dynamic optimal balance between comprehensive fault coverage, timely repair, and service continuity.
[0053] In one embodiment, the tiered processing execution module is deployed inside the monitoring Flash-FPGA and connected to the configuration memory of the signal processing SRAM-type FPGA via the SelectMap interface. It supports 8-bit and 16-bit parallel configuration modes, with a configuration clock frequency of 50 MHz. This module supports three tiered repair modes: dynamic refresh mode for the configuration memory, full refresh mode for the configuration memory, and global reset mode for the FPGA.
[0054] The execution process of the configuration memory dynamic refresh mode is as follows: The complete configuration bitstream of the FPGA is pre-stored in the Flash-FPGA. Refresh data is written to only the configuration storage bits, excluding distributed RAM, shift register chains, and BRAM, in address order via the SelectMap interface. During the dynamic refresh process, other FPGA logic functions operate normally, and services are completely uninterrupted. The total refresh time is approximately two thousand milliseconds.
[0055] The execution process of the full refresh mode for the configuration memory is as follows: A page-by-page incremental refresh is performed on all configuration memories of the FPGA, including distributed RAM, shift register chain (SRL), and BRAM. The Flash-FPGA is monitored to write all configuration pages sequentially according to their addresses, and the total refresh time for the entire chip does not exceed three thousand milliseconds. During the refresh process, most of the FPGA's logic functions are unaffected. Functional modules that only use distributed RAM, SRL, and BRAM resources will experience a brief interruption of no more than one millisecond, ensuring minimal service interruption.
[0056] The execution process of the FPGA global reset mode is as follows: the monitoring Flash-FPGA sends a low-level reset signal to the reset pin of the signal processing SRAM-type FPGA. The reset duration is 200 milliseconds. At the same time, the complete configuration bitstream is reloaded through the SelectMap interface. The total time for reset and loading does not exceed five seconds, during which the service is interrupted. This mode is only executed when both dynamic refresh and full refresh fail.
[0057] In this embodiment, three repair modes form a progressive repair chain from light to heavy. Compared with the single full-chip global reset method in the prior art, this module precisely selects the repair granularity according to the hierarchical instructions issued by the decision control module: the first-level local refresh only refreshes the configuration logic area that is susceptible to single-event upsets, without touching the working distributed RAM and BRAM data areas; the second-level full refresh completes the repair of the entire configuration area while maintaining the continuity of most services; the third-level global reset is only activated in a very few unrecoverable scenarios. This progressive design enables more than 90% of single-event soft errors to be repaired under zero service interruption conditions, significantly improving the on-orbit availability of the spaceborne communication device.
[0058] In one embodiment, the service status awareness module is deployed inside the monitoring Flash-FPGA. It receives satellite ephemeris data from the satellite computer via the CAN bus external interface and obtains service status information from the communication unit via the user management interface. This module includes a satellite transit status judgment unit and a service load status judgment unit.
[0059] The satellite transit status determination unit calculates the relative elevation angle between the satellite and the ground station based on ephemeris data, and determines that the satellite is currently in transit status when the relative elevation angle meets a preset coverage threshold. When the satellite elevation angle is greater than 10 degrees, it is determined to be in transit status; when the satellite elevation angle is less than 5 degrees, it is determined to be in non-transit status.
[0060] The service load status judgment unit is used to determine whether the communication device is in a preset peak service period state based on the uplink signal lock status, automatic gain control value, demodulation signal-to-noise ratio, and downlink transmitter switch status. When the uplink signal is locked, the AGC value is in the range of 0.3V to 2.5V, the SNR value is greater than 20, and the downlink transmitter switch is in the on state, the communication device is determined to be in a peak service period.
[0061] When a 12-hour scheduled repair task or a configuration memory full refresh task is triggered, and the service status awareness module reports that the communication device is in peak service conditions, the decision control module issues a delayed execution instruction to the hierarchical handling execution module. The delay time is half an hour, during which the status data fed back by the service status awareness module is continuously monitored. If the service is still in peak service conditions after half an hour, the delay continues, with a maximum delay of one hour. If an immediate emergency repair request is triggered by the multi-dimensional fault monitoring module during the delay period, the decision control module prioritizes outputting the corresponding immediate repair instruction to the hierarchical handling execution module. The delay mechanism does not apply to cases of uplink and downlink service function anomalies; in such cases, a fault-triggered immediate refresh is triggered.
[0062] The innovation of this embodiment lies in the first-time integration of satellite orbital operating parameters and real-time communication service status into the decision-making process for single-event soft fault (SOF) repair. The transit time of a single ground station for a low-Earth orbit satellite is only about ten minutes, a peak service period. Existing timed reset technologies do not differentiate between satellite operating states, often performing high-impact repair operations during peak service periods, leading to simultaneous disconnections for a large number of users. This module, through dual sensing of satellite transit status and service load status, ensures that high-impact repair operations—i.e., full refresh and global reset—automatically avoid peak service periods. Even if a new SOF occurs during the delay period, fault-triggered immediate dynamic refresh takes priority, ensuring both timely fault repair and maximizing the protection of communication service quality during the transit period.
[0063] In one embodiment, the collaborative working mechanism implemented by the functional modules through the internal high-speed bus fully describes the complete processing flow of the system from fault detection to repair closed loop.
[0064] After the system is powered on, the monitoring Flash-FPGA first completes its own initialization, and then loads the configuration bitstream to the signal processing SRAM-type FPGA through the SelectMap interface. After the loading is completed, the system enters normal operation.
[0065] The multi-dimensional fault monitoring module collects fault monitoring data from various dimensions according to a preset cycle and outputs it to the decision control module. In the preset cycle, the phase-locked loop status monitoring is 1 second, the CAN bus interaction monitoring is 1 second, the AGC-SNR monitoring is 1 second, the A / B machine uplink status comparison is 300ms, and the service channel locking monitoring is continuous real-time monitoring.
[0066] The decision control module receives monitoring data in real time and continuously scans the data status of each monitoring point under normal operating conditions. When a suspected abnormal state is detected, the decision control module jumps from the normal operating state to the fault verification stage, comparing other monitoring data associated with the suspected abnormal state to confirm whether a real fault has occurred.
[0067] After confirming the fault, the decision control module selects the corresponding tiered handling instruction based on the fault's impact scope and severity, transitioning from the fault verification state to the repair execution state, and transmitting the tiered handling instruction to the tiered handling execution module. After completing the tiered repair operation, the tiered handling execution module sends a completion signal back to the decision control module, which then transitions to the repair verification state, triggering the multi-dimensional fault monitoring module to retest the repaired data to verify its effectiveness. If the repair verification is successful, the system returns to normal operation; if the repair verification fails, the decision control module automatically escalates the repair level (e.g., from level one to level two, level two to level three) and outputs the corresponding repair instruction again.
[0068] Throughout the operation, the service status awareness module feeds back the satellite transit status and service load status to the decision control module in real time, which is used to adjust the timing of the output of tiered handling instructions. When the 12-hour periodic repair task or the full refresh task is triggered, the decision control module decides whether to delay the output of the repair instruction based on the current status fed back by the service status awareness module.
[0069] The aforementioned collaborative working mechanism realizes a complete automated processing chain from fault detection to repair: monitoring, verification, classification, execution, retesting, and escalation when necessary, with each link forming a closed-loop feedback with the preceding and following links. Traditional solutions typically separate fault monitoring and fault repair into two independent subsystems, issuing only alarms after a fault is detected, while repair operations rely on manual intervention on the ground. This system tightly couples the four major functional modules through an internal high-speed bus, automating the entire process from fault detection to repair completion with millisecond-level response, completely eliminating the time delay caused by manual intervention.
[0070] System Validation and Effect Comparison To verify the beneficial effects of the technical solution of this application compared with the prior art, based on the typical working environment and single-event effect test conditions of low-orbit satellite communication equipment, a comparative test was conducted on the existing single monitoring system and the multi-dimensional fault monitoring system of this application. At the same time, a comparative test was conducted on the existing single full reset repair strategy and the three-level graded repair strategy of this application.
[0071] Table 1 Comparison of Fault Detection Capabilities
[0072] Table 2 Comparison of Repair Modes and Business Impact
[0073] Table 3 Comparison of Overall System Performance
[0074] The comparison results in Tables 1, 2, and 3 show that the technical solution proposed in this application is significantly superior to existing solutions in terms of comprehensive fault monitoring coverage, timely repair response, minimal service interruption, and intelligent selection of repair timing. The multi-dimensional fault monitoring system reduces the fault miss rate from over 35% to below 2%, the three-level graded repair strategy enables over 90% of single-event soft errors to be repaired with zero service interruption, and the satellite transit and service status awareness mechanism ensures that high-impact repair operations are automatically avoided during peak business periods.
[0075] The above verification results show that this application constructs a multi-dimensional fault monitoring system covering hardware status, service links, and FPGA dedicated interfaces, establishes a dual response mechanism combining fault-triggered immediate repair and timed hierarchical repair, realizes three hierarchical repair modes: dynamic refresh, full refresh, and global reset, and introduces a satellite transit and service status awareness mechanism to optimize the timing of repair execution. It systematically solves the core technical problems in the existing technology, such as the high missed detection rate of the single fault monitoring system, the passive and delayed response of the refresh mechanism, the single repair mode with a large impact on services, and the unreasonable refresh timing. It significantly improves the on-orbit reliability and service quality of the spaceborne communication device.
[0076] It should be understood that, although Figure 1 The processes are shown sequentially as indicated by the arrows, but these processes are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order in which these processes are executed; they can be executed in other orders. Furthermore, Figure 1At least a portion of the process may include multiple sub-processes or multiple stages. These sub-processes or stages are not necessarily executed at the same time, but may be executed at different times. The execution order of these sub-processes or stages is not necessarily sequential, but may be executed in turn or alternately with other processes or at least a portion of the sub-processes or stages of other processes.
[0077] In one embodiment, a single-event soft error classification monitoring and processing device for a spaceborne communication device is provided, comprising: a multi-dimensional fault monitoring module, a decision control module, a classification processing execution module, and a service status perception module.
[0078] The multi-dimensional fault monitoring module is used to collect fault monitoring data from multiple dimensions in real time from the signal processing SRAM-type FPGA, RF module and peer communication interface through a dedicated interface, and output single-event soft error monitoring results.
[0079] The decision control module is used to receive monitoring results, cross-validate the monitoring results to eliminate false alarms, and generate graded handling instructions based on the error impact range and severity of single-event soft errors.
[0080] The graded processing execution module is used to receive graded processing instructions and perform graded repair operations on the signal processing SRAM-type FPGA through a dedicated interface.
[0081] The service status awareness module is used to acquire the satellite transit status and the service load status of the communication device, and feeds the perceived status data back to the decision control module. The decision control module optimizes the timing of the hierarchical handling execution module to perform repair operations based on the status data.
[0082] For specific limitations regarding the aforementioned device, please refer to the system limitations described above, which will not be repeated here. Each module in the aforementioned device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware within or independently of the processor in the computer device, or stored in software within the memory of the computer device, so that the processor can invoke and execute the operations corresponding to each module.
[0083] In one embodiment, a computer device is provided, which may be the processing platform where the monitoring Flash-FPGA of a spaceborne communication device is located, and its internal structure diagram may be as follows: Figure 2As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores FPGA configuration bitstream data and fault monitoring history data. The network interface communicates with the spaceborne computer via a CAN bus. When executed by the processor, the computer program implements the system functions described in any of the above embodiments.
[0084] Those skilled in the art will understand that Figure 2 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0085] In one embodiment, a computer device is provided, which may be a ground test platform or a spaceborne computer, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the system functions of any of the above embodiments.
[0086] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the system functions of any of the above embodiments.
[0087] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0088] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0089] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A self-monitoring and hierarchical single-event soft error handling system for spaceborne communication equipment, characterized in that, The system adopts an A / B dual-machine hot backup hardware architecture. Both machine A and machine B include a monitoring Flash-FPGA and a signal processing SRAM-type FPGA. The monitoring Flash-FPGA establishes a physical connection with the signal processing SRAM-type FPGA, the radio frequency module, and the peer communication interface through a dedicated interface. The monitoring Flash-FPGA is internally deployed with a multi-dimensional fault monitoring module, a decision control module, a hierarchical handling and execution module, and a business status perception module. Each module is interconnected through an internal high-speed bus and forms a collaborative working mechanism. The multi-dimensional fault monitoring module is used to collect fault monitoring data from the signal processing SRAM-type FPGA, the radio frequency module and the peer communication interface in real time through the dedicated interface, and send the single-event soft error monitoring results to the decision control module. The decision control module is used to receive the monitoring results, cross-validate the data from multiple related sources in the monitoring results, generate graded handling instructions based on the error impact range and severity of single-event soft errors, and output the graded handling instructions to the graded handling execution module. The graded processing execution module is used to receive the graded processing instruction and perform graded repair operations on the signal processing SRAM type FPGA through the dedicated interface; The service status perception module is used to acquire the satellite transit status and the service load status of the communication device, and feed the perceived status data back to the decision control module. The decision control module optimizes the timing of the hierarchical handling execution module to perform repair operations based on the status data.
2. The system according to claim 1, characterized in that, The multi-dimensional fault monitoring module includes a phase-locked loop status monitoring unit, a communication link interaction monitoring unit, a service channel locking monitoring unit, a radio frequency link status comparison unit, a dual-machine uplink status comparison unit, and a configuration interface status monitoring unit. The phase-locked loop status monitoring unit is used to monitor the phase-locked loop and the local oscillator lock-in indication status of the signal processing SRAM-type FPGA internal phase-locked loop and the RF channel local oscillator lock-in indication status. The communication link interaction monitoring unit is used to monitor the continuity of communication interaction between the signal processing SRAM-type FPGA and the external spacecraft computer. The service channel lock monitoring unit is used to monitor the effectiveness of the carrier lock, bit synchronization, and frame synchronization indication signals output by the signal processing SRAM-type FPGA. The RF link status comparison unit is used to read and compare the automatic gain control value of the RF module with the baseband demodulation signal-to-noise ratio result; The dual-machine uplink status comparison unit is used to obtain the uplink status data of machine A and machine B respectively through the peer communication interface, and compare the status differences between machine A and machine B. The configuration interface status monitoring unit is used to read the status register and configuration register data of the signal processing SRAM FPGA before configuration refresh.
3. The system according to claim 1, characterized in that, The decision control module includes a fault cross-verification unit, a hierarchical handling decision unit, and a timed task management unit. The fault cross-verification unit is used to synchronously call the monitoring results of other monitoring units that have physical logical association with the fault alarm when receiving an abnormal alarm from any monitoring unit, and confirm the occurrence of a single-event fault when all associated monitoring results meet the preset abnormal conditions. The hierarchical handling decision unit is used to divide the fault into three handling levels after confirming a single-event fault, and generate repair instructions corresponding to the handling levels; wherein, the first level instruction is used to instruct the hierarchical handling execution module to perform a local module reset and immediately trigger a dynamic refresh, the second level instruction is used to instruct the execution of a full refresh of the configuration memory, and the third level instruction is used to instruct the execution of a global reset and reload the configuration bit stream. The scheduled task management unit is used to maintain scheduled repair tasks with a preset cycle, and to determine whether to adjust the triggering timing of the scheduled repair tasks based on the satellite transit and service load status fed back by the service status perception module.
4. The system according to claim 1, characterized in that, The hierarchical processing execution module is connected to the configuration memory of the signal processing SRAM-type FPGA through the SelectMap interface, and is used to configure the dynamic refresh mode, full refresh mode, and global reset mode of the configuration memory. The configuration memory dynamic refresh mode is used to dynamically refresh other configuration memory bits in the signal processing SRAM type FPGA, excluding distributed RAM, shift register chain and BRAM. The configuration memory full refresh mode is used to perform page-by-page incremental refresh of all configuration memory bits of the signal processing SRAM type FPGA. The FPGA global reset mode is used to send a reset signal to the reset pin of the signal processing SRAM type FPGA and control the signal processing SRAM type FPGA to reload the complete configuration bit stream.
5. The system according to claim 1, characterized in that, The service status perception module includes a satellite transit status judgment unit and a service load status judgment unit. The satellite transit status determination unit is used to calculate the relative elevation angle between the satellite and the ground station based on ephemeris data, and to determine that the satellite is currently in transit status when the relative elevation angle meets a preset coverage threshold. The service load status judgment unit is used to determine whether the communication device is in a preset peak service period state based on the uplink signal lock status, automatic gain control value, demodulation signal-to-noise ratio, and downlink transmitter switch status of the communication device.
6. The system according to claim 3, characterized in that, The decision control module optimizes the timing of repair execution based on the status data as follows: When a pre-set periodic repair task or a full refresh task of the configuration memory is triggered, and the service status perception module reports that the communication device is in a peak service period, a delayed execution instruction is sent to the hierarchical processing execution module. During the delay, the status data fed back by the service status awareness module is continuously monitored. Once the status data is restored to the service idle state, the pending repair operation is executed. If an immediate emergency repair request is triggered by the multi-dimensional fault monitoring module during the delay period, the decision control module will prioritize outputting the corresponding immediate repair instruction to the hierarchical handling execution module.
7. The system according to claim 1, characterized in that, The collaborative working mechanism implemented by each functional module through the internal high-speed bus is as follows: The multi-dimensional fault monitoring module collects monitoring data of faults in each dimension according to a preset cycle and outputs it to the decision control module; The decision control module receives the monitoring data in real time and jumps to the verification stage when a suspected abnormal state is detected. It compares the other monitoring data associated with the suspected abnormal state to confirm whether a real fault has occurred. After confirming the fault, the decision control module selects the corresponding graded handling instruction based on the scope and severity of the fault and transmits it to the graded handling execution module; After the graded treatment execution module completes the graded repair operation, it sends a completion signal to the decision control module. The decision control module then triggers the multi-dimensional fault monitoring module to retest the repaired data to verify the effectiveness of the repair. If the verification and repair fail, the decision control module will automatically upgrade the repair level and output the corresponding repair command again; The service status perception module feeds back the satellite transit status and service load status to the decision control module in real time, which is used to adjust the timing of the output of the graded handling instructions.
8. A self-monitoring and hierarchical single-event soft error handling device for spaceborne communication equipment, characterized in that, The device includes: The multi-dimensional fault monitoring module is used to collect fault monitoring data from multiple dimensions from the signal processing SRAM-type FPGA, RF module and peer communication interface in real time through a dedicated interface, and output single-event soft error monitoring results. The decision control module is used to receive the monitoring results, perform cross-validation on the monitoring results to eliminate false alarms, and generate graded handling instructions based on the error impact range and severity of single-event soft errors. The graded processing execution module is used to receive the graded processing instructions and perform graded repair operations on the signal processing SRAM type FPGA through the dedicated interface; The service status perception module is used to acquire the satellite transit status and the service load status of the communication device, and feed the perceived status data back to the decision control module. The decision control module optimizes the timing of the hierarchical handling execution module to perform repair operations based on the status data.