Zero-interruption recovery method based on signal full-link level lockstep running state mirror redundancy

CN122570252BActive Publication Date: 2026-09-11HUNAN SIBEITU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611045915.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-11
Estimated Expiration
2046-07-14

AI Technical Summary

Technical Problem

[0005]然而,现有技术均围绕“故障发生后处置恢复”的核心思路设计,完全未考虑星载硬实时任务的零中断、确定性时延约束,在实际工程应用中存在无法克服的固有缺陷

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122570252B_ABST
    Figure CN122570252B_ABST
Patent Text Reader

Abstract

This application relates to the field of radiation hardening technology for spaceborne electronic systems, and in particular to a zero-interruption recovery method based on signal end-link level lockstep operational state mirror redundancy. The method includes: using a homogeneous high-precision clock to drive two completely identical end-link signal processing systems to operate in clock-level lockstep mode, with completely homogeneous input excitation and real-time mirrored operational states; in normal mode, the main system outputs externally, while the mirror system provides hot backup; single-event soft errors and single-event functional interruption faults are identified through digital and radio frequency full-dimensional real-time monitoring; fault isolation and seamless link switching are completed at the nanosecond level by pure hardware logic; after switching, the background system performs hierarchical recovery of the faulty link and synchronizes the full operational state, reconstructing the dual-path lockstep redundancy mode. This method can achieve zero-interruption, non-disruptive recovery of hard real-time tasks, with high latency determinism, no intrusive modification of components, and is adaptable to commercial aerospace COTS devices, suitable for hard real-time scenarios such as spaceborne communication and navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of radiation hardening technology for spaceborne electronic systems, and in particular to a zero-interruption recovery method based on signal end-link level lockstep operational state mirror redundancy. Background Technology

[0002] With the rapid advancement of low-Earth orbit satellite internet constellation construction and deep space exploration projects, as well as the continuous improvement of the autonomy and intelligence of onboard systems, the proportion of tasks undertaken by onboard electronic systems, such as real-time onboard power supply communication, real-time communication of high-speed service signal links, and onboard autonomous navigation, has increased significantly. These tasks all fall under the category of hard real-time tasks, with strict deterministic time constraints. They must be completed within a preset fixed deadline. Any execution interruption, excessive latency jitter, or calculation errors will directly lead to user task failure, and may even cause catastrophic failures such as spacecraft attitude loss, interruption of satellite-to-ground communication links, and loss of inter-satellite communication. This places extremely high demands on the system's continuous operation capability and radiation resistance reliability.

[0003] The space environment contains high-energy radiation sources such as Earth's radiation belts trapping particles, solar cosmic rays, and galactic cosmic rays. The single-event effects (SEEs) they induce are the primary cause of unplanned resets and functional failures in spacecraft electronic systems, accounting for over 60% of on-orbit anomalies. As semiconductor processes iterate towards deep submicron and nanometer nodes, critical device dimensions continue to shrink, and operating voltages decrease, leading to an exponential decrease in the critical charge of sensitive nodes and a significant increase in sensitivity to SEEs. Two types of recoverable radiation effects without permanent physical damage pose the core threats to the reliable operation of hard real-time missions in orbit: one is single-event soft errors, which refer to recoverable errors caused by high-energy charged particles bombarding sensitive nodes of semiconductor devices. Core examples include single-event flips and single-event transients, which can lead to register state flips, data anomalies, and disordered logic execution flow, causing errors in hard real-time mission calculations and process out of control. This is the main source of silent failures in spacecraft systems in orbit. Another type is single-event interrupt, a special type of single-event effect. It refers to the bombardment of the core control module of a device by high-energy particles, causing the device to enter an unexpected abnormal working state. There is no permanent hardware damage, but it will directly cause the device to malfunction and the system to be suspended. It is a key source of failure for sudden on-orbit failure of hard real-time tasks.

[0004] To address the two types of radiation effects mentioned above, existing technologies have developed a series of protection and recovery solutions, which have been applied to a certain extent in aerospace engineering: For single-event soft error protection, the mainstream hardware solutions include triple-modular redundancy (TMR), partial triple-modular redundancy (PTMR), DICE hardened triggers, and SEC-DED / ECC error correction coding. The mainstream software solutions include periodic cache refresh, checkpoint rollback recovery, redundancy calculation, and code signature verification. For SEFI monitoring and recovery, the mainstream monitoring solutions include independent hardware watchdogs, FPGA configuration completion DONE signal level monitoring, configuration frame readback CRC verification, periodic inspection of key registers, and two-level telemetry anomaly detection (both system-wide and single-machine levels). The mainstream recovery strategy is a tiered reset approach: minor anomalies trigger local module resets / dynamic configuration refreshes; severe anomalies trigger system-wide hot resets / master / standby switching; and extreme anomalies trigger power-off restarts.

[0005] However, existing technologies are all designed around the core idea of ​​"fault handling and recovery", and do not take into account the zero-interruption and deterministic latency constraints of spaceborne hard real-time tasks. In practical engineering applications, they have inherent defects that cannot be overcome.

[0006] Specifically, existing single-event soft error protection schemes cannot simultaneously address fault tolerance and the deterministic timing requirements of real-time tasks, resulting in unavoidable risks of task interruption and latency exceeding limits. TMR / PTMR can only shield single-bit SEUs. For logic anomalies caused by multi-bit flips, cascaded SETs, and illegal state machine transitions, task re-execution and register reloading must be triggered, directly leading to the interruption of hard real-time task execution flow. ECC error correction can only be effective for memory SEUs and cannot cover logic circuit SETs or control logic SEUs. Furthermore, the error correction process introduces memory access latency jitter, disrupting the deterministic timing of hard real-time systems. Software checkpoint rollback recovery requires interrupting normal task execution, and the rollback process after an error occurs will directly lead to task interruption, making it completely unsuitable for continuously running hard real-time scenarios.

[0007] Existing SEFI recovery solutions all rely on interrupting core tasks, presenting inherent technical bottlenecks and failing to achieve continuous operation of hard real-time tasks. Partial reset / partial reconfiguration solutions interrupt the real-time data stream of the corresponding module, leading to communication frame loss and PLL lockout, directly causing hard real-time task failure. Global reset / system reconfiguration solutions have latency in the millisecond or even second range, completely interrupting hard real-time tasks. Primary / backup redundancy switching solutions employ a passive "fault-based switching" mode, with switching latency in the microsecond to millisecond range throughout the process, and issues such as data stream interruptions and asynchrony between primary and backup states. In warm / cold backup modes, task states need to be reinitialized after switching, further exacerbating task interruptions.

[0008] Existing technologies cannot achieve coordinated handling of single-event soft errors and SEFI, posing risks of fault propagation and secondary interruptions, and failing to meet the end-to-end protection requirements of hard real-time scenarios. Most existing solutions are designed independently for single-type single-event effects, with soft error protection and SEFI recovery being isolated from each other. When soft errors cascade and trigger SEFI, only a step-by-step reset operation can be triggered, leading to secondary task interruptions. At the same time, it is impossible to accurately distinguish between SEFI and functional anomalies caused by total dose effect (TID), single-event lockout (SEL), and electromagnetic interference (EMI), resulting in frequent false resets. The misjudgment rate of existing engineering solutions can reach more than 30%, seriously disrupting the continuous operation of hard real-time tasks.

[0009] Furthermore, existing technologies are ill-suited to the application requirements of emerging hard real-time scenarios on spacecraft and COTS devices in commercial aerospace, leaving a significant technological gap. Emerging hard real-time tasks such as onboard AI real-time inference, inter-satellite laser communication, and distributed constellation collaborative control have requirements for mission continuity and deterministic latency at the nanosecond to microsecond level, which current technologies cannot fully meet. At the same time, industrial-grade COTS devices widely used in commercial aerospace have soft error and SEFI rates more than 10 times higher than those of aerospace-grade devices. Existing aerospace-grade protection solutions suffer from high costs, high power consumption, and the need for intrusive modifications to the internal design of the devices, making them unsuitable for the low-power, miniaturized, and non-intrusive application requirements of commercial aerospace.

[0010] In summary, existing technologies have been unable to resolve the core contradiction between high reliability against single-event faults and zero-interruption continuous operation of hard real-time missions. There is currently no solution that can achieve in-situ transparent fault tolerance for single-event soft errors and seamless, non-disruptive recovery of SEFI, and thus cannot meet the high reliability and high continuity requirements of current and future spaceborne hard real-time missions. Summary of the Invention

[0011] Therefore, it is necessary to provide a zero-interruption recovery method based on signal end-link level lockstep running state mirror redundancy to address the above-mentioned technical problems. This method can achieve zero-interruption, non-disruptive recovery from single-event soft errors and single-event functional interruptions, and ensure the continuous and stable operation of spaceborne hard real-time tasks.

[0012] A zero-interruption recovery method based on full-link-level lockstep operational state mirror redundancy is applied to a spaceborne hard real-time signal processing system. The spaceborne hard real-time signal processing system includes a main operational system with full-link-level lockstep and a mirror operational system with full-link-level lockstep. The method includes: Step 1: After the onboard hard real-time signal processing system is powered on, the same-source high-precision clock synchronization unit starts up and outputs the same-source synchronized system clock and synchronization trigger signal to the main operating system, the mirror operating system and the global monitoring and execution module, completing the hardware initialization, logic loading and configuration parameter writing of the dual-path system, so that the initial operating state of the main operating system and the mirror operating system is completely consistent. Step 2: The main operating system and the mirror operating system enter a clock-level lockstep parallel operation state driven by the same system clock. They receive the same source input service signals, perform the same signal processing operations on the same clock edge, and generate mirror-consistent operating status data and output service signals in real time. In normal operation mode, the output terminal of the main operating system is turned on and outputs service signals to the outside, while the output terminal of the mirror operating system is in a high-impedance isolation state and does not output to the outside. Step 3: The global monitoring and execution module, based on the synchronous trigger signal, collects the digital link operating status and RF link operating parameters of the main operating system and the mirror operating system, as well as the output service signals, clock cycle by clock cycle. At each clock edge, it compares the operating status of the dual-path system bit by bit to determine whether a single-event soft error or a single-event function interruption fault has occurred; if there is no fault, it returns to step 2. Step 4: When a fault is detected in the main operating system or the mirror operating system, the global monitoring and execution module uses pure hardware logic circuits to synchronously switch the state of the two output paths on the same clock edge, shutting off the conduction of the output end of the faulty link, isolating the faulty link, and at the same time turning on the output end of the normal link, so that the normal link can seamlessly take over the output of external business signals. During the switching process, the external business signals are continuous and uninterrupted. Step 5: After the switchover is completed, the global monitoring and execution module performs a graded recovery operation on the faulty link in the background. During the recovery process, the normal link continues to carry out external services. After the faulty link is restored to normal, the current complete running state data of the normal link is batch mirrored and synchronized to the restored faulty link, so that the running state of the two systems is completely consistent. Step 6: After completing the running state mirror synchronization, the global monitoring and execution module controls the recovered faulty link and normal link to re-enter the clock-level lockstep running state through the same source synchronization trigger signal, restore the dual-path mirror redundancy backup mode, and return to step 2.

[0013] The aforementioned zero-interruption recovery method based on the full-link-level lockstep operation state mirror redundancy, in this application, achieves high-precision clock synchronization of the same source in step 1, providing a zero-deviation clock with a phase deviation of less than one clock edge setup / hold time window for the dual-path system, ensuring the consistency of the clock reference for subsequent lockstep operation; step 2 enables the main operating system and the mirror operating system to achieve clock-level lockstep parallel operation across the entire link, with real-time complete mirroring of the dual-path system's operating state, data state, and control state, ensuring that the dual-path system states are completely consistent at any time, with no timing deviations or state differences, providing a prerequisite for state consistency for subsequent zero-interruption switching; step 3 performs full-dimensional real-time monitoring of the dual-path system's digital and RF links clock-cycle based on the synchronization trigger signal, enabling the detection of dual-path state inconsistencies and the location of faulty link nodes and fault types within one clock cycle after a fault occurs, achieving early fault detection and early fault location; utilizing step 4 in... When a fault is detected, the operation of shutting down the faulty link and turning on the normal link is synchronously executed on the same clock edge through pure hardware logic circuits. The entire switching process is software-free, avoiding the latency uncertainty caused by software interruptions. The switching latency is only in the nanosecond range. There are no phase jumps, amplitude changes, or data frame drops in external service signals, achieving true zero-interruption seamless takeover and completely avoiding the execution interruption and latency jitter of hard real-time tasks. Step 5: After the switch is completed, while the normal link continues to carry the service, the faulty link is restored in stages in the background. The restoration operation does not affect the continuous operation of the foreground service. After the faulty link is restored to normal, the full link operation state of the normal link is completely mirrored and synchronized to the restored faulty link to ensure that the dual-path system state is completely consistent. Step 6: The restored dual-path system re-enters the clock-level lockstep operation state, restores the dual-path mirrored redundancy backup mode, completes the complete fault handling closed loop, and prepares for the next fault. The entire process forms a complete closed loop from fault detection, zero-interruption switch to background recovery and re-locking, fundamentally solving the fundamental defect of mission interruption caused by the "fault handling and recovery" of existing technologies. It achieves in-situ transparent fault tolerance for single-event soft errors and nanosecond-level seamless and non-intrusive recovery of SEFI. At the same time, the non-intrusive architecture design does not require modification of the internal logic of the device and can be directly adapted to aerospace COTS devices. Attached Figure Description

[0014] Figure 1 This is a block diagram of the onboard hard real-time signal processing system in one embodiment; Figure 2 This is a flowchart illustrating a zero-interruption recovery method based on signal end-link level lockstep running state mirror redundancy in one embodiment; Figure 3 This is a timing diagram for zero-interruption switching in one embodiment. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0016] The zero-interruption recovery method based on signal end-link level lockstep operational state mirror redundancy provided in this application can be applied to, for example... Figure 1 The illustrated spaceborne hard real-time signal processing system is divided into three layers from top to bottom: the service signal end-to-end layer, the monitoring and execution layer, and the clock synchronization layer. The core units include a homogeneous high-precision clock synchronization unit, an end-to-end lockstep main operation system, an end-to-end lockstep mirror operation system, and a global monitoring and execution module, along with supporting RF self-calibration coupling monitoring branches and cross-machine collaborative monitoring units. The system adopts a three-layer architecture: dual-path end-to-end homogeneous clock lockstep operation, hardware-level nanosecond-level fault bypass switching, and background fault autonomous recovery and resynchronization. The core achieves seamless takeover with zero interruption in the event of a fault through end-to-end clock-level synchronization mirroring. After fault switching, the background autonomously recovers and resynchronizes, ensuring continuous service operation throughout the process.

[0017] In one embodiment, such as Figure 2 As shown, a zero-interruption recovery method based on signal end-link level lockstep running state mirror redundancy is provided, including the following steps: Step 1: After the onboard hard real-time signal processing system is powered on, the same-source high-precision clock synchronization unit starts up and outputs the same-source synchronized system clock and synchronization trigger signal to the main operating system, the mirror operating system and the global monitoring and execution module, completing the hardware initialization, logic loading and configuration parameter writing of the dual-path system, so that the initial operating state of the main operating system and the mirror operating system is completely consistent.

[0018] The homogeneous high-precision clock synchronization unit is the foundation for clock-level lockstep operation in the dual-path system. This unit uses a single aerospace-grade high-stability crystal oscillator as the sole clock source, splitting it into two completely equal-length, equal-phase differential clock signals via a low-jitter clock buffer. These signals are then fed into the clock inputs of the FPGA / SoC, ADC / DAC, and RF chips in the main operating system and the mirror operating system, respectively. The clock PCB traces employ a strict equal-length design, with trace length deviations controlled within 50ps, ensuring the phase difference between the two clocks is less than the setup / hold time window of one clock edge. Simultaneously, a homogeneous synchronous trigger clock is output to the global monitoring and execution module, serving as a synchronization reference for dual-path system state comparison and fault monitoring, ensuring all monitoring, comparison, and switching actions are completed on the synchronous clock edge. Complete consistency of the initial states of the two systems is a prerequisite for subsequent lockstep operation, ensuring that from the initial moment, the operating state, data state, and control state of the two systems are completely mirrored, with no timing deviations or state differences.

[0019] Step 2: The main operating system and the mirror operating system enter a clock-level lockstep parallel operation state driven by the same system clock. They receive the same source input service signals, perform the same signal processing operations on the same clock edge, and generate mirror-consistent operating status data and output service signals in real time. In normal operation mode, the output terminal of the main operating system is turned on and outputs service signals to the outside, while the output terminal of the mirror operating system is in a high-impedance isolation state and does not output to the outside.

[0020] The dual-path system employs identical hardware design, logic code, configuration parameters, and PCB layout, covering the end-to-end full signal link for spaceborne communication / data processing tasks. This includes: a baseband data processing module (FPGA / SoC), a digital-to-analog converter (DAC), an RF transmit link, an RF receive link, an analog-to-digital converter (ADC), and a baseband demodulation and decoding module, with no bottlenecks in the link. The dual-path system adopts a parallel synchronous lock-step operation mode. Driven by the same clock source, it executes identical instructions, processes identical input data, and updates identical internal register states and output signals on the same clock edge, ensuring that the operating state, data state, and control state of the dual-path system are completely mirrored at any given time, with no timing deviations or state differences. The input signals of the dual-path system are completely from the same source. Service input data, external received signals, and configuration instructions are all simultaneously sent to the corresponding input terminals of the main operating system and the mirror operating system through the same driver circuit, ensuring 100% consistency of input stimuli between the two systems. The outputs of both systems are connected to the nanosecond-level switching control and fault isolation submodule. In normal operation mode, the output of the main operating system is turned on and outputs service signals to the outside; the output of the mirror system is in a high-impedance isolation state and is only used as a real-time hot mirror backup. It does not output to the outside to avoid signal conflicts.

[0021] Step 3: The global monitoring and execution module collects the digital link operating status and RF link operating parameters of the main operating system and the mirror operating system, as well as the output service signals, on a clock cycle based on the synchronous trigger signal. At each clock edge, it compares the operating status of the dual-path system bit by bit to determine whether a single-event soft error or single-event function interruption fault has occurred; if there is no fault, it returns to step 2.

[0022] The global monitoring and execution module is implemented using pure hardware logic on an FPGA, without software intervention, ensuring nanosecond-level response speed. The global monitoring and execution module comprises three core sub-modules: a real-time end-to-end status monitoring sub-module, a nanosecond-level switching control and fault isolation sub-module, and a fault recovery and resynchronization sub-module. The real-time end-to-end status monitoring sub-module enables real-time fault detection and location across the entire dual-path system without blind spots, and is divided into three dimensions: digital link monitoring, radio frequency link monitoring, and cross-machine collaborative monitoring.

[0023] In terms of digital link monitoring, the system collects baseband processing output data, key register status, state machine operation status, interface communication signals, and configuration register values ​​of the dual-channel system in real time. Bit-by-bit comparison is performed at each clock edge. When inconsistencies occur between the two channels' data / status, a fault flag is immediately triggered. At the same time, the system locates the link node where the fault occurred and the fault type - single-event soft error or SEFI. The detection delay is ≤10 system clock cycles.

[0024] In terms of RF link monitoring, the RF transmission signals of the main operating system and the mirror system are coupled to the corresponding receiving branches through equal-length coupling lines via RF self-calibration coupling circuits to form a self-loop monitoring link. The FPGA performs demodulation, frequency offset detection, power detection, spectrum analysis, and EVM-error vector amplitude calculation on the loop signal in real time. When the RF signal parameters of a certain link exceed the preset threshold, it is determined that the link has an SEFI or soft error, triggering a fault flag. The detection delay is ≤1μs.

[0025] For cross-machine collaborative monitoring, for electrical / optical data communication links with external single machines, the integrity and correctness of data sent by the master / mirror systems are compared in real time through the loopback verification mechanism of the external single machine, realizing fault monitoring of external communication links and avoiding missed faults at the end of the link. When no fault is detected, the dual-path system continues to maintain a lockstep operation state to ensure continuous and stable service output.

[0026] Step 4: When a fault is detected in the main operating system or the mirror operating system, the global monitoring and execution module uses pure hardware logic circuits to synchronously switch the state of the two output paths on the same clock edge, shutting off the conduction of the output end of the faulty link, isolating the faulty link, and at the same time turning on the output end of the normal link, so that the normal link can seamlessly take over the output of external service signals. During the switching process, the external service signals are continuous and uninterrupted.

[0027] The switching of the output path is executed by the nanosecond-level switching control and fault isolation submodule. This submodule is the core execution unit for zero-interrupt recovery, implemented using pure hardware logic circuits. Its core is a high-speed analog switch array combined with digital bypass control logic. The switching speed is ≤3ns, with no software involvement, avoiding the latency uncertainties caused by software interruptions. When the monitoring submodule triggers a fault flag, the hardware logic synchronously executes two actions within nanoseconds: first, it shuts down the output of the faulty link, achieving electrical isolation of the faulty module through isolation circuits to prevent the fault from spreading to other parts of the system; second, it connects the output of the normal link, allowing the normal link to seamlessly take over external service output. The switching process involves no task pauses, data stream interruptions, or clock interruptions. The external output signal exhibits no phase jumps, amplitude abrupt changes, or data frame drops, achieving true zero-interrupt seamless takeover and completely avoiding execution interruptions and latency jitter in hard real-time tasks. The switching sequence of this step can be referenced. Figure 3 , Figure 3 This is a timing diagram for zero-interruption fault switching in one embodiment. The system clock signal is a homogeneous clock with a period of T, serving as the synchronization reference for all actions. The fault comparison result signal is normally low, but jumps to high when the two paths are inconsistent, with a trigger delay ≤1 clock cycle. The main system output enable signal is normally high, but jumps to low during fault switching, with a switching delay ≤3ns. The mirror system output enable signal is normally low, but jumps to high during fault switching, perfectly synchronized with the shutdown edge of the main system enable signal, with no time difference, ensuring uninterrupted output signal. The external service output signal is continuous and distortion-free, with no phase transitions, data loss, or amplitude changes before and after switching, achieving zero-interruption output. The total delay from fault detection to switching completion is ≤10 system clock cycles, fully meeting the nanosecond-level switching requirements of onboard hard real-time tasks.

[0028] Step 5: After the switchover is completed, the global monitoring and execution module performs a graded recovery operation on the faulty link in the background. During the recovery process, the normal link continues to carry out external services. After the faulty link is restored to normal, the current complete running state data of the normal link is batch mirrored and synchronized to the restored faulty link, so that the running state of the two systems is completely consistent.

[0029] After the switchover is complete, services continue to be carried by the normal link, operating without interruption. This step is executed by the fault recovery and resynchronization submodule. After the switchover, services continue to be carried by the normal link, operating without interruption. This submodule performs tiered recovery operations on the faulty link in the background, without affecting the continuous operation of normal services. The tiered recovery operation first distinguishes the fault type and level. For single-event soft errors, register refresh and configuration frame dynamic refresh operations are performed; for single-event functional interruptions, according to the fault level, partial module reset, full chip reset, program reconfiguration, and power-off restart operations are performed sequentially. After the faulty link is restored to normal, the submodule mirrors and synchronizes the current full-link operating state of the normal link, including register state, data buffer, state machine position, carrier tracking phase, frequency offset compensation parameters, link calibration parameters, etc., to the restored faulty link in batches via a high-speed parallel bus, ensuring that the dual-system states are completely consistent, providing a basis for subsequent re-entry into lockstep operation.

[0030] Step 6: After completing the running state mirror synchronization, the global monitoring and execution module controls the recovered faulty link and normal link to re-enter the clock-level lockstep running state through the same source synchronization trigger signal, restore the dual-path mirror redundancy backup mode, and return to step 2.

[0031] After the state mirroring is completed, the recovered link is started synchronously with the normal link on the same clock edge by triggering the synchronous trigger signal through the same source clock. It then re-enters the clock-level lockstep operation state, restores the dual-path mirroring redundancy backup mode, completes the complete fault handling and system redundancy recovery closed loop, and waits for the next fault handling.

[0032] In the aforementioned zero-interruption recovery method based on signal end-to-end lockstep operation state mirror redundancy, in steps 1 and 2 of this application, two systems covering the entire end-to-end link from baseband to radio frequency are driven into clock-level lockstep parallel operation through a homogeneous high-precision clock synchronization unit. This, combined with the same driving circuit, achieves 100% homogeneity of input excitation, enabling real-time complete mirroring of the dual-path system's operating state, data state, and control state. This establishes the prerequisite that a hot mirror link, completely consistent with the main system's state, can seamlessly take over services at any time. This transforms fault response from the traditional passive mode of interruption followed by recovery to an active takeover by synchronous bypassing to the already synchronized mirror link during a fault. This model eliminates the root cause of task interruption at the architectural level. Based on this, step 3 uses digital link monitoring in the full-link real-time status monitoring submodule to compare the values ​​of dual-channel baseband output data, key registers, state machines, interface communication signals, and configuration registers bit by bit clock cycle, detecting latency ≤10 clock cycles. It also monitors the RF link, forming a self-looping monitoring link through an RF self-calibration coupling circuit. This real-time demodulation and detection of the loopback signal's frequency offset, power, spectrum, and EVM, with a detection latency ≤1μs, and cross-machine collaborative monitoring, using external single-machine loopback verification to compare the integrity and correctness of transmitted data in real time. This three-dimensional collaboration establishes a connection between the digital and RF domains. The comprehensive, blind-spot-free monitoring mechanism can detect inconsistencies between two paths within a single clock cycle and accurately pinpoint the faulty link node and fault type. This avoids the risk of false triggering and secondary interruptions caused by misjudging SEFI as a common soft error. Furthermore, the detection process runs in parallel with the business processing pipeline, introducing no additional latency jitter. In step 4, the global monitoring and execution module employs a high-speed analog switch array and digital bypass control logic implemented entirely in hardware. Upon detecting a fault, it synchronously performs the operation of shutting down the faulty link and connecting the normal link on the same clock edge. The entire switching process is software-free, avoiding latency uncertainties caused by software interruptions. Moreover, the digital bypass... While switching, the control logic disables the output enable signal of the faulty link and switches its output to a high-impedance state through a high-speed analog switch, achieving complete electrical isolation between the faulty link and other parts of the system. This effectively prevents fault signals from backflowing or coupling to the normal link, fundamentally eliminating the risk of fault propagation. During the switching process, the off edge and the on edge completely coincide, and there are no phase jumps, amplitude changes, or data frame drops in the external service signals. This achieves seamless takeover with zero interruptions at the nanosecond level, fully meeting the deterministic latency constraints of nanosecond to microsecond levels for hard real-time tasks. Moreover, the above switching is completed by hardware logic at the clock cycle level without introducing any additional latency jitter, ensuring the time determinism of the hard real-time system.In steps 5 and 6, after the switchover is completed, the normal link continues to carry the service with zero interruption. At the same time, a hierarchical recovery operation is performed on the faulty link in the background. Single-event soft error (SOFE) is performed to refresh the register and dynamically refresh the configuration frame. SEFI sequentially performs partial module reset, full chip reset, program reconfiguration, and power failure restart. The entire recovery process does not affect the continuous operation of the foreground service. After the faulty link is restored to normal, the full link operation state of the normal link is batch mirrored and synchronized to the restored faulty link through a high-speed parallel bus. Then, the dual-path system is re-entered into clock-level lockstep operation through the same-source synchronization trigger signal. Thus, a complete fault handling closed loop and redundant state recovery are completed under the premise of zero service interruption. This enables the system to withstand an unlimited number of continuous single-event soft errors and SEFI attacks. All of the above technical means do not require modification of the internal logic or trigger design of FPGA / SoC devices. Instead, they can be implemented by externally deploying mature commercial devices such as same-source clock units, equal-length line drive distribution networks, high-speed analog switch arrays, and monitoring FPGAs. Therefore, it can be directly and non-intrusively adapted to commercial aerospace COTS devices and is suitable for various emerging hard real-time scenarios such as spaceborne space-to-ground communication, inter-satellite communication, and space-to-ground power supply. Throughout the entire process, the total latency from fault detection to handover completion is ≤10 system clock cycles. The handover process is free of any service interruption, data stream interruption, or phase jump, fully meeting the nanosecond-level handover requirements and deterministic latency constraints of spaceborne hard real-time tasks. This achieves in-situ transparent fault tolerance for single-event soft errors and seamless, non-intrusive recovery from SEFI.

[0033] In one embodiment, the acquisition and comparison of the digital link operating status includes real-time acquisition of baseband processing output data, key register status, state machine operating status, interface communication signals and configuration register values ​​of the dual-path system, and bit-by-bit comparison at each clock edge; the acquisition and monitoring of radio frequency link operating parameters includes coupling the radio frequency transmission signals of the main operating system and the mirror operating system to the corresponding receiving branch through a radio frequency self-calibration coupling circuit to form a self-loop monitoring link, and performing real-time demodulation, frequency offset detection, power detection, spectrum analysis and error vector amplitude calculation on the loop signal. When the radio frequency signal parameter of a certain link exceeds a preset threshold, it is determined that the link has failed.

[0034] Specifically, digital link monitoring and RF link monitoring complement each other, forming a comprehensive, real-time fault detection and location capability across the entire link. Digital link monitoring covers various soft errors in the digital domain, including baseband processing, register status, state machine operation, and interface communication. RF link monitoring, through its RF self-calibration coupling monitoring branch, covers faults in the analog / digital signal processing domain, including RF transmit links, receive links, and modulation / demodulation. Monitoring data from both dimensions are aggregated and analyzed in the global monitoring and execution module. When inconsistencies occur between the two data / status streams or when a certain RF signal parameter exceeds a preset threshold, a fault flag is immediately triggered, and the faulty link node and fault type (single-event soft error or SEFI) are located. The digital link detection delay is ≤10 system clock cycles, and the RF link detection delay is ≤1μs. Through collaborative monitoring of both digital and RF dimensions, single-event soft errors and SEFI faults can be accurately distinguished, avoiding misclassification of SEFI as a regular soft error or vice versa, effectively reducing the false trigger rate and eliminating the risk of fault propagation and secondary interruption. Upon fault detection, the fault flag is immediately transmitted to the nanosecond-level switching control and fault isolation submodule to trigger a switching action.

[0035] In one embodiment, the pure hardware logic circuit includes a high-speed analog switch array and digital bypass control logic. The switching speed of the high-speed analog switch array is ≤3ns. The output terminal of the faulty link shutdown and the output terminal of the normal link are executed synchronously on the same clock edge. The output terminal of the high-speed analog switch array is connected to the output terminal of the main operating system and the output terminal of the mirror operating system, respectively. The digital bypass control logic outputs a gating control signal according to the fault detection result, and switches the high-speed analog switch array from turning on the faulty link to turning on the normal link on the same clock edge.

[0036] Specifically, the pure hardware logic circuitry, without software involvement, avoids the latency uncertainties caused by software interruptions. The switching speed of the high-speed analog switch array is ≤3ns, and combined with the digital link detection latency of ≤10 system clock cycles, the total latency from fault detection to switching completion is ≤10 system clock cycles. The digital bypass control logic receives the fault flag output by the full-link status real-time monitoring submodule and synchronously outputs two selection control signals on the same clock edge: one to turn off the high-speed analog switch of the faulty link, and the other to turn on the high-speed analog switch of the normal link. The two control signals are strictly synchronized, ensuring that the turn-off edge and turn-on edge completely coincide without time difference, guaranteeing that the external service output signal has no gaps, phase jumps, or amplitude abrupt changes during switching. Under normal operating conditions, the control logic turns on the full-link output of the main system and isolates the output of the mirror system; under fault conditions, the control logic synchronously switches, and the normal link seamlessly takes over the external service output. The entire switching process ensures continuous and distortion-free external service signals, achieving true zero-interruption seamless takeover.

[0037] In one embodiment, the faulty link is isolated by switching the output of the faulty link to a high-impedance state and disconnecting its electrical connection with the external load through a high-speed analog switch array, while the output enable signal of the faulty link is disabled through digital bypass control logic, thereby achieving electrical isolation between the faulty link and other parts of the system.

[0038] Specifically, when the monitoring submodule triggers the fault flag, the digital bypass control logic sends an output disable signal to the faulty link simultaneously with the output enable control signal, disabling the output of the faulty link. The high-speed analog switch array switches the output of the faulty link to a high-impedance state, disconnecting its electrical connection to the external load. Through this dual isolation mechanism of output enable / disable and high-impedance switching, complete electrical isolation between the faulty link and other parts of the system is ensured, preventing fault signals from backflowing or coupling to the normal link or external load through the output, fundamentally eliminating the risk of fault propagation. The isolation action and the conduction action of the normal link are executed synchronously on the same clock edge, without introducing additional delay.

[0039] In one embodiment, the graded recovery operation is to perform a differentiated recovery strategy based on the fault type and fault level: for single-event soft errors, register refresh and configuration frame dynamic refresh operations are performed; for single-event functional interruptions, partial module reset, full chip reset, program reconfiguration, and power-off restart operations are performed in sequence according to the fault level until the fault link is restored to normal.

[0040] Specifically, the fault recovery and resynchronization submodule first distinguishes the fault type and level based on the fault location information provided by the full-link status real-time monitoring submodule. If it is determined to be a single-event soft error (SEU / SET), a lightweight recovery operation is performed: first, register refresh is performed to restore critical registers to preset safe values; if register refresh is ineffective, configuration frame dynamic refresh is performed, using SelectMap or a similar interface to perform frame-level refresh of the FPGA configuration memory to repair the configuration bit flip caused by SEU. If it is determined to be SEFI, recovery is performed step by step according to the severity of the fault: the first level performs a partial module reset, resetting only the functional module that failed; if the partial reset is ineffective, a full chip reset is performed; if the full chip reset is ineffective, program reconfiguration is performed, reloading the complete configuration bit stream; if reconfiguration is still ineffective, a power-off restart is performed. After each level of recovery operation is completed, the full-link status real-time monitoring submodule verifies whether the fault has been eliminated. If it has been eliminated, the recovery process stops; if it has not been eliminated, it is upgraded to the next level of recovery operation. The entire hierarchical recovery process is executed in the background, without affecting the continuous carrying of external services by the normal link, and any momentary interruption caused by the recovery operation does not affect the output of external services.

[0041] In one embodiment, the complete operational data includes the register state of the normal link, data buffer content, the current position of the state machine, carrier tracking phase value, frequency offset compensation parameter value, link calibration parameter value, and signal correlator calculated pulses. The complete operational data of the current normal link is batch mirrored and synchronized to the recovered faulty link. This is performed through a high-speed parallel bus between the normal link and the faulty link. After the mirroring and synchronization is completed, the register state, data buffer content, state machine position, carrier tracking phase, frequency offset compensation parameter, and link calibration parameter of the recovered faulty link are completely consistent with those of the normal link.

[0042] Specifically, after the faulty link recovers, to enable the dual-path system to re-enter clock-level lockstep operation, the operating state of the faulty link must be completely consistent with that of the normal link. Operating state data includes not only the register states, data buffer contents, and state machine positions in the digital control domain, but also the carrier tracking phase value (32-bit), frequency offset compensation parameter value (32-bit), link calibration parameter value (32-bit), and correlator calculation pulses (1-bit pulse width per clock cycle) in the digital signal processing domain. This operating state data is transmitted in batches in real time via a high-speed parallel bus between the normal and faulty links. During transmission, the normal link continuously carries services, and the transmission operation does not occupy the service processing bandwidth of the normal link or is only performed during idle intervals. The fault mirror system immediately sets the correct carrier tracking phase value (32-bit), frequency offset compensation parameter value (32-bit), and link calibration parameter value (32-bit) into the mirror system based on the pulses calculated by the correlator from the normal link (main operating system). After the mirror system synchronization is completed, all operating state parameters of the recovered faulty link and the normal link are completely consistent, laying the state foundation for re-entering clock-level lockstep operation.

[0043] In one embodiment, the same-source high-precision clock synchronization unit uses a single high-stability crystal oscillator as the sole clock source. It is divided into two differential clock signals through a first-level low-jitter clock buffer and sent to the clock input terminals of the main operating system and the mirror operating system, respectively. The PCB traces of the two differential clock signals are designed with equal lengths, and the phase difference between the two differential clock signals is less than the setup / hold time window of one clock edge.

[0044] Specifically, a single aerospace-grade high-stability crystal oscillator is used as the sole clock source, avoiding frequency deviation and phase drift between multiple independent crystal oscillators. A low-jitter clock buffer splits the single clock signal into two completely equal-length, equal-phase differential clock signals, which are then fed into the clock inputs of the FPGA / SoC, ADC / DAC, and RF chips in the main operating system and the mirror operating system, respectively. The clock PCB traces employ a strict equal-length design, with trace length deviation controlled within 50ps, ensuring that the phase difference between the two clock signals is less than the setup / hold time window of one clock edge. Simultaneously, a synchronous trigger clock from the same source is output to the global monitoring and execution module as a synchronization reference for dual-system status comparison and fault monitoring, ensuring that all monitoring, comparison, and switching actions are completed on the synchronous clock edge.

[0045] In one embodiment, the service signals from the same source are sent to the corresponding input terminals of the main operating system and the mirror operating system simultaneously through the same driving circuit, so that the input excitation of the dual-path system is completely consistent.

[0046] Specifically, all input signals, including business input data, external received signals, and configuration commands, are simultaneously sent to the corresponding input terminals of the main operating system and the mirror operating system through the same driver circuit. The same driver circuit ensures that the level, timing, and phase of the two output signals are completely consistent, and the equal-length traces ensure that the signal transmission delay is the same, thereby guaranteeing that the input excitation of the two systems is 100% consistent, providing the same input conditions for lockstep operation.

[0047] In one embodiment, both the main operating system and the mirror operating system are processing systems that cover the entire signal link from end to end. Each system includes a baseband data processing module, a digital-to-analog conversion module, an RF transmit link, an RF receive link, an analog-to-digital conversion module, and a baseband demodulation and decoding module. The hardware design, logic code, configuration parameters, and PCB layout of the two systems are completely identical.

[0048] Specifically, the dual-path system employs identical hardware design, logic code, configuration parameters, and PCB layout, covering the end-to-end full signal link for spaceborne communication / data processing tasks, including: baseband data processing module (FPGA / SoC), digital-to-analog converter (DAC), RF transmit link, RF receive link, analog-to-digital converter (ADC), and baseband demodulation and decoding module, with no bottlenecks in the link. The completely symmetrical design of the dual-path system ensures that under the same input stimulus and clock drive, the two systems operate in completely identical states, without any state deviations caused by hardware or logic differences.

[0049] In one embodiment, a cross-machine collaborative monitoring step is also included. The cross-machine collaborative monitoring step is to use the loopback verification mechanism of the external single machine to compare the integrity and correctness of the data sent from the main running system and the mirror running system to the external single machine in real time, so as to realize the fault monitoring of the external communication link.

[0050] For the electrical / optical data communication link with an external unit, a loopback verification mechanism on the external unit is used to compare the integrity and correctness of data sent by the main operating system and the mirror operating system in real time. After receiving data, the external unit returns the data through the loopback channel. The global monitoring and execution module compares the returned data bit by bit with the original sent data. If the comparisons are inconsistent for several consecutive cycles, a fault is determined to have occurred in the external communication link, triggering a fault flag. Cross-machine collaborative monitoring enables fault detection of the external communication link, avoiding missed faults at the end of the link and further expanding the coverage of the entire link monitoring.

[0051] It should be understood that, although Figure 2 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0052] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0053] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A zero-interruption recovery method based on full-link-level lockstep operational state mirror redundancy, applied to a spaceborne hard real-time signal processing system, wherein the spaceborne hard real-time signal processing system comprises a full-link-level lockstep main operational system and a full-link-level lockstep mirror operational system, characterized in that, The method includes: Step 1: After the onboard hard real-time signal processing system is powered on, the same-source high-precision clock synchronization unit starts up and outputs the same-source synchronized system clock and synchronization trigger signal to the main operating system, the mirror operating system and the global monitoring and execution module, and completes the hardware initialization, logic loading and configuration parameter writing of the dual-path system, so that the initial operating state of the main operating system and the mirror operating system are completely consistent. Step 2: The main operating system and the mirror operating system enter a clock-level lockstep parallel operation state under the drive of the same system clock, receive the same source input service signals, perform the same signal processing operation on the same clock edge, and generate mirror-consistent operating status data and output service signals in real time; wherein, in normal operation mode, the output terminal of the main operating system is turned on and outputs service signals to the outside, while the output terminal of the mirror operating system is in a high-impedance isolation state and does not output to the outside. Step 3: Based on the synchronization trigger signal, the global monitoring and execution module collects the digital link operating status and RF link operating parameters of the main operating system and the mirror operating system, as well as the output service signals, clock cycle by clock cycle. At each clock edge, it compares the operating status of the dual-path system bit by bit to determine whether a single-event soft error or single-event function interruption fault has occurred; if there is no fault, it returns to step 2. Step 4: When a fault is detected in the main operating system or the mirror operating system, the global monitoring and execution module uses pure hardware logic circuits to synchronously switch the state of the two output paths on the same clock edge, shutting off the conduction of the output end of the faulty link, isolating the faulty link, and simultaneously connecting the output end of the normal link, so that the normal link can seamlessly take over the output of external service signals, and the external service signals are continuous and uninterrupted during the switching process. Step 5: After the switchover is completed, the global monitoring and execution module performs a graded recovery operation on the faulty link in the background. During the recovery process, the normal link continues to carry out external services. After the faulty link is restored to normal, the current complete running state data of the normal link is batch mirrored and synchronized to the restored faulty link, so that the running state of the two systems is completely consistent. Step 6: After completing the running state mirror synchronization, the global monitoring and execution module controls the recovered faulty link and the normal link to re-enter the clock-level lockstep running state through the same source synchronization trigger signal, restore the dual-path mirror redundancy backup mode, and return to step 2.

2. The method according to claim 1, characterized in that, The acquisition and comparison of the digital link operating status includes real-time acquisition of baseband processing output data, key register status, state machine operating status, interface communication signals and configuration register values ​​of the dual-path system, and bit-by-bit comparison at each clock edge; the acquisition and monitoring of the radio frequency link operating parameters includes coupling the radio frequency transmission signals of the main operating system and the mirror operating system to the corresponding receiving branch through the radio frequency self-calibration coupling circuit to form a self-loop monitoring link, and performing real-time demodulation, frequency offset detection, power detection, spectrum analysis and error vector amplitude calculation on the loop signal. When the radio frequency signal parameter of a certain link exceeds the preset threshold, it is determined that the link has failed.

3. The method according to claim 1, characterized in that, The pure hardware logic circuit includes a high-speed analog switch array and digital bypass control logic. The switching speed of the high-speed analog switch array is ≤3ns. The output terminal for shutting down the faulty link and the output terminal for turning on the normal link are executed synchronously on the same clock edge. The output terminals of the high-speed analog switch array are respectively connected to the output terminals of the main operating system and the mirror operating system. The digital bypass control logic outputs a gating control signal based on the fault detection result, and switches the high-speed analog switch array from conducting a faulty link to conducting a normal link on the same clock edge.

4. The method according to claim 3, characterized in that, The isolation of the faulty link is achieved by switching the output of the faulty link to a high-impedance state and disconnecting its electrical connection with the external load through the high-speed analog switch array, and by disabling the output enable signal of the faulty link through the digital bypass control logic, thereby realizing the electrical isolation of the faulty link from other parts of the system.

5. The method according to claim 1, characterized in that, The graded recovery operation is to execute differentiated recovery strategies according to the fault type and fault level: for single-event soft errors, register refresh and configuration frame dynamic refresh operations are performed; for single-event functional interrupts, partial module reset, full chip reset, program reconfiguration, and power-off restart operations are performed in sequence according to the fault level until the fault link is restored to normal.

6. The method according to claim 1, characterized in that, The complete operational data includes the register state of the normal link, the data buffer content, the current position of the state machine, the carrier tracking phase value, the frequency offset compensation parameter value, the link calibration parameter value, and the signal correlator calculated pulse; The complete operational data of the normal link is batch mirrored and synchronized to the recovered faulty link. This is performed through the high-speed parallel bus between the normal link and the faulty link. After the mirroring and synchronization is completed, the register state, data buffer content, state machine position, carrier tracking phase, frequency offset compensation parameters and link calibration parameters of the recovered faulty link are completely consistent with those of the normal link.

7. The method according to claim 1, characterized in that, The same high-precision clock synchronization unit uses a single high-stability crystal oscillator as the sole clock source. It is divided into two differential clock signals by a first-level low-jitter clock buffer and sent to the clock input terminals of the main operating system and the mirror operating system, respectively. The PCB traces of the two differential clock signals are designed with equal lengths, and the phase difference between the two differential clock signals is less than the setup / hold time window of one clock edge.

8. The method according to claim 1, characterized in that, The service signals from the same source are sent to the corresponding input terminals of the main operating system and the mirror operating system simultaneously through the same driving circuit, so that the input excitation of the dual-path system is completely consistent.

9. The method according to claim 1, characterized in that, Both the main operating system and the mirror operating system are processing systems covering the entire signal link from end to end. Each system includes a baseband data processing module, a digital-to-analog conversion module, an RF transmission link, an RF reception link, an analog-to-digital conversion module, and a baseband demodulation and decoding module. The hardware design, logic code, configuration parameters, and PCB layout of the two systems are completely identical.

10. The method according to claim 1, characterized in that, It also includes a cross-machine collaborative monitoring step, in which the global monitoring and execution module compares in real time the integrity and correctness of the data sent from the main operating system and the mirror operating system to the external single machine through the loopback verification mechanism of the external single machine, so as to realize the fault monitoring of the external communication link.

Citation Information

Patent Citations

  • Real-time quantitative evaluation method and system for cross-domain collaborative efficiency of avionics system

    CN122110766A

  • Radio-Control Board For Software-Defined Radio Platform

    US20110078355A1