An autonomous fault tolerance and failure recovery method for on-orbit processing systems

By implementing TMR and EDAC technologies in FPGA, combined with real-time data comparison and NOR memory update, the problem of satellite FPGA logic state being susceptible to single-particle upsets was solved, autonomous fault tolerance and fault recovery were achieved, and the stable operation of the satellite system was ensured.

CN113918386BActive Publication Date: 2025-09-26BEIJING INST OF TECH LEIKE AEROSPACE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111239294.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-25
Publication Date
2025-09-26
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

The FPGA logic state in the satellite is easily affected by single-particle effects, resulting in reduced availability and reliability. The potential state jump caused by single-particle upsets affects the satellite functional performance and system operation stability.

Method used

A JTAG burner is used to copy three copies of the application file to different address segments of the PROM memory. The FPGA compares the program data in the SRAM area with the PROM data in real time, performs automatic error correction, and updates the application file in the NOR memory through ground injection when necessary. Combined with TMR and EDAC technology, autonomous fault tolerance and fault recovery are achieved.

Benefits of technology

It improves the reliability and stability of satellite on-orbit operation, reduces circuit power consumption, simplifies circuit complexity, and achieves autonomous fault tolerance and fault recovery without the need for external controller intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113918386B_ABST
    Figure CN113918386B_ABST
Patent Text Reader

Abstract

The autonomous fault tolerance and fault recovery method for an on-orbit processing system of the present invention can verify and correct the on-orbit processing system based on the internal software of the FPGA to ensure the normal operation of the satellite. Verification and correction based on the internal software of the FPGA to address the impact of single-event upsets include TMR (triple modular redundancy), EDAC (single-bit error correction), and program reloading. TMR mainly ensures that the system can still be loaded after the application file burned by the satellite is subjected to a single-event upset; EDAC mainly ensures that all key data such as configuration variables can still be corrected after a single-event upset during the application loading and operation of the single machine, ensuring normal operation. The present invention also includes a program reloading process, which mainly ensures that the ground-based on-orbit injection can still be achieved after the three application files burned by the satellite are all subjected to a single-event upset, ensuring the normal operation of the satellite.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reliability of on-orbit fault emergency plans, and in particular to an autonomous fault tolerance and fault recovery method for an on-orbit processing system. Background Art

[0002] With the continuous advancement of technology, the demand for satellite data processing capabilities is increasing. FPGAs are gradually being used in aerospace engineering, becoming core components for satellite data processing and control. However, the unique operating environment of satellites in space is subject to various external factors, including solar electromagnetic radiation, high-energy particles such as neutral particles and plasma at various temperatures, and the radiation effects caused by high-energy charged particles incident on electronic devices are called single-event effects. Depending on the mechanism of the effect, it can range from single-event upsets (SEUs) to single-event lockups.

[0003] Judging from the current application of satellite products, the availability and reliability of FPGA logic states are often limited due to single-particle effects. When a single-particle upset occurs, the potential state jumps, "0" becomes "1", or "1" becomes "0". When a single-particle upset occurs in a device, although data can be resent, the device or functional module can be restarted or powered off, the functional performance of the satellite product at that time will be greatly affected, posing a certain risk to the system's on-orbit operation. Summary of the Invention

[0004] In view of this, the present invention provides an autonomous fault tolerance and fault recovery method for an on-orbit processing system, which can verify and correct the on-orbit processing system based on the internal software of the FPGA to ensure the normal operation of the satellite.

[0005] To achieve the above purpose, the technical solution of the present invention is as follows:

[0006] An autonomous fault tolerance and fault recovery method for an on-orbit processing system comprises the following steps:

[0007] Step 1: Use a JTAG programmer to copy three identical copies of the application file and program them into different address segments of the PROM memory.

[0008] Step 2: After the stand-alone machine is powered on, the FPGA reads the PROM program data into the SRAM area, reads the SRAM area program data in real time and loads it, and the stand-alone machine performs real-time IP core refresh;

[0009] Step 3: When the PROM program data is normal, the FPGA initiates a readback refresh function to the processing board at set intervals. According to the instruction scheduling, it reads the program data in the SRAM area once and compares it with the PROM data. If a single bit of data is abnormal after the readback comparison, it automatically performs error correction processing. The FPGA replaces the abnormal data with the original data to ensure normal execution of the single machine.

[0010] When the PROM program data is abnormal, the application file is uploaded to the NOR memory of the FPGA external device using the ground upload method and loaded; after the FPGA is reloaded, all modules are initialized and work starts again.

[0011] In the step 3, the application file is uploaded to the NOR memory of the FPGA plug-in using the ground upload method and loaded, which includes the following steps:

[0012] Step 21: The ground application file is uploaded to the NOR memory plugged into the FPGA inside the on-orbit stand-alone unit through the LVDS interface. After the ground upload, the NOR memory continues to store the application file until the next ground upload overwrites it.

[0013] Step 22: After the circuit is powered on, the FPGA loads and initializes all modules;

[0014] Step 23: Send a command to load the program using the above program via the ground or satellite autonomous mission command;

[0015] Step 24: FPGA reads the application program in the NOR memory and reloads the application program injected on the ground into the FPGA through the JTAG interface.

[0016] Among them, in the step 2, the specific method of reading the SRAM area program data in real time and loading is:

[0017] After the circuit is powered on, the FPGA reads the PROM configuration and application files according to the reset signal; the FPGA compares the three application files in pairs, and votes on the comparison output results by taking two out of three. The result with the higher similarity is output and loaded for operation.

[0018] Among them, in step 2, the specific method of performing real-time IP core refresh on a single machine is: after the FPGA is loaded and initialized, it starts normal operation and starts the SEU refresh module; the FPGA enters the main process and starts the detection refresh module. The refresh module controls the corresponding UART, configures the baud rate to 115200, the bit width to 8 bits, and the parity bit to 1.

[0019] Among them, in step 3, the specific implementation method of automatic error correction processing is as follows:

[0020] The FPGA configures the EDAC to ACM mode, sets AUTO_CORRECT_MODE to a high level, and sets End_of_scan to an output signal to confirm that the device reads back the CRC and scans the device to find configuration changes caused by SEUs. The FPGA enters automatic detection and error correction mode. The detection module controls the corresponding UART, periodically reads the RAM configuration data, and performs a CRC check. If a CRC anomaly is detected, the FPGA will read back the configuration bit error based on the check anomaly value and correct it.

[0021] Among them, in the step three, after the single machine detects that there is a single bit of data anomaly and corrects the error in real time, it is fed back to the single machine telemetry.

[0022] The NOR memory uses 21 address lines and 16 data lines to store ground application files.

[0023] Beneficial effects:

[0024] The present invention's autonomous fault-tolerance and fault-recovery method for an on-orbit processing system uses internal FPGA software for verification and correction to address the impact of single-event upsets. The method includes TMR (triple modular redundancy), EDAC (single-bit error correction), and program reloading. TMR primarily ensures that even after a single-event upset occurs when burning application files on a single satellite, the system can still be loaded. EDAC primarily ensures that even after a single-event upset occurs during application loading and operation, all configuration variables and other critical data can still be corrected, ensuring normal operation. The present invention also includes a program reloading process, which primarily ensures that even after a single-event upset occurs when burning three copies of application files on a single satellite, ground-based onboard injection can still be performed, ensuring normal satellite operation. The autonomous fault-tolerance, namely the TMR (triple modular redundancy) and EDAC methods described in the present invention, are integrated within the FPGA, eliminating the need for additional circuitry, reducing circuit power consumption, and eliminating the need for an external controller. The system is low in complexity, highly reliable, and requires no ground intervention, achieving autonomous fault-tolerance for the on-orbit processing system. The program reloading process only adds a NOR memory chip, simplifying the circuitry.

[0025] The NOR memory described in this invention uses 21 address lines and 16 data lines to store ground-based application files. The application program of these files does not affect the normal loading of the FPGA. This type of memory has high radiation resistance and is already widely used in spacecraft image processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is the TMR three-out-two flow chart of the present invention.

[0027] Figure 2 This is a flow chart of the autonomous detection and error correction of the on-orbit processing system of the present invention.

[0028] Figure 3 This is the CRC-16 bit encoding principle of the present invention.

[0029] Figure 4 This is a block diagram of the autonomous detection and error correction principles of the on-orbit processing system of the present invention.

[0030] Figure 5 This is a simulation block diagram of the autonomous detection and error correction of the on-orbit processing system of the present invention.

[0031] Figure 6 This is a flow chart of the fault recovery method of the present invention. DETAILED DESCRIPTION

[0032] The present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0033] The present invention provides an autonomous fault tolerance and fault recovery method for an on-orbit processing system. The method performs verification and correction based on FPGA internal software to address the impact of single-event upsets. The method includes TMR (triple modular redundancy), EDAC (single-bit error correction), and program reloading. TMR mainly ensures that after a satellite single-machine burns an application file that is subject to a single-event upset, system loading can still be achieved. EDAC mainly ensures that during the application loading and running of the single-machine, all configuration variables and other key data can still be corrected after a single-event upset occurs, thereby ensuring normal operation. Program reloading mainly ensures that after a satellite single-machine burns three application files that are subject to a single-event upset, ground-based on-orbit injection can still be achieved, thereby ensuring normal operation of the satellite. Figure 2 The flow chart of the present invention comprises the following steps:

[0034] Step 1: Use a JTAG programmer to copy three identical copies of the application file and program them into different address segments of the PROM memory (performing TMR, or triple modular redundancy).

[0035] Step 2: After the stand-alone machine is powered on, the FPGA reads the PROM program data into the SRAM area, reads the SRAM area program data in real time, performs relevant configuration, and performs real-time IP core refresh on the stand-alone machine; the specific method is:

[0036] After the circuit is powered on, the FPGA reads the PROM configuration and application files according to the reset signal; the FPGA compares the three application files in pairs, and votes on the comparison output results out of three, outputting the result with higher similarity and loading it for operation; the TMR three-out-of-two flow chart is shown in the figure. Figure 1 As shown;

[0037] After the FPGA is loaded and initialized, it starts to work normally and starts the SEU refresh module. The FPGA enters the main process and starts the detection refresh module. The refresh module controls the corresponding UART and configures the baud rate to 115200, the bit width to 8 bits, and the parity bit to 1.

[0038] Step 3: When the PROM program data is normal, the SRAM program may experience data anomalies due to factors such as single particles in space, resulting in a single bit flip of inconsistent program data. This will cause the standalone machine to operate abnormally. Autonomous fault tolerance can be used to implement error correction to ensure normal execution of the standalone machine. Although PROM program data is less affected by external factors, there is still a small risk of data anomalies, resulting in a bit flip in the soldering program data. When the PROM program data is abnormal, autonomous fault tolerance cannot be used to implement error correction, and a program ground update and upload is required.

[0039] The specific process of automatic error correction is as follows: FPGA reads the program data in the SRAM area according to instruction scheduling, compares it with the PROM data, and performs single-machine detection. If a single-bit data anomaly is found after the readback comparison, the FPGA replaces the anomaly with the original data to ensure normal execution of the single machine. In addition, after the single machine detects a single-bit data anomaly and corrects the error in real time, the error is fed back to the single machine telemetry to count the number of flips. The specific implementation of automatic error correction is as follows:

[0040] FPGA configures EDAC to ACM (automatic correction mode), sets AUTO_CORRECT_MODE to high, and sets End_of_scan to output signal. Verify that the device reads back CRC and scans the device to find configuration changes caused by SEU. FPGA enters automatic detection and correction mode. The detection module controls the corresponding UART, periodically reads RAM configuration data, and performs CRC check. For details, see Figure 3 If a CRC anomaly is detected, the FPGA will read back the configuration bit error based on the checksum anomaly value and correct it.

[0041] In the present invention, the problem of bit flipping of the welding program data is that the system cannot realize error correction through autonomous fault tolerance. At this time, the program is updated and loaded on the ground. The process is as follows: Figure 5 As shown, it includes the following steps:

[0042] Step 21: Upload the application file to the NOR memory of the FPGA external device in a ground-based upload manner;

[0043] After ground filling, the NOR memory continues to keep the application file until the next ground filling overwrites;

[0044] Step 22: After the circuit is powered on, the FPGA loads and initializes all modules and then sends a command to load the program using the above program through the ground or satellite autonomous mission command;

[0045] Step 24: FPGA reads the application program of the NOR memory application file and reloads the application program injected on the ground into the FPGA through the JTAG interface;

[0046] Step 25: After the FPGA is reloaded, all modules are initialized and work is resumed.

[0047] Adopting this approach significantly reduces the risk of single-machine system failures and ensures normal operation. Furthermore, the autonomous fault-tolerance components are integrated within the FPGA, eliminating the need for an external controller and reducing circuit power consumption. Ground-based reloading is achieved by adding a NOR memory. Furthermore, the NOR memory described in this invention uses 21 address lines and 16 data lines to store ground-based application files. The application of these files does not affect the normal loading of the FPGA. This type of memory has high radiation resistance and has been widely used in spacecraft image processing.

[0048] In order to illustrate the effectiveness of the present invention, the following experimental demonstration is carried out. The experimental data uses a single bit of data in the configuration space to randomly change. The experimental results are as follows: Figure 6 As shown in the figure, FPGA simulation results demonstrate that the module has detected and corrected the error. This experiment, using the method proposed in this invention, focused on analyzing the effectiveness of an autonomous fault-tolerance method for an on-orbit processing system. The experimental results demonstrate that the invention can accurately correct single-bit errors, thereby correcting subsequent errors. Furthermore, this on-orbit fault recovery method has been used multiple times in stand-alone on-orbit systems, ensuring their normal operation.

[0049] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for autonomous fault tolerance and failure recovery of an on-orbit processing system, characterized in that: The steps include: Step 1: Use a JTAG programmer to copy three identical copies of the application file and program them into different address segments of the PROM memory. Step 2: After the stand-alone machine is powered on, the FPGA reads the PROM program data into the SRAM area, reads the SRAM area program data in real time and loads it, and the stand-alone machine performs real-time IP core refresh; Step 3: When the PROM program data is normal, the FPGA initiates a readback refresh function to the processing board at set intervals. According to the instruction scheduling, it reads the program data in the SRAM area once and compares it with the PROM data. If a single bit of data is abnormal after the readback comparison, it automatically performs error correction processing. The FPGA replaces the abnormal data with the original data to ensure normal execution of the single machine. When the PROM program data is abnormal, the application file is uploaded to the NOR memory of the FPGA using the ground upload method and loaded; after the FPGA is reloaded, all modules are initialized and work is resumed; In step 3, the application file is uploaded to the NOR memory of the FPGA plug-in using the ground upload method and loaded, which includes the following steps: Step 21: The ground application file is uploaded to the NOR memory plugged into the FPGA inside the on-orbit stand-alone unit through the LVDS interface. After the ground upload, the NOR memory continues to store the application file until the next ground upload overwrites it. Step 22: After the circuit is powered on, the FPGA loads and initializes all modules; Step 23: Send a command to load the program using the above program via the ground or satellite autonomous mission command; Step 24: FPGA reads the application program in the NOR memory and reloads the application program injected on the ground into the FPGA through the JTAG interface; In the step 2, the specific method of reading the SRAM area program data in real time and loading is: After the circuit is powered on, the FPGA reads the PROM configuration and application files according to the reset signal; the FPGA compares the three application files in pairs, and votes on the comparison output results by taking two out of three. The result with the higher similarity is output and loaded for operation.

2. The method according to claim 1, wherein In step 2, the specific method of performing real-time IP core refresh on a single machine is as follows: after the FPGA is loaded and initialized, it starts normal operation and starts the SEU refresh module; the FPGA enters the main process and starts the detection refresh module. The refresh module controls the corresponding UART, configures the baud rate to 115200, the bit width to 8 bits, and the parity bit to 1.

3. The method according to claim 2, wherein In step 3, the automatic error correction process is specifically implemented as follows: The FPGA configures the EDAC to ACM mode, sets AUTO_CORRECT_MODE to a high level, and sets End_of_scan to an output signal to confirm that the device reads back the CRC and scans the device to find configuration changes caused by SEUs. The FPGA enters automatic detection and error correction mode. The detection module controls the corresponding UART, periodically reads the RAM configuration data, and performs a CRC check. If a CRC anomaly is detected, the FPGA will read back the configuration bit error based on the check anomaly value and correct it.

4. The method according to claim 2 or 3, wherein: In step 3, the single machine detects a single bit of data anomaly and corrects the error in real time, and then feeds it back to the single machine telemetry.

5. The method according to claim 2 or 3, wherein: The NOR memory uses 21 address lines and 16 data lines to store ground application files.

Citation Information

Patent Citations

  • Satellite-borne embedded software fault-tolerant starting system and method

    CN108446189A

  • Program on-orbit loading refreshing method based on triple modular redundancy

    CN111176908A