A fault-tolerant loading system and method based on experience addressing evolution

CN122593870APending Publication Date: 2026-08-18SHANGHAI GESI AEROSPACE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610800621.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]如:互为冷备份的Norflash不做三模冗余的策略,由于NorFlash 互为冷备份、未采用三模冗余,无法对存储数据实时表决纠错,单粒子诱发的存储位错无法在线识别与修复,隐性数据故障易导致 FPGA 配置加载异常;Flash存在共源失效风险,受电磁脉冲、簇射粒子、电源故障影响可同步损坏,整机无第三路存储兜底,丧失配置恢复路径;故障判定依赖单一数据源,易出现误切换或切换失效,切换过程中断基带业务,无法实现故障无缝自愈;固件升级易造成主备版本不一致,无冗余参考基准无法甄别有效固件,切换后易出现软硬件适配故障

Benefits of technology

[0038] In this invention, monitoring the FPGA's access to NorFlash is limited to sending read commands and the starting address, fundamentally eliminating the risk of NorFlash being accidentally erased due to timing disorder of the error correction algorithm, and changing the system's on-orbit survival logic from "error correction" to "area replacement".

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122593870A_ABST
    Figure CN122593870A_ABST
Patent Text Reader

Abstract

The application provides a fault-tolerant loading system and method based on experience addressing evolution, comprising a baseband FPGA, a monitoring FPGA, a plurality of independent redundant memories and a non-volatile memory unit connected in communication; wherein the baseband FPGA and the monitoring FPGA are interconnected in communication; the redundant memories are internally divided into a plurality of independent storage partitions, forming a multi-dimensional discrete loading space; the monitoring FPGA is internally provided with a dynamic addressing pointer for positioning a current loading combination in the discrete loading space; and the non-volatile memory unit solidifies pointer coordinates of previous successful loading. The system and method realize "natural elimination" of fault partitions and "extremely fast convergence" of loading paths.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of high-reliability design technology for satellite-borne electronic equipment. Specifically, it relates to a fault-tolerant loading system and method based on empirical addressing evolution. More specifically, it relates to a fault-tolerant loading method and system for SRAM-type FPGAs based on radiation-resistant non-volatile MRAM and dynamic pointer progression in a space radiation environment. Background Technology

[0002] In core components such as satellite telemetry and control systems, SRAM-based FPGAs must load bitstreams from external NorFlash upon power-up. Existing technologies typically employ strategies such as "multiple NorFlash chips acting as cold backups without triple-modulus redundancy" or "multiple NorFlash chips storing only a single copy of the loaded information for triple-modulus redundancy and error correction." However, these existing technologies have the following drawbacks:

[0003] For example, Norflash chips that serve as cold backups do not employ a triple-redundancy strategy. Because of this, real-time voting and error correction of stored data is impossible, and single-event-induced storage faults cannot be identified and repaired online. Latent data faults can easily lead to abnormal FPGA configuration loading. Flash memory is susceptible to common-source failures, and can be simultaneously damaged by electromagnetic pulses, shower particles, and power supply failures. The entire system lacks a third-path storage backup, resulting in a lost configuration recovery path. Fault determination relies on a single data source, making accidental switching or switching failures likely. Baseband services are interrupted during switching, preventing seamless self-healing. Firmware upgrades can easily cause inconsistencies between primary and backup versions; without redundant reference benchmarks, valid firmware cannot be identified, and hardware / software compatibility issues are likely after switching. If enhanced reliability is required, multiple Norflash chips need to be stacked, resulting in a significant increase in board space.

[0004] For example, the strategy of storing only a single load information for multiple NorFlash chips to perform triple redundancy and error correction has the following problems: (1) Error correction requires deep timing intervention of NorFlash. Under the bombardment of high-energy particles in space, if the timing logic is subject to single-event flip, it is very easy to cause the timing disorder of the error correction algorithm, which in turn leads to the on-orbit erroneous erasure and write of NorFlash. (2) Each NorFlash chip stores only one load information. If the critical address of a Flash chip is permanently damaged (SEU that does not recover after power failure), triple redundancy will directly degrade to binary redundancy or even fail. (3) Like multiple cold backups, this method also lacks on-orbit state memory. Every time the power is lost and restarted, voting must start from the default address, which wastes the measurement and control arc time. Summary of the Invention

[0005] This invention aims to overcome the aforementioned defects, avoid on-orbit erroneous erasure and write, improve fault tolerance, and overcome the shortcomings of existing fault-tolerant loading methods, such as lack of state memory and the need for mechanical repetition of invalid attempts with each power-on. It provides a fault-tolerant loading method and system for SRAM-type FPGAs based on empirical addressing evolution, which can effectively improve the traditional triple modular redundancy (TMR) fault-tolerant design, enhance the ability of aerospace / military equipment to resist single-event upsets, crashes, and configuration loss, and provide a more streamlined board area for loading SRAM-type FPGAs, achieving "natural elimination" of fault partitions and "rapid convergence" of loading paths.

[0006] The present invention provides a fault-tolerant loading system based on empirical addressing evolution, characterized in that it includes a baseband FPGA with communication connection, a monitoring FPGA, several independent redundant memories, and non-volatile memory units;

[0007] Among them, the baseband FPGA and the monitoring FPGA are interconnected and communicate with each other;

[0008] The redundant memory is internally divided into several independent storage partitions, forming a multi-dimensional discrete loading space;

[0009] The monitoring FPGA has a dynamic addressing pointer inside, which locates the current load combination in the discrete load space;

[0010] Non-volatile memory units that store the pointer coordinates of each successful load.

[0011] Furthermore, the fault-tolerant loading system based on empirical addressing evolution provided by the present invention is further characterized in that:

[0012] The FPGA is monitored to jump between pointer coordinates using a step-increment algorithm.

[0013] Furthermore, the fault-tolerant loading system based on empirical addressing evolution provided by the present invention is further characterized in that:

[0014] The monitoring FPGA authorizes the baseband FPGA to read data from the corresponding redundant memory according to the dynamic addressing pointer;

[0015] The FPGA monitors the Done status signal for reading data from the baseband FPGA.

[0016] Furthermore, the fault-tolerant loading system based on empirical addressing evolution provided by the present invention is further characterized in that:

[0017] Only two control lines are provided between the monitoring FPGA and the baseband FPGA for monitoring purposes: resetting the baseband and providing feedback on the start / stop results of the baseband.

[0018] Furthermore, the fault-tolerant loading system based on empirical addressing evolution provided by the present invention is further characterized in that:

[0019] The monitoring FPGA is of Flash type;

[0020] The baseband FPGA is of SRAM type.

[0021] In addition, the present invention also provides a method for adapting the above-mentioned system, namely, a fault-tolerant loading method based on empirical addressing evolution, characterized in that: a number of pointer coordinates are formed based on redundant memories that are set independently of each other. When the baseband FPGA cannot successfully load data based on the current pointer coordinate, the monitoring FPGA does not diagnose, locate, or repair bad areas, but directly jumps to the next pointer coordinate in an incremental manner until the data of the baseband FPGA is successfully loaded.

[0022] Furthermore, the fault-tolerant loading method based on empirical addressing evolution provided by the present invention is further characterized in that:

[0023] The fault-tolerant loading method is run whenever the baseband FPGA is reloaded.

[0024] Furthermore, the fault-tolerant loading method based on empirical addressing evolution provided by the present invention is characterized by comprising the following steps:

[0025] S1. When the system reloads the baseband FPGA, the monitoring FPGA requests the historical success pointer coordinates stored in the non-volatile memory cell.

[0026] When a historical success pointer coordinate exists, assign that historical success pointer coordinate to the dynamic addressing pointer;

[0027] When "the coordinates of the historical success pointer do not exist", assign the default initial coordinates to the dynamically addressed pointer;

[0028] S2. Monitor the FPGA licensed baseband FPGA, read data from the corresponding redundant memory storage partition based on the current dynamic addressing pointer, and load it.

[0029] S3. Detect the Done status signal of the baseband FPGA within the preset window.

[0030] When a "Done valid signal is detected", the loading is considered successful, and the process proceeds to S4.

[0031] When "Done signal invalid detected", the monitoring FPGA executes a preset step increment algorithm on the dynamic addressing pointer, forcibly jumps to the next combination of pointer coordinates, and returns to S2;

[0032] S4. Monitor the FPGA to overwrite the coordinates of the currently successful dynamic addressing pointer into the non-volatile memory cell.

[0033] Furthermore, the fault-tolerant loading method based on empirical addressing evolution provided by the present invention is further characterized in that:

[0034] In S2, the monitoring FPGA monitors the Done status signal of the baseband FPGA, without intercepting or analyzing the data flow during the loading process.

[0035] Furthermore, the fault-tolerant loading method based on empirical addressing evolution provided by the present invention is further characterized in that:

[0036] In S3, the FPGA is monitored without performing any data repair or fault tracing operations.

[0037] The function and effects of this invention:

[0038] In this invention, monitoring the FPGA's access to NorFlash is limited to sending read commands and the starting address, fundamentally eliminating the risk of NorFlash being accidentally erased due to timing disorder of the error correction algorithm, and changing the system's on-orbit survival logic from "error correction" to "area replacement".

[0039] Traditional triple redundancy is limited by the fact that "each Flash chip stores only a single piece of information," resulting in extremely low fault tolerance (tolerating a maximum of one chip failure). This invention, however, divides each Flash chip into multiple partitions (e.g., four partitions), creating 64 possible loading combinations from three Flash chips. Even if a large area of ​​a Flash chip fails, as long as the remaining partitions can be pieced together to form the correct bit stream, the system can select the correct combination by stepping through pointers, significantly increasing fault tolerance.

[0040] The "pointer stepping mechanism" of this invention achieves natural avoidance of damaged partitions—no physical marking of bad blocks is required, no time-consuming repair is needed, and if the current combination fails to load, the failed combination is directly discarded, the pointer step is incremented, and the next combination to load is selected. It solves the problem of accumulated damage from spatial radiation with the simplest logic.

[0041] In traditional solutions, if a critical address in a Flash chip suffers permanent damage (a non-recoverable SEU after a power outage), the error bit remains permanently, requiring time to perform invalid voting and repair from the default address each time power is restored. This invention employs radiation-resistant non-volatile memory cells, whose write cycles far exceed the maximum reload cycles within the satellite's design life, meeting the long-lifespan requirements of space applications. It can store historical success pointer coordinates. In the event of a sudden power outage and restart, the system can skip all historical failure areas and quickly complete the loading, allowing more time for satellite tracking and control. Attached Figure Description

[0042] Figure 1 This is a hardware topology diagram of the fault-tolerant loading system in this embodiment.

[0043] Figure 2This is a schematic diagram of the redundant memory matrix space in this embodiment.

[0044] Figure 3 This is a flowchart of the loading control process based on empirical addressing evolution in this embodiment. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] This embodiment proposes an improved triple redundancy scheme that uses a dual-chip architecture of a monitoring FPGA (Flash type) and a baseband FPGA (SRAM type) with an external non-volatile memory unit, thereby achieving the effects of "natural elimination" of fault partitions and "rapid convergence" of loading paths.

[0047] Specifically, this embodiment provides a fault-tolerant loading system based on empirical addressing evolution, including: a baseband FPGA, a monitoring FPGA, several (at least three) independent redundant memories (such as NorFlash) and a radiation-resistant non-volatile memory unit;

[0048] Generally, the baseband FPGA and the monitoring FPGA are interconnected and communicate with each other;

[0049] Each redundant memory is connected to the monitoring FPGA in a multi-channel discrete SPI star topology;

[0050] The non-volatile memory unit is directly connected to the monitoring FPGA via a parallel asynchronous bus;

[0051] Each redundant memory chip is internally divided into at least two independent storage partitions, forming a multi-dimensional discrete loading space;

[0052] The monitoring FPGA has an internal "dynamic addressing pointer" used to locate the current load combination in a discrete load space;

[0053] Non-volatile memory units are used to store the pointer coordinates of each successful load.

[0054] like Figure 3 As shown, the method for implementing fault-tolerant loading using the above system includes the following steps:

[0055] S1. Experience Wake-up and Priority Loading: When the system powers on and initializes, the monitoring FPGA first reads the "historical success pointer coordinates" stored in the non-volatile memory cell. If historical coordinates exist, the dynamic addressing pointer is directly assigned to these coordinates, skipping the default initial address. If it is the first power-on and there are no historical records, the default initial coordinates are assigned.

[0056] In practice, S1 is performed whenever the baseband FPGA needs to be reloaded, such as: reloading of a single telemetry and control unit, powering on or off of the entire satellite, soft reset, hard reset, watchdog overflow reset, sending FPGA reload commands from the ground, automatic retrying upon loading failure, and various other baseband FPGA reload scenarios.

[0057] S2. Transparent Loading and Simplified Decision-Making: The monitoring FPGA, based on the current dynamic addressing pointer, reads and loads data from the corresponding multi-redundant memory partitions of the baseband FPGA; the monitoring FPGA only monitors the Done status signal of the baseband FPGA and does not intercept or analyze the data flow during the loading process.

[0058] Regarding the monitoring FPGA, it only monitors the Done status signal of the baseband FPGA, without intercepting or analyzing the data flow during the loading process. Specifically, after the dynamic pointer selects a specific partition of a Flash chip, the monitoring FPGA's internal SPI switch directly connects the Flash's SPI pin to the baseband FPGA, bypassing the monitoring FPGA's internal registers. The monitoring FPGA only retains two control lines: Prog (reset baseband) and Done (baseband feedback start / stop result), and the data flow does not pass through the monitoring FPGA's sampling, buffering, or parsing.

[0059] The advantages of this setup are: it significantly simplifies the monitoring FPGA logic and reduces the probability of irradiation errors. If the monitoring system were to capture SPI data streams, verify firmware, and parse data, it would require on-chip storage and computing resources. In such cases, a single-event fault (SIF) in space would more easily cause errors in the monitoring logic and bus jamming. By only receiving the Done level, the monitoring logic only involves level detection. The pin for detecting Done consumes almost no resources, resulting in extremely simple logic and the lowest failure rate. Furthermore, since the monitoring system does not read the data stream, there is no need for CRC checks, fault location, etc. If a fault occurs, it can be directly retried at a different address, saving diagnostic costs. Moreover, assuming that Flash partition data corruption only invalidates the baseband Done level, the fault is confined to the baseband and memory links. The monitoring system is not affected by abnormal data streams and can operate stably, continuously completing pointer jump retry.

[0060] Step S3, pointer stepping and natural avoidance: If the Done signal is detected as valid within the preset window, the loading is determined to be successful and proceed to step S4; if the Done signal is detected as invalid, it is determined that there is an unacceptable error in the partition combination pointed to by the current pointer. The monitoring FPGA does not perform any data repair or fault tracing operations, but only executes the preset step increment algorithm on the dynamic addressing pointer, forcibly jumps to the next combination, and repeats step S2.

[0061] No data repair or fault tracing operations are performed on the monitoring FPGA. Specifically, locating errors and attempting data repair can take milliseconds or even hundreds of milliseconds, while incrementing the pointer immediately switches to a memory combination and initiates loading again. The satellite powers on quickly and does not get stuck in the fault partition for a long time while waiting for repair. By not reading or parsing the erroneous data stream, abnormal Flash bad data will not enter the monitoring internal logic, avoiding interference from erroneous data with the pointer register and MRAM storage. In this way, baseband / Flash faults are isolated by the fault black box, and the fault will not spread in a chain reaction.

[0062] Step S4, Experience Consolidation and Path Convergence: After successful loading, the monitoring FPGA overwrites the current successful "dynamic addressing pointer coordinates" into the non-volatile memory cell; as the satellite's on-orbit lifespan progresses, the system automatically achieves adaptive convergence of the loading path to the healthy partition through multiple cycles of "failure step-success consolidation".

[0063] Application test example:

[0064] like Figure 1 and Figure 2 As shown, this embodiment uses three NorFlash chips (denoted as NorFlash_1, NorFlash_2, and NorFlash_3) as the loading source for the baseband FPGA. Each Flash chip is divided into four partitions (Part 0~3). The combination of the three Flash chips constitutes 4×4×4=64 loading combinations (this combination changes when the number of partitions and the number of NorFlash chips change, and can be adjusted according to actual needs in practical applications). Radiation-resistant MRAM is connected to the monitoring FPGA as a non-volatile memory unit.

[0065] The core of this invention lies in monitoring how the FPGA manages the "evolution" process of these 64 combinations, as detailed below:

[0066] Phase 1: When the satellite first enters orbit, the MRAM is empty (or all 0s). After the monitoring FPGA is powered on, it reads that the MRAM is empty and initializes the dynamic addressing pointer to the default [combination #0: A0+B0+C0].

[0067] The monitoring FPGA pulls the Prog signal low, releasing the SPI bus to the baseband FPGA. The baseband FPGA reads data from A0, B0, and C0. Assuming everything is normal at this point, the baseband FPGA loads successfully and pulls the Done signal high. After detecting Done, the monitoring FPGA immediately writes the coordinate data "combination #0" into the MRAM. Thus, the system's "initial memory" is formed.

[0068] Phase Two: Suppose that the [Combination #0: A0+B0+C0] region encounters a single-event flip (SEU).

[0069] During the reload of the single-unit measurement and control system, the monitoring FPGA reads the MRAM and still obtains the memory "combination #0", prompting the baseband FPGA to load. However, this time, the baseband FPGA reads a garbled bit stream, internal verification fails, and the Done signal is not pulled high within the specified time, allowing the monitoring FPGA to take over the bus. The monitoring FPGA internally performs a simple pointer increment operation, forcibly jumping the address pointer to [combination #1: A0+B0+C1], and then pulls the Prog signal low again to load the baseband FPGA.

[0070] If combination #1 still fails to load, the monitoring FPGA continues to increment by 1, jumping to combination #2, #3... until it jumps to [combination #4: A0+B0+C4]. The baseband FPGA loads successfully and pulls "Done" high. The monitoring FPGA immediately writes "combination #4" into MRAM, overwriting the previous "combination #0".

[0071] Phase Three:

[0072] The adaptive convergence of Phase Two is repeated throughout the entire on-orbit cycle. The system never actively detects which partition is damaged. It relies solely on a primitive and error-prone mechanism: "If it fails, move the partition one combination forward; if it succeeds, remember the current position." This mechanism, with a total of 64 possible partition combinations (as shown in the diagram), effectively keeps damaged partitions permanently behind the pointer. The longer the on-orbit time, the closer the system pointer is to the end of the partition combination list, increasing the determinism of loading. This transforms the complex problem of "fault diagnosis" into a pure "state machine step-by-step memory" problem.

Claims

1. A fault-tolerant loading system based on experience addressing evolution, characterized by: It includes a baseband FPGA for communication connections, a monitoring FPGA, several independent redundant memories, and non-volatile memory units; The baseband FPGA and the monitoring FPGA are interconnected and communicate with each other. The redundant memory is internally divided into several independent storage partitions, forming a multi-dimensional discrete loading space; The monitoring FPGA is equipped with a dynamic addressing pointer to locate the current loading combination in a discrete loading space; The non-volatile memory unit stores the pointer coordinates of each successful load.

2. The fault-tolerant loading system based on empirical addressing evolution as described in claim 1, characterized in that: The monitoring FPGA uses a step-increment algorithm to jump between pointer coordinates.

3. The fault-tolerant loading system based on empirical addressing evolution as described in claim 1, characterized in that: Only two control lines for monitoring purposes are provided between the monitoring FPGA and the baseband FPGA: one for resetting the baseband and the other for baseband feedback start / stop results.

4. The fault-tolerant loading system based on empirical addressing evolution as described in claim 1, characterized in that: The monitoring FPGA authorizes the baseband FPGA to read data from the corresponding redundant memory according to the dynamic addressing pointer; The monitoring FPGA monitors the Done status signal of the baseband FPGA data reading.

5. A fault-tolerant loading system based on empirical addressing evolution as described in claim 1, characterized in that: The monitoring FPGA is of Flash type; The baseband FPGA is of SRAM type.

6. A fault-tolerant loading method based on empirical addressing evolution, characterized in that: Several pointer coordinates are formed based on redundant memory that is set up independently. When the baseband FPGA fails to load data based on the current pointer coordinate, the monitoring FPGA does not diagnose, locate, or repair bad areas, but directly jumps to the next pointer coordinate in an incremental manner until the data of the baseband FPGA is successfully loaded.

7. The fault-tolerant loading method based on empirical addressing evolution as described in claim 6, characterized in that: The fault-tolerant loading method is run whenever the baseband FPGA is reloaded.

8. The fault-tolerant loading method based on empirical addressing evolution as described in claim 6, characterized in that, Includes the following steps: S1. When the system reloads the baseband FPGA, the monitoring FPGA requests the historical success pointer coordinates stored in the non-volatile memory cell. When a historical success pointer coordinate exists, assign that historical success pointer coordinate to the dynamic addressing pointer; When "the coordinates of the historical success pointer do not exist", assign the default initial coordinates to the dynamically addressed pointer; S2. Monitor the FPGA licensed baseband FPGA, read data from the corresponding redundant memory storage partition based on the current dynamic addressing pointer, and load it. S3. Detect the Done status signal of the baseband FPGA within the preset window. When "Done valid signal detected", the loading is considered successful, and the process proceeds to S4; When "Done signal invalid detected", the monitoring FPGA executes a preset step increment algorithm on the dynamic addressing pointer, forcibly jumps to the next combination of pointer coordinates, and returns to S2; S4. Monitor the FPGA to overwrite the coordinates of the currently successful dynamic addressing pointer into the non-volatile memory cell.

9. The fault-tolerant loading method based on empirical addressing evolution as described in claim 8, characterized in that, In S2, the monitoring FPGA monitors the Done status signal of the baseband FPGA, without intercepting or analyzing the data flow during the loading process.

10. The fault-tolerant loading method based on empirical addressing evolution as described in claim 8, characterized in that, In S3, the FPGA is monitored without performing any data repair or fault tracing operations.