ZYNQ series FPGA self-repairing structure based on self-adaptive redundancy strategy

By designing a self-repair structure based on adaptive redundancy strategies in ZYNQ series FPGAs, combining the processor-side controller and readback refresh controller, the problem of insufficient fault tolerance in the prior art is solved, and efficient fault repair and resource consumption balance in different radiation environments are achieved.

CN120020731APending Publication Date: 2025-05-20NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311542293.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

While improving the fault tolerance capability of FPGAs, the prior art has problems such as high power consumption, slow error detection speed and poor flexibility, making it difficult to achieve balance of resource consumption while ensuring performance.

Method used

A self-healing structure based on adaptive redundancy strategy was designed. By dividing multiple logical functional areas in the ZYNQ series FPGA, each area contains three functional cells with the same function and an interconnected cell. Different redundancy modes (single-cell redundancy, two-cell cold backup redundancy, and three-cell redundancy) are adopted to adapt to different radiation environments, combining the self-healing controller on the processor side and the readback refresh controller to achieve rapid detection and repair of faults.

Benefits of technology

Through adaptive redundancy strategies and dynamic reconfigurable functions, the FPGA's self-repair capability in a radiated environment is improved, power consumption is reduced, error detection speed and flexibility are improved, while maintaining the high performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020731A_ABST
    Figure CN120020731A_ABST
Patent Text Reader

Abstract

The invention discloses a ZYNQ series FPGA self-repairing structure based on a self-adaptive redundancy strategy, which comprises a logic end and a processor end, the logic end and the processor end perform signal interaction through an AXI bus, and a PCAP interface is adopted to complete the configuration of the logic end. The logic end is divided into a plurality of logic function areas, and each area comprises three identical function cells and an interconnection cell. And the processor end comprises a self-repairing controller, a fault responder, a read-back refreshing controller and an external memory DDR (Double Data Rate). The invention provides a ZYNQ self-repairing design method based on the structure, functional cells can change redundancy modes according to different radiation environments, detection and repairing of cell-level and molecular-level faults are achieved, reliability and flexibility of operation of a ZYNQ series FPGA system are improved, and therefore the universal ZYNQ self-repairing design method is provided for design and development personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention utilizes the dynamic reconfigurability of the ZYNQ series FPGAs to design a self-repair structure based on an adaptive redundancy strategy and a corresponding self-repair method, belonging to the field of design of reconfigurable hardware fault tolerance methods. Background Art

[0002] Commercial off-the-shelf SRAM-based FPGAs are increasingly used in the aerospace field. Among them, the ZYNQ series FPGAs of XILINX are a fully programmable system on a chip, which essentially integrates dual ARM Cortex-A9 hard-core processors and SRAM-based FPGA chips, and can be flexibly customized by users for different functions. Due to its high performance, flexible structure, short design cycle and other characteristics, it has received wide attention in recent years. However, compared with aerospace-grade FPGAs, since it does not adopt a high-radiation-resistant design at the process level, it is prone to transient and permanent faults when operating in the space environment. A transient fault refers to a bit flip in the configuration information of the logic circuit in the device, resulting in abnormal circuit function. This type of fault can be repaired without damaging the device. A permanent fault refers to an irreparable fault in the device. Therefore, it is very crucial to study fault tolerance methods for FPGAs.

[0003] At present, experts at home and abroad have proposed a series of fault tolerance methods for SRAM-based FPGAs, which can be essentially divided into two directions. One is to study the manufacturing process of FPGAs, and the other direction is to design error detection and correction methods for FPGAs at different abstraction levels, including hardware redundancy, configuration information read-back and refresh, error detection and correction codes, etc.

[0004] Manufacturing process hardening refers to using SRAM memory cells manufactured by radiation-resistant processes to essentially reduce the impact of radiation on FPGAs, but this method is costly.

[0005] Hardware redundancy backs up redundant circuits with the same function. Dual-mode redundancy can only detect faults, and triple-mode redundancy can locate the faulty module and at the same time ensure the output of the correct result. Combining with the dynamic reconfigurability function of FPGAs, the reconstruction and repair of the faulty module can be completed, but fixed hardware redundancy increases resource waste and system operating power consumption.

[0006] Configuration information read-back and refresh detects faults by reading back the device configuration information and comparing it with the original configuration file. When a fault is detected, the correct configuration information is written back to the frame address of the configuration memory where the fault occurred to complete the refresh repair. However, the cycle for a single read-back detection of the complete configuration information is too long.

[0007] Error detection and correction codes mainly add check bits to the stored information and can locate or correct data errors with a limited number of bits. One application is the design of error detection and correction codes for configuration information, and the other is for the problem that the internal data of BRAM changes dynamically and it is impossible to use read-back comparison for fault tolerance design. However, the number of data bits that can be repaired by this method is limited.

[0008] In summary, while the existing fault tolerance technologies can improve the reliability of FPGAs, they also have problems such as high power consumption, slow error detection speed, and poor flexibility. Therefore, how to maintain the balance between performance and resource consumption while performing fault tolerance on FPGAs is a new challenge for the design of FPGA fault tolerance methods. Summary of the Invention

[0009] Object of the Invention: In order to improve the self-repair ability of ZYNQ series FPGAs after failures caused by radiation, after comprehensively considering system performance and power consumption, the present invention proposes a ZYNQ self-repair structure based on an adaptive redundancy strategy and a corresponding self-repair strategy, which improves the reliability during system operation while ensuring the high performance of the ZYNQ system.

[0010] Technical Solution:

[0011] The ZYNQ on-chip includes a processor side and a logic side. The two parts interact with signals through the AXI bus.

[0012] The logic side is divided into multiple logic function regions, and each region contains three identical functional cells and an interconnection cell.

[0013] The functional cell, that is, the reconfigurable module, decomposes the functions required by the user into multiple tasks, and each functional cell is responsible for one task. The functional cells can work in different redundancy modes, including single-cell redundancy, dual-cell cold backup redundancy, and triple-cell redundancy. In the single-cell redundancy mode, each functional cell is independent and can load different task configurations; in the dual-cell cold backup redundancy mode, the three cells in the organization load the same task configuration, and only two functional cells work simultaneously at the same time. The cold backup module is enabled when the outputs of the working cells are different; in the triple-cell redundancy mode, three cells with the same function work simultaneously.

[0014] The interconnection cell consists of an FSM state machine and a fault voter. The FSM state machine has seven states and is used to output the enable signals (1 enables, 0 masks) of the three functional cells after receiving the mode switching signal or fault signal sent by the processor side through the AXI bus. The fault voter is responsible for judging and locating the faulty cell.

[0015] The functions of the processor side include a self-repair controller, a fault responder, a read-back refresh controller, and an external memory DDR.

[0016] The self - repair controller is responsible for sending pattern signals to interconnected cells at different radiation levels to control the functional cells to work in different redundant modes. The self - repair controller reads the configuration file in the DDR to the PCAP interface to implement the configuration of the functional cell tasks and the reconstruction and repair of cell - level instantaneous faults.

[0017] The fault responder is responsible for responding to the functional module fault signals of the interconnected cells, interrupting the read - back detection of the configuration information of the functional cells at the processor end, and quickly entering the cell - level fault repair state.

[0018] The read - back and refresh controller is responsible for reading back the configuration frames in the functional cells through the PCAP interface and comparing them with the corresponding configuration frames in the functional cell configuration files stored in the memory DDR. When a molecular - level fault is detected, the fault configuration frame is configured and refreshed with the configuration frame in the DDR. If there is still an error in the repaired configuration frame after refreshing, it is determined as a permanent fault and enters the permanent fault repair mode.

[0019] The memory DDR is used to store the original configuration bit files of all functional cells at the logic end.

[0020] Among the three functional cells in the logical function area, according to the priority of internal tasks, there is a main cell and two secondary cells. In the dual - cell cold - backup redundancy mode and the three - cell redundancy mode, the tasks in the main cell are given priority, and the tasks in the secondary cells are assigned to other idle logical function areas and work as their main cells in the current redundancy mode.

[0021] The present invention aims at two types of faults, namely instantaneous faults and permanent faults, which are further divided into cell - level faults and molecular - level faults according to different fault - tolerance granularities. Molecular - level faults are internal configuration information faults of functional modules, and cell - level faults are functional faults of functional modules. In different radiation environments, the self - repair method will change accordingly.

[0022] Environment 1: In a low - radiation environment, the functional cells work in the single - cell redundancy mode. The read - back and refresh controller at the processor end detects whether there are faults inside the functional cells. After detecting a molecular - level fault, it is refreshed and repaired. If there is still an error when reading back the configuration frame again after repair, it is determined as a permanent fault and the logical function organization is masked.

[0023] Environment 2: When the radiation level increases to a medium radiation environment, the functional cells switch to the dual-cell cold backup redundancy mode. The interconnected cells compare the outputs of the working cells. When the outputs are different, the third functional cell is immediately enabled to locate the faulty functional cell and send a cell-level fault signal to the fault responder. The fault responder responds with an interruption, pauses the readback detection of the readback refresh controller, and the self-repair controller completes the reconstruction and repair of the faulty functional cell.

[0024] Environment 3: When the high radiation level is reached, three working cells are enabled to perform the same task. When a cell-level fault occurs, the repair process is similar to that in the medium radiation level.

[0025] Compared with the prior art, the significant effects of the present invention are as follows: 1. By dividing the functional cells for different tasks, the resources occupied by the functional circuit can be more concentrated, and at the same time, the number of configuration frames for readback detection is reduced; 2. The functional cells adopt an adaptive redundancy structure, combined with readback refresh, enabling the system to adapt to different radiation environments; 3. The control module is mainly arranged at the processor end, and at the same time, the PCAP interface is used for readback refresh and reconfiguration, reducing the use of logic resources; 4. The structure is built on the existing commercial ZYNQ series FPGA platform, facilitating transplantation and popularization. Brief Description of the Drawings

[0026] Figure 1 It is a block diagram of the ZYNQ self-repair system based on the adaptive redundancy strategy of the present invention;

[0027] Figure 2 (a) is the working state of the functional cell under low radiation level;

[0028] (b) is the working state of the functional cell under medium radiation level;

[0029] (c) is the working state of the functional cell under high radiation level;

[0030] Figure 3 It is a schematic diagram of the interconnected cell structure;

[0031] Figure 4 It is the state transition diagram of the FSM state machine in the interconnected cell Detailed Embodiment

[0032] The present invention will be further described in detail below in conjunction with the drawings of the specification and the specific embodiments.

[0033] As Figure 1 shown is the system block diagram of the present invention, which includes two parts: the processor end and the logic end. The two parts interact with signals through the AXI bus, and the PCAP interface is used to complete the configuration of the logic end.

[0034] The logic terminals are divided into multiple logic function regions, each region contains three identical functional cells and one interconnection cell. The functional cells can operate in different redundancy modes. The interconnection cell can switch the redundancy mode of the functional cells, locate the faulty cells and send fault information to the processor terminal.

[0035] The functions of the processor terminal include a self-healing controller, a fault responder, a read-back refresh controller, and an external memory DDR. The self-healing controller is responsible for sending mode signals to the interconnection cell under different radiation levels, and completing the repair of faults by reading the configuration file from the DDR through the PCAP interface according to the fault signal. The fault responder is responsible for responding to the functional module fault signal of the interconnection cell, interrupting the read-back detection of the configuration information of the functional cells at the processor terminal, and quickly entering the cell-level fault repair state. The read-back refresh controller is responsible for the read-back refresh repair of the configuration information in the functional cells through the PCAP interface, and reading the error frame again after the repair to determine whether it is a permanent fault. The memory DDR is used to store the original configuration bit files of all functional cells at the logic terminal.

[0036] The functional cells, namely the reconfigurable modules, decompose the functions required by the user into multiple tasks, and each functional cell is responsible for one task. The functional cells can operate in different redundancy modes, including single-cell redundancy, dual-cell cold backup redundancy, and triple-cell redundancy. In the single-cell redundancy mode, each functional cell is independent and can load different task configurations; in the dual-cell cold backup redundancy mode, the three cells in the organization load the same task configuration, and only two functional cells work at the same time, and the cold backup module is enabled when the outputs of the working cells are different; in the triple-cell redundancy mode, three cells with the same function work at the same time.

[0037] Such as Figure 2The working states of functional cells under three radiation levels of high, medium, and low, and the fault tolerance methods. The functional cell MC is the main cell in the tissue, and SC1 and SC2 are secondary cells. In a low-radiation environment, each functional cell is independent of each other. MC, SC1, and SC2 can respectively complete different tasks. The processor end detects and repairs molecular-level transient faults in the configuration information of the functional cells through read-back refresh. When a permanent fault occurs, the functional area is shielded. When the radiation intensity reaches the medium radiation level, the task of the main cell in the tissue is given the highest priority. In Fig. (b) TISSUE, the main cell MC1 and the secondary cell SC1 fuse to jointly complete the task of MC in Fig. (a). The output signals of MC1 and SC1 are output after being compared by the interconnection cells. The secondary cell SC12 enters the standby state. After detecting a fault through the fault voter in the interconnection cells, the standby cells are immediately enabled to assist in locating the faulty cells and replacing and repairing them. The tasks in SC1 and SC2 in (a) are configured to other functional areas. When resources are scarce, the convenience of the on-chip processor end can also be utilized to arrange tasks with lower priorities to the processor end for implementation. When the radiation intensity is at the high radiation level, all three functional cells in the logic functional area are enabled, and the task in the main cell MC is undertaken by an entire logic functional area.

[0038] Figure 3 It is the internal structure diagram of the interconnection cell, which consists of an FSM state machine and a fault voter. The FSM state machine is used to output the enable signals (1 enables, 0 shields) of the three functional cells correspondingly after receiving the mode switching signal or fault signal sent by the processor end through the AXI bus. The fault voter is responsible for judging and locating the faulty cells.

[0039] Figure 4 It is the state transition diagram of the state machine in the interconnection cell. The following explains each state of the state machine in combination with the function of the fault voter.

[0040] (1) Idle mode S0: In the S0 state, all internal functional modules in the logic functional area are empty-configured. After waiting for the system to allocate tasks to the functional cells, it enters the single-cell working state S1.

[0041] (2) Single-cell working state S1: The system initially works in the S1 state. The functional cells in the logic functional area have had their internal circuits configured by dynamic partial reconfiguration. At this time, the logic functional area is in a low-radiation level state. The interconnection cell enables the corresponding number of functional cells according to the number of tasks (enable signal 0 shields, 1 enables), and the fault voter connects the corresponding input and output to the input and output interfaces of the functional cells. When receiving the molecular-level transient fault signal sent by the self-repair controller, it jumps to the molecular-level transient fault repair waiting state S5. When receiving the radiation level change signal, it enters the mode switching waiting state S2.

[0042] (3) Mode Switching Wait State S2: Due to the change in radiation level, the cells within the logic function area need to have their tasks reassigned by the self-repair controller through dynamic partial reconfiguration, with the tasks of the primary cells taking precedence, and changing the internal functional circuits of the secondary cells to complete the switching of the functional cell redundancy mode. Therefore, in the mode switching wait state S3, the input and output signals of the secondary cells are cut off, and after waiting for the functional cells to be configured, it is switched to the single-cell working state S1, the dual-cell cold backup working state S3, or the triple-cell redundancy working state S4 according to the current radiation level.

[0043] (4) Dual-cell Cold Backup Working State S3: In the medium working level mode S3, the primary cell MC and a secondary cell SC1 work together to complete the task. The enable signals of MC and SC1 are pulled high. The fault voter connects the input din_1 to the input interfaces of MC and SC1. The outputs of the two cells MC and SC1 are compared and then connected to the output interface dout_1. When a molecular-level instantaneous fault is detected, it enters the molecular-level instantaneous fault repair state S5. When the high radiation level is reached or the outputs of the two cells are inconsistent, the interconnection cell enables the secondary cell SC2, that is, it jumps to the triple-cell redundancy working state S4.

[0044] (5) Triple-cell Redundancy Working State S4: When the system operates in a high radiation level environment, a high system reliability is required to ensure the normal operation of the system. Therefore, three functional cells within the logic function area are enabled to work together. The three cells share the signal input signal din_1. The fault voter outputs the signal after comparison through the dout_1 interface, and at the same time sends the faulty cell number Error_cell (00 means no fault, 01 means MC fault, 10 means SC1 fault, 11 means SC2 fault) and the error signal to the state machine. When the error signal is at a high level, it means that a cell has a fault, and it will jump to the cell-level instantaneous fault repair wait state S6.

[0045] (6) Molecular-level Instantaneous Fault Repair Wait State S5: The S5 state represents that the self-repair control organization detects an instantaneous fault in a certain functional cell. First, the enable signal of this functional cell is pulled low, and it waits for the processor side to repair the molecular layer configuration information in the functional cell. After writing the configuration information, it reads this frame again for comparison. If it is the same as the corresponding frame in the original configuration file, the repair is completed, and it returns to the original radiation level working mode to continue working. If this frame still reports an error after comparison, it means that this frame has a permanent fault, and it enters the permanent fault masking state S7.

[0046] (7) Cell-level transient fault repair waiting state S6: After entering the S6 state, according to Error_cell, the enable signal of the corresponding faulty cell is pulled low, and after waiting for the self-repair control organization to reconfigure and repair the faulty cell, it returns to continue working in the original redundant mode according to the radiation level mode signal. If the fault remains unrepaired, it is judged that a permanent fault has occurred and enters the permanent fault shielding state S7.

[0047] (8) Permanent fault shielding state S7

[0048] In the S7 state, an irreparable permanent fault has occurred in the functional organization. To prevent this part of the circuit from affecting the system function, this logical function area is shielded by pulling low the enable signals of three functional cells, and the processor allocates the tasks of this area to the idle logical function area to complete the repair.

Claims

1. A ZYNQ series FPGA self-repair structure based on an adaptive redundancy strategy, characterized in that: It includes the processor side and the logic side. The two parts exchange signals through the AXI bus, and the PCAP interface is used to complete the configuration of the logic side. The logic end is divided into multiple logic function areas, each area contains three identical function cells and one interconnection cell; the function cells can work in different redundancy modes; the interconnection cell can switch the redundancy mode of the function cells, locate the faulty cells and send fault information to the processor end; Processor-side functions include self-repair controller, fault responder, read-back refresh controller, and external memory DDR; The self-repair controller is responsible for sending mode signals to interconnected cells at different radiation levels, and completing fault repair by reading configuration files from DDR through the PCAP interface according to the fault signal; The fault responder is responsible for responding to the fault signal of the functional module of the interconnected cell, interrupting the processor end to read back the configuration information of the functional cell, and quickly entering the cell-level fault repair state; The readback refresh controller is responsible for reading back, refreshing and repairing the configuration information in the functional cell through the PCAP interface, and reading back the error frame again after the repair to determine whether it is a permanent fault; The memory DDR is used to store the original configuration bit files of all functional cells of the logic end.

2. According to claim 1, a ZYNQ series FPGA self-repair structure based on an adaptive redundancy strategy is characterized in that: The three functional cells in the logical function area are divided into a main cell and two secondary cells according to the priority of the internal tasks. In the dual-cell cold backup redundancy mode and the three-cell redundancy mode, the tasks in the main cell are given priority, and the tasks in the secondary cells are assigned to other idle logical function areas, working as their main cells in the current redundancy mode.

3. According to claim 1, a ZYNQ series FPGA self-repair structure based on an adaptive redundancy strategy is characterized in that: The interconnected cell consists of an FSM state machine and a fault voter; the FSM state machine has seven states, which is used to output the enable signals (1 enable, 0 shield) of the three functional cells accordingly after receiving the mode switching signal or fault signal sent by the processor through the AXI bus; the fault voter is responsible for judging and locating the faulty cell and sending the fault information to the processor.

4. The ZYNQ series FPGA self-repairing structure based on the adaptive redundancy strategy according to claim 1, characterized in that: The present invention targets two types of faults, namely transient faults and permanent faults, which are further divided into cellular-level faults and molecular-level faults according to the different fault-tolerance granularities. Molecular-level faults are internal configuration information faults of functional modules, while cellular-level faults are functional faults of functional modules. Under different radiation environments, the self-repair method will have corresponding changes. Environment 1: In a low-radiation environment, the functional cells work in single-cell redundancy mode. The processor-side readback refresh controller detects whether there is a fault inside the functional cell. After a molecular-level fault is detected, it is refreshed and repaired. If the configuration frame is still wrong after the repair, it is determined to be a permanent fault and the logical function organization is shielded. Environment 2: When the radiation level increases to a medium radiation environment, the functional cell switches to a dual-cell cold backup redundancy mode, and the interconnected cells compare the outputs of the working cells. When the outputs are different, the third functional cell is immediately enabled to locate the faulty functional cell and send a cell-level fault signal to the fault responder. The fault responder responds to the interrupt, suspends the readback detection of the readback refresh controller, and the self-repair controller completes the reconstruction and repair of the faulty functional cell. Environment 3: When high radiation levels are reached, three working cells are enabled to perform the same tasks, and when a cell-level failure occurs, the repair process is similar to that under medium radiation levels.

Citation Information

Cited By

  • Interconnection efficient on-chip automatic repair circuit and method based on multidirectional shift

    CN122044964A

  • An efficient on-chip repair circuit and method based on multi-directional shifting interconnects

    CN122044964B