Protection Method and System Based on Single-Effect Error State Coding Characterization of Key Signals

By calibrating the ID numbers of key signals in the FPGA circuit and designing STMR, combined with the monitoring unit program of the PS system, real-time detection and recovery of single-event soft errors were achieved. This solved the problem that traditional TMR could not characterize the error state, reduced resource overhead and transmission bandwidth, and improved system reliability.

CN120074747BActive Publication Date: 2025-12-02NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411832487.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-12-02
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Traditional TMR technology can only tolerate errors but cannot characterize error states, leading to redundant failures of critical signals in user functional circuits due to error accumulation. When a large number of critical signals in complex logic circuits malfunction, it is difficult to accurately detect the operating status of the FPGA circuit system and fully characterize its functional state. At the same time, the overhead of signal transmission interfaces and storage resources is too large.

Method used

A protection method based on single-event error state coding characterization of key signals is adopted. By calibrating the ID number and designing STMR for key signals of user function circuits in the PL system, an error state detection circuit is inserted. The monitoring unit program in the PS system is used to analyze data frames and control frames to achieve real-time detection and recovery of error states.

Benefits of technology

It reduces the resource overhead and data transmission bandwidth of user function circuits, improves the reliability and real-time performance of FPGA circuit systems, and can detect and recover from single-event soft errors in a timely manner, making it suitable for digital circuit systems of aerospace electronic equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074747B_ABST
    Figure CN120074747B_ABST
Patent Text Reader

Abstract

This invention provides a protection method and system based on single-event error state encoding representation of critical signals. It encodes critical signals (STMRs) and error critical signal IDs, and uses data frames and control frames to represent and store the error states of these critical signals. Finally, the monitoring unit program in the PS system performs error state analysis and formulates protection decisions. This invention can reduce the resource overhead and data transmission bandwidth of user function circuits, while improving the reliability of user function circuit systems implemented on SoC-type FPGAs. It is applicable and operable to digital circuit systems for aerospace electronic equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to hardening technology for space radiation effects, and more particularly to a protection method based on the coding characterization of critical signal error states. Background Technology

[0002] SRAM-based FPGAs, or Field-Programmable Gate Arrays based on Static Random Access Memory, contain a large number of configuration memory cells. They are programmed to implement specific functions required by the user and offer advantages such as high integration, strong parallel computing capabilities, short development cycles, and reprogrammability, making them widely used in spacecraft electronics. With the continuous development of microelectronics technology, FPGAs have evolved from single programmable logic resources to heterogeneous SoC-type FPGAs integrating components based on different architectures and processes. Xilinx's heterogeneous SoC-type FPGAs integrate a Programmable Logic (PL) system and a Processing System (PS) system; some models even integrate radio frequency components. The PL is the programmable logic part, allowing users to implement specific hardware logic functions according to their needs. The PS is the processor system part, typically referring to one or more processor cores integrated within the FPGA, such as the ARM Cortex-A series processors. The PS provides a software-programmable environment for users to run operating systems and applications.

[0003] SRAM-based FPGAs are single-event upset (SEU) sensitive devices. Their numerous internal memory cells are highly susceptible to SEUs under space irradiation, manifesting as soft-error coupling in sequential logic systems and ultimately causing malfunctions in user functional circuits. SRAM-based FPGAs typically employ a combination of Triple Module Redundancy (TMR) and refresh for SEU protection. A schematic diagram of a traditional TMR implementation is shown in Figure X. However, TMR incurs a 3-4 times increase in resource overhead. With the increasing complexity of user functional circuit designs, performing TMR on the entire design is not feasible. A viable approach is to perform partial TMR on critical modules and signals. Since registers and block memories in user functional circuits cannot be restored to correct data after an SEU, TMR also suffers from error accumulation leading to failure. To avoid TMR failure, critical signals in user functional circuits not only require TMR but must also be promptly restored after a SEU to ensure that redundant critical signals output correct results.

[0004] Current research on critical signal protection methods for SRAM-based FPGA sequential logic circuits focuses on the selection of critical signals. To avoid transmission interference between signals, critical signals are often transmitted through independent signal lines and device pins. In complex designs, detecting the state of critical signals consumes significant interface resources or data transmission bandwidth, increasing data storage overhead and the risk of single-event soft errors, ultimately affecting the reliability of critical signal-based detection and protection decisions.

[0005] Since the heterogeneous SoC-type FPGA provides two systems, PL and PS, the processor in the PS system can be used to check the key signals of the user function circuit in the PL system for single-event upset detection and recovery, thereby improving the user function circuit's resistance to single-event soft errors without increasing system hardware overhead. Summary of the Invention

[0006] To overcome the shortcomings of existing technologies, this invention provides a protection method and system based on single-event error (SEE) state encoding and characterization of critical signals. Existing technologies primarily focus on the analysis and identification of critical signals in user function circuits within FPGAs, neglecting the excessive transmission bandwidth and storage overhead caused by monitoring the states of numerous critical signals in complex user function circuits. Furthermore, traditional TMR (Transmission Moderator) technology can only achieve fault-tolerant design and cannot promptly detect and respond to single-event soft errors in multiple redundant circuits, ultimately leading to user function anomalies due to the accumulation of SEE. This invention provides a protection system and method based on SEE state encoding and characterization of critical signals. It encodes critical signals (STMR) and error critical signal IDs, and uses data frames and control frames to characterize and store the error states of critical signals. Finally, the monitoring unit program in the PS (Power Supply System) completes error state analysis and forms protection decisions. This invention can reduce the resource overhead and data transmission bandwidth of user function circuits, while improving the reliability of user function circuit systems implemented on SoC-type FPGAs. It is applicable and operable for digital circuit systems in aerospace electronic equipment.

[0007] The main technical problems solved by this invention are: 1) Traditional TMR can only tolerate errors but cannot characterize error states, leading to redundant failures of key signals in user functional circuits due to error accumulation; 2) When a large number of key signals in complex logic circuits malfunction, how to achieve accurate detection of the operating status of FPGA circuit systems and comprehensive characterization of functional states is solved; 3) By performing information fusion-based characterization encoding on a large number of key signals, the problem of excessive overhead of signal transmission interfaces and storage resources is solved.

[0008] The technical solution adopted by this invention to solve its technical problem is:

[0009] A protection method based on single-event error state coding representation of key signals includes the following steps:

[0010] Step S10: The key signals of the user function circuit in the PL system are assigned ID numbers in descending order of function weight. The ID numbers from small to large correspond to the function weights from high to low. A maximum of 65536 key signals are supported. 0 is the highest priority ID number and 65535 is the lowest priority ID number.

[0011] Step S20: Perform state triple module redundancy (STMR) design on the key signals of the user function circuit in the PL system. Based on copying the registers storing the key signals three times, insert an error state detection circuit. When any register causes a data error due to a single-event flip, the error indication signal is pulled high and sent to the key signal encoding and characterization module.

[0012] In step S30, the key signal encoding and characterization module in the PL system performs binary encoding on the key signal ID number that causes single-event flip and sends the encoding result to the framing module.

[0013] In step S40, the framing module in the PL system performs single-event upset monitoring on key signals in the user logic circuit. When a single-event upset occurs on a key signal, a data frame and a control frame are generated and written into an external DDR memory as the basis for the monitoring program in the PS system to make protection decisions.

[0014] In step S50, the monitoring unit program running on the processor in the PS system periodically accesses the DDR memory to read control frames. If the control frame is updated, the data frame is read according to the control frame information. The error key signal ID number in the data frame is extracted, and the number of key signal errors is analyzed and calculated. When the number of errors is greater than the threshold, a protection decision is issued to realize single-event soft error recovery of user function circuits.

[0015] The data frame and control frame generation process in step S40 includes the following steps:

[0016] In step S401, under the control of the internal data frame timestamp counter, the framing module in the PL system generates a data frame and writes it to the external DDR memory whenever the counter increments by 1 if the error indication signal of the critical signal STMR circuit is detected to be high. The error critical signal ID number code and the error occurrence time information are recorded. When the data frame timestamp counter is full, it indicates that one round of critical signal detection is completed. After the counter is cleared, it starts counting again and continues to perform error status detection of critical signals.

[0017] Step 402: Whenever a data frame is generated, the framing module in the PL system increments the error_cnt value in the control frame by 1. When the data frame timestamp counter is full, it increments the control frame timestamp counter by 1, performs framing encoding according to the control frame format, and writes the control frame to the DDR memory. When the control frame timestamp counter is full, it is reset to zero and the counting starts again.

[0018] In step S50, the process of the monitoring unit reading control frames and data frames, performing data parsing, and generating protection decisions includes the following steps:

[0019] Step S501: The monitoring unit program in the PS system periodically reads the control frame; first, it determines whether the timestamp in the control frame has been updated. If the timestamp has been updated, it means that a round of key signal detection has been completed and data frame extraction and parsing are required; otherwise, there is no need to read the data frame.

[0020] In step S502, the monitoring unit program in the PS system extracts the error_cnt value from the control frame and determines whether the error_cnt signal is 0. If the signal is 0, it means that there is no critical signal error at present, and returns to step S501 to continue execution; if the signal is not 0, it means that a critical signal error has occurred, and jumps to step S503 to continue execution.

[0021] Step S503: The monitoring unit program in the PS system reads error_cnt data frames from the DDR memory using error_cnt as the offset address, and extracts the time information and error key signal ID number from the data frames;

[0022] Step S504: The monitoring unit program in the PS system analyzes the error key signal ID number in the data frame and removes duplicate ID numbers. That is, when the error key signal ID number is the same in multiple data frames, the error is counted only once. Then the number of error key signals is summed to obtain the total number of errors.

[0023] Step S505: The monitoring unit program in the PS system determines whether the total number of errors is greater than the protection decision threshold. If not, it jumps to step S501 to continue execution; if so, it jumps to step S506 to continue execution.

[0024] Step S506: If the number of errors exceeds the protection decision trigger threshold, it indicates that there are many critical signal errors, which are very likely to affect the normal function of the user functional circuit. Single-event soft error recovery needs to be performed immediately. The monitoring unit program in the PS system forms a protection decision and issues it, thereby realizing single-event soft error protection for the user functional circuit in the PL system.

[0025] More preferably, in step S10, assigning ID numbers to key signals is based on the functional weight of the key signal in the user functional circuit, assigning different ID numbers to the key signals. A maximum of 65536 key signals are supported. 0 is the highest priority ID number, and 65535 is the lowest priority ID number.

[0026] More preferably, in step S20, the key signal STMR is an improved circuit structure based on traditional triple modular redundancy, adding an error state detector. This aims to overcome the drawback of the classic TMR, which cannot clearly indicate whether an error has occurred. In the specific implementation, the output logic unit of the key signal (generally represented as a LUT or register in netlist logic) is first copied three times. Then, the output of each module is input to the majority voter and the error state detector respectively. The majority voter operates in the same way as the classic triple modular redundancy, i.e., the minority obeys the majority principle. The error state detector's default output is 0, indicating no error has occurred. When any of the three modules differs from the other two, the error state detector outputs 1, indicating that the key signal has flipped, and drives the subsequent error state encoding module to perform encoding operations. The error state indication signal of each key signal has a bit width of 1 bit, where '1' represents an abnormal state and '0' represents a normal state. A maximum of 65536 key signals are supported.

[0027] More preferably, in step S30, the key signal ID number encoding maps the ID number from a decimal number to a binary number to reduce the data capacity for representing and transmitting a large number of key signal error states, which is then used for subsequent frame generation of data frames. The ID number encoding information is a 16-bit binary number, supporting an ID number range of 0 to 65535.

[0028] More preferably, in step S40, the framing encoding module detects and frames the error states of key signals and user-defined signals, defining two types of frame formats: data frames and control frames.

[0029] The data frame records the key signal ID number and the time information of the error occurrence within one error detection cycle. The data frame format is shown in Table 1.

[0030] Table 1 Data Frame Structure

[0031] Frame header Timestamp Key signal ID encoding information Reserved bits CRC check 4’b1010 28bit 16bit 8bit 8bit

[0032] The header of the data frame is 4'b1010, where b represents binary, the 4 before b represents 4 binary numbers, and 1010 represents the actual binary number content.

[0033] The data frame timestamp counter is 28 bits long and is used to record the time when the error occurs. It is driven by the internal clock of the heterogeneous SoC-type FPGAPL system, and the start time of counting is synchronized with the start time of the user function circuit. After the data frame timestamp is full, it is cleared to zero and the counting starts again. The data frame timestamp starts counting from 0 and continues until it is full, which constitutes one critical signal detection cycle.

[0034] The critical signal ID encoding information consists of 16 bits, supporting ID numbers from 0 to 65535, and is the binary encoding of the critical signal DI number that caused the error. If multiple critical signals are in error at the same time, only the critical signal ID number with the highest weight is encoded.

[0035] The reserved bits are 8 bits, all of which are filled with 1'b0.

[0036] The CRC check value is 8 bits, which is a cyclic redundancy check code that corrects one error and checks two errors. The check range starts from the frame header and ends at the reserved bits.

[0037] The control frame uses the key signal detection cycle of the data frame as the time counting unit, and periodically feeds back program data information to the monitoring unit. This data information records the cycle counting timestamp, the error count value (error_cnt) for key signals in the data frame, and the CRC check value. The control frame format is shown in Table 2:

[0038] Table 2 Control Frame Structure

[0039] Frame header Timestamp Error_cnt CRC check 4’b0101 28bit 24bit 8bit

[0040] The frame header of the control frame is 4'b0101, where b represents binary, the 4 before b represents 4 binary numbers, and 0101 represents the specific content of the binary number.

[0041] The control frame timestamp is 28 bits. When the timestamp of the data frame is full, the timestamp count of the control frame is incremented by 1'b1. After the timestamp of the control frame is full, it is reset to zero and the counting starts again.

[0042] The error_cnt is a data frame count, which is incremented by 1 each time a data frame is generated. error_cnt-1 is the address offset used by the monitoring unit to control the reading of data frames from the DDR memory.

[0043] The CRC check value is 8 bits, which is a cyclic redundancy check code that corrects one error and checks two errors. The check range starts from the frame header and ends at error_cnt.

[0044] The present invention also provides a protection system based on single-event error state encoding characterization of key signals, which includes: a PL system, a PS system and an external DDR memory.

[0045] The PL and PS systems are connected to an external DDR memory. They are responsible for designing critical signal STMRs in user function circuits, marking critical signal IDs, and encoding critical signal IDs for error states. They are also responsible for assembling data frames and control frames and writing them to the external DDR memory.

[0046] The external DDR memory is connected to the PL and PS systems and is responsible for storing data frames and control frames. The data frame storage space address range is 0 to 65535, and a maximum of 65536 frames can be stored; the control frame storage space address is 65536, and only one frame is stored.

[0047] The PS system and PL system are connected to the external DDR memory. The PS system is responsible for reading control frames and data frames from the external DDR memory, extracting error status key signal ID information and user-defined signal status information from the data frames, performing deduplication processing, counting the number of errors, comparing them with the protection decision threshold, and finally generating a protection decision to complete the single-event soft error recovery of the PL system user function circuit.

[0048] The PL system includes one or more key signal STMR modules; a key signal encoding and characterization module; a custom signal statistics module; and a framing module.

[0049] One or more critical signal STMR modules perform triple modulo redundancy and insert error detectors on critical signals in the user functional circuits. If a single-event soft error occurs in any of the three circuits, the error detector will pull up the output error status indication signal and transmit it to the critical signal encoding and characterization module. Up to 65,536 critical signal STMR modules can be inserted to perform triple modulo redundancy and error detection on 65,536 critical signals.

[0050] A key signal encoding and characterization module receives the error state indication signal output by the key signal STMR module, marks the key signals with ID numbers according to their functional weights, performs binary encoding of the ID numbers for the error state key signals, and transmits the encoding results to the framing module. Key signals in normal state are not encoded with ID numbers.

[0051] A custom signal statistics module detects whether errors have occurred in user-defined signals and transmits the results to the framing module.

[0052] A framing module, controlled by a data frame timestamp counter, completes one data frame assembly and writes it to DDR memory each time the timestamp counter increments by 1. The data frame timestamp counter starts counting from 0 and reaches full at 65535, indicating the completion of one critical signal error state detection cycle. At this time, the control frame timestamp counter increments by 1, and a control frame is generated and written to DDR memory. Both the data frame and control frame timestamp counters are reset to zero and the counting restarts after they are full.

[0053] The PS system includes one or more processors and a monitoring unit program.

[0054] One or more processors: The monitoring unit program is configured to be executed by said one or more processors.

[0055] A monitoring unit program periodically reads control frames from DDR memory and reads data frames from DDR memory using the `error_cnt` value in the control frame as the address offset. The monitoring unit program counts the number of critical signal errors in the data frames, deduplicates critical signals with the same ID number across multiple data frames to avoid counting a persistently erroneous critical signal as multiple errors. It then compares the count of error-state critical signals with a user-preset threshold. When the number of errors exceeds the threshold, it indicates that too many critical signals are being processed for error states, and the user's functional circuitry is highly likely to malfunction. Protective measures should be taken, and the monitoring unit generates and issues protective decisions. These decisions include resetting the PL system, reloading the PL system, and simultaneously reloading both the PL and PS systems.

[0056] The beneficial effects of this invention are as follows:

[0057] STMR is performed on key signals in user functional circuits. Based on the traditional three-mode redundancy, an error state detector is added. When a single-event soft error occurs in one of the modes, the error state can be indicated, which serves as the basis for the monitoring unit to make protection decisions. This can prevent the accumulation of errors from ultimately causing the redundancy measures to fail.

[0058] Key signals are encoded with ID numbers according to their functional weights. The ID numbers reflect the degree of influence of key signals on the functions of user functional circuits. The error status information of 65,536 key signals is converted into 16-bit binary encoded data, which realizes the compression of large data information and reduces the logic resource overhead of user functional circuits. At the same time, the on-chip high-speed shared bus interface of heterogeneous SoC FPGA is used to realize high-speed data transmission, reduce bandwidth occupancy, and improve the real-time performance of key signal error status detection.

[0059] By using control frame timestamps to identify whether control frames have been updated, the runtime of the monitoring unit is reduced. The `error_cnt` information is only extracted when the control frame timestamp differs from the previously read timestamp. Furthermore, using `error_cnt` as the offset address for the monitoring unit to read data frames avoids the significant time overhead of reading 65,536 data frames. The combination of timestamps and `error_cnt` in the control frames reduces the time overhead for the monitoring unit to read data and control frames, improving the real-time performance of the monitoring unit in acquiring effective information and making decisions.

[0060] By leveraging the architecture of heterogeneous SoC-type FPGAs that integrate PL and PS systems, and utilizing one or more processor cores in the PS system to run monitoring unit programs, the operating status of user function circuits running in the PL system can be detected without adding additional components, enabling timely protection decisions.

[0061] The present invention is applied to the design of single-event soft error detection and protection for user function circuits based on heterogeneous SoC-type FPGAs. The design is highly targeted, and its system architecture covers the design of mainstream aerospace digital circuit systems. The method has a certain degree of universality and can be extended to the design of single-event soft error protection for traditional architecture FPGAs. Attached Figure Description

[0062] Figure 1 This invention provides a flowchart for the coding representation and protection decision-making of critical signal error states.

[0063] Figure 2 This is a schematic diagram of a single-event soft error protection system for key signal encoding and characterization provided by the present invention.

[0064] Figure 3 This invention provides a schematic diagram of a key signal STMR circuit.

[0065] Figure 4 This is a schematic diagram of a key signal ID number encoding provided by the present invention. Detailed Implementation

[0066] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0067] This invention provides a protection method based on single-event error state coding representation of key signals, such as... Figure 1 As shown, the method includes the following steps:

[0068] Step S10: The key signals of the user function circuits in the PL system are assigned ID numbers according to their functional weights from highest to lowest. The ID numbers, from smallest to largest, correspond to functional weights from highest to lowest. A maximum of 65536 key signals are supported. 0 represents the highest priority ID number, and 65535 represents the lowest priority ID number.

[0069] Step S20: Perform state triple module redundancy (STMR) design on the key signals of the user function circuit in the PL system. Based on copying the registers storing the key signals three times, insert an error state detection circuit. When any register causes a data error due to a single-event flip, the error indication signal is pulled high and sent to the key signal encoding and characterization module.

[0070] In step S30, the key signal encoding and characterization module in the PL system performs binary encoding on the key signal ID number that causes single-event flip and sends the encoding result to the framing module.

[0071] In step S40, the framing module in the PL system performs single-event upset monitoring on key signals in the user logic circuit. When a single-event upset occurs on a key signal, a data frame and a control frame are generated and written to an external DDR memory as the basis for the monitoring program in the PS system to make protection decisions.

[0072] In step S50, the monitoring unit program running on the processor in the PS system periodically accesses the DDR memory to read control frames. If the control frame is updated, the data frame is read according to the control frame information. The error key signal ID number in the data frame is extracted, and the number of key signal errors is analyzed and calculated. When the number of errors is greater than the threshold, a protection decision is issued to realize single-event soft error recovery of user function circuits.

[0073] The data frame and control frame generation process in step S40 includes the following steps:

[0074] In step S401, under the control of the internal data frame timestamp counter, the framing module in the PL system generates a data frame and writes it to the external DDR memory whenever the counter increments by 1 if the error indication signal of the critical signal STMR circuit is detected to be high. The error critical signal ID number and the time of error occurrence are recorded. When the data frame timestamp counter is full, it indicates that one round of critical signal detection is completed. After the counter is cleared, it starts counting again and continues to perform error status detection of critical signals.

[0075] Step 402: Whenever a data frame is generated, the framing module in the PL system increments the error_cnt value in the control frame by 1. When the data frame timestamp counter is full, it increments the control frame timestamp counter by 1, performs framing encoding according to the control frame format, and writes the control frame to the DDR memory. When the control frame timestamp counter is full, it is reset to zero and the counting starts again.

[0076] The process of the monitoring unit reading control frames and data frames, performing data parsing, and generating protection decisions in step S50 includes the following steps:

[0077] In step S501, the monitoring unit program in the PS system periodically reads the control frame. First, it determines whether the timestamp in the control frame has been updated. If the timestamp has been updated, it indicates that a round of key signal detection has been completed, and data frame extraction and parsing are required; otherwise, there is no need to read the data frame.

[0078] In step S502, the monitoring unit program in the PS system extracts the value of error_cnt from the control frame and determines whether the error_cnt signal is 0. If it is 0, it means that there is no critical signal error, and the program returns to step S501 to continue execution; if it is not 0, it means that a critical signal error has occurred, and the program jumps to step S503 to continue execution.

[0079] Step S503: The monitoring unit program in the PS system reads error_cnt data frames from the DDR memory using error_cnt as the offset address, and extracts the time information and error key signal ID number from the data frames.

[0080] Step S504: The monitoring unit program in the PS system analyzes the error key signal ID number in the data frame and removes duplicate ID numbers. That is, when the error key signal ID number is the same in multiple data frames, the error is counted only once. Then, the number of error key signals is summed to obtain the total number of errors.

[0081] Step S505: The monitoring unit program in the PS system determines whether the total number of errors is greater than the protection decision threshold. If not, it jumps to step S501 to continue execution; if so, it jumps to step S506 to continue execution.

[0082] Step S506: If the number of errors exceeds the protection decision trigger threshold, it indicates that there are many critical signal errors, which are very likely to affect the normal function of the user functional circuit. Single-event soft error recovery needs to be performed immediately. The monitoring unit program in the PS system forms a protection decision and issues it, thereby realizing single-event soft error protection for the user functional circuit in the PL system.

[0083] In this implementation example, the key signal STMR in step S20 is an improvement on the traditional triple-modular redundancy circuit structure, adding an error state detector. Its purpose is to overcome the drawback of the classic TMR, which cannot clearly indicate whether an error has occurred. The STMR circuit schematic is shown below. Figure 2 As shown. In the specific implementation, the output logic unit of the key signal (generally represented as a LUT or register in the netlist logic) is first copied three times. Then, the output of each module is input to the majority voter and the error status detector respectively. The majority voter operates in the same way as the classic three-module redundancy, i.e., the minority obeys the majority principle. For the error status detector, its default output is 0, indicating that no error has occurred. When any of the three modules is different from the other two, the error status detector outputs 1, indicating that this key signal has been toggled, and drives the subsequent error status encoding module to perform the encoding operation. The error status indication signal of each key signal has a bit width of 1 bit, where '1' represents an abnormal state and '0' represents a normal state. A maximum of 65536 key signals are supported.

[0084] In this implementation example, step S20 involves binary encoding of the key signal ID number, as illustrated in the following diagram. Figure 3 As shown.

[0085] This invention also provides a single-event soft error (SEM) protection system based on key signal encoding and characterization. The system supports encoding and characterizing key signals of user-defined functional circuits in a power PL (Power Producer) circuit, and performs frame encoding on key signals and user-defined signals to generate data frames and control frames. By analyzing the data frames and control frames, a protection strategy for the PL system is generated, achieving SEM protection for the PL system. The corresponding system structure of this invention is as follows: Figure 4 As shown, it includes: PL system, external DDR memory, and PS system.

[0086] The PL system is connected to the PS system and external DDR memory. It is responsible for designing key signals STMR in user function circuits, marking key signal IDs, and encoding key signal IDs for error states. It is also responsible for assembling data frames and control frames and writing them to external DDR memory.

[0087] External DDR memory: Connected to the PL and PS systems, responsible for storing data frames and control frames. The data frame storage space address ranges from 0 to 65535, and can store up to 65536 frames; the control frame storage space address is 65536, and only one frame is stored.

[0088] PS system: Connected to the PL system and external DDR memory, it is responsible for reading control frames and data frames from the external DDR memory, extracting error status key signal ID information and user-defined signal status information from the data frames, performing deduplication processing, counting the number of errors, comparing them with the protection decision threshold, and finally generating a protection decision to complete the single-event soft error recovery of the PL system user function circuit.

[0089] The aforementioned PL system includes: one or more key signal STMR modules; a key signal encoding and characterization module; a custom signal statistics module; and a framing module.

[0090] One or more critical signal STMR modules: These modules perform triple modulo redundancy and insert error detectors for critical signals in the user functional circuits. If a single-event soft error occurs in any of the three circuits, the error detector will pull up the output error status indication signal and transmit it to the critical signal encoding and characterization module. Up to 65,536 critical signal STMR modules can be inserted to perform triple modulo redundancy and error detection on 65,536 critical signals.

[0091] A key signal encoding and characterization module: receives the error state indication signal output from the key signal STMR module, marks the key signals with ID numbers according to their functional weights, performs binary encoding of the ID numbers for the error state key signals, and transmits the encoding results to the framing module. Key signals in normal states are not encoded with ID numbers.

[0092] A custom signal statistics module: responsible for detecting whether errors have occurred in user-defined signals and transmitting the results to the framing module.

[0093] A framing module: Under the control of a data frame timestamp counter, each time the timestamp counter increments by 1, a data frame is assembled and written to the DDR memory. The data frame timestamp counter starts counting from 0 and reaches full at 65535, indicating the completion of a critical signal error state detection cycle. At this time, the control frame timestamp counter increments by 1, and a control frame is generated and written to the DDR memory. Both the data frame and control frame timestamp counters are reset to zero and the counting restarts after they are full.

[0094] The aforementioned PS system includes one or more processors and a monitoring unit program.

[0095] One or more processors: The monitoring unit program is configured to be executed by said one or more processors.

[0096] A monitoring unit program periodically reads control frames from DDR memory and reads data frames from DDR memory using the `error_cnt` value in the control frame as the address offset. The monitoring unit program counts the number of critical signal errors in the data frames, deduplicates critical signals with the same ID number across multiple data frames to avoid counting a persistently erroneous critical signal as multiple errors. It then compares the count of error-state critical signals with a user-preset threshold. When the number of errors exceeds the threshold, it indicates that too many critical signals are being processed for error states, and the user's functional circuitry is highly likely to malfunction. Protective measures should be taken, and the monitoring unit generates and issues protective decisions. These decisions include resetting the PL system, reloading the PL system, and simultaneously reloading both the PL and PS systems.

Claims

1. A protection method based on single-event error state coding representation of key signals, characterized in that... Includes the following steps: Step S10: The key signals of the user function circuit in the PL system are assigned ID numbers in descending order of function weight. The ID numbers from small to large correspond to the function weights from high to low. A maximum of 65536 key signals are supported. 0 is the highest priority ID number and 65535 is the lowest priority ID number. Step S20: Perform state three-mode redundancy design on the key signals of the user function circuit in the PL system. Based on copying the register storing the key signals three times, insert an error state detection circuit. When any register causes a data error due to a single-event flip, the error indication signal is pulled high and sent to the key signal encoding and characterization module. In step S30, the key signal encoding and characterization module in the PL system performs binary encoding on the key signal ID number that causes single-event flip and sends the encoding result to the framing module. In step S40, the framing module in the PL system performs single-event upset monitoring on key signals in the user logic circuit. When a single-event upset occurs on a key signal, a data frame and a control frame are generated and written into an external DDR memory as the basis for the monitoring program in the PS system to make protection decisions. In step S50, the monitoring unit program running on the processor in the PS system periodically accesses the DDR memory to read control frames. If the control frame is updated, the data frame is read according to the control frame information. The error key signal ID number in the data frame is extracted, and the number of key signal errors is analyzed and calculated. When the number of errors is greater than the threshold, a protection decision is issued to realize single-event soft error recovery of user function circuits.

2. The protection method based on single-event error state coding representation of key signals according to claim 1, characterized in that: The data frame and control frame generation process in step S40 includes the following steps: In step S401, under the control of the internal data frame timestamp counter, the framing module in the PL system generates a data frame and writes it to the external DDR memory whenever the counter increments by 1 if the error indication signal of the critical signal STMR circuit is detected to be high. The error critical signal ID number code and the error occurrence time information are recorded. When the data frame timestamp counter is full, it indicates that one round of critical signal detection is completed. After the counter is cleared, it starts counting again and continues to perform error status detection of critical signals. Step 402: Whenever a data frame is generated, the framing module in the PL system increments the error_cnt in the control frame by 1. When the data frame timestamp counter is full, the control frame timestamp counter is incremented by 1. The frame is then encoded according to the control frame format and written to the DDR memory. When the control frame timestamp counter is full, it is cleared and the counting starts again.

3. The protection method based on single-event error state coding representation of key signals according to claim 1, characterized in that: In step S50, the process of the monitoring unit reading control frames and data frames, performing data parsing, and generating protection decisions includes the following steps: Step S501: The monitoring unit program in the PS system periodically reads the control frame; first, it determines whether the timestamp in the control frame has been updated. If the timestamp has been updated, it means that a round of key signal detection has been completed and data frame extraction and parsing are required; otherwise, there is no need to read the data frame. In step S502, the monitoring unit program in the PS system extracts the error_cnt value from the control frame and determines whether the error_cnt signal is 0. If the signal is 0, it means that there is no critical signal error at present, and returns to step S501 to continue execution; if the signal is not 0, it means that a critical signal error has occurred, and jumps to step S503 to continue execution. Step S503: The monitoring unit program in the PS system reads error_cnt data frames from the DDR memory using error_cnt as the offset address, and extracts the time information and error key signal ID number from the data frames; Step S504: The monitoring unit program in the PS system analyzes the error key signal ID number in the data frame and removes duplicate ID numbers. That is, when the error key signal ID number is the same in multiple data frames, the error is counted only once. Then the number of error key signals is summed to obtain the total number of errors. Step S505: The monitoring unit program in the PS system determines whether the total number of errors is greater than the protection decision threshold. If not, it jumps to step S501 to continue execution; if so, it jumps to step S506 to continue execution. Step S506: If the number of errors exceeds the protection decision trigger threshold, it indicates that there are many critical signal errors, which are very likely to affect the normal function of the user functional circuit. Single-event soft error recovery needs to be performed immediately. The monitoring unit program in the PS system forms a protection decision and issues it, thereby realizing single-event soft error protection for the user functional circuit in the PL system.

4. The protection method based on single-event error state coding representation of key signals according to claim 1, characterized in that: In step S20, the output logic unit of the key signal is first copied three times. Then, the output of each module is input into the majority voter and the error state detector respectively. The operation of the majority voter is the same as that of the classic three-module redundancy, that is, the minority obeys the majority principle. For the error state detector, its default output is 0, indicating that no error has occurred. When any one of the three modules is different from the other two modules, the error state detector outputs 1, indicating that the key signal has been flipped, and drives the subsequent error state encoding module to perform the encoding operation. The error state indication signal of each key signal has a bit width of 1 bit, where '1' represents an abnormal state and '0' represents a normal state, and a maximum of 65536 key signals are supported.

5. The protection method based on single-event error state coding representation of key signals according to claim 1, characterized in that: In step S30, the key signal ID number encoding is to map the ID number from a decimal number to a binary number to reduce the data capacity for representing and transmitting a large number of key signal error states. This is used for subsequent frame generation of data frames. The ID number encoding information is a 16-bit binary number, and the supported ID number range is 0 to 65535.

6. The protection method based on single-event error state coding representation of key signals according to claim 1, characterized in that: In step S40, the framing and encoding module detects and frames the error states of key signals and user-defined signals, defining two types of frame formats: data frames and control frames. The data frame records the key signal ID number and the time information of the error occurrence within one error detection cycle. The data frame format is shown in Table 1. Table 1 Data Frame Structure The header of the data frame is 4'b1010, where b represents binary, the 4 before b represents 4 binary numbers, and 1010 represents the specific content of the binary number. The data frame timestamp counter has a total of 28 bits and is used to record the time when the error occurs. It is driven by the internal clock edge of the heterogeneous SoC FPGA PL system. The start time of counting is synchronized with the start time of the user function circuit. After the data frame timestamp is full, it is cleared to zero and the counting starts again. The data frame timestamp starts counting from 0 until it is full, which is one round of key signal detection cycle. The key signal ID encoding information consists of 16 bits and supports ID numbers from 0 to 65535. It is a binary encoding of the key signal ID number that has erred. If multiple key signals are erroneous at the same time, only the key signal ID number with the highest weight is encoded. The reserved bits are 8 bits, all of which are filled with 1'b0; The CRC check value is 8 bits, which is a cyclic redundancy check code that corrects one error and checks two errors. The check range starts from the frame header and ends at the reserved bits. The control frame uses the key signal detection cycle of the data frame as the time counting unit and periodically feeds back the program data information to the monitoring unit. This data information records the cycle counting timestamp, the error count value of the key signal in the data frame, and the CRC check value. The control frame format is shown in Table 2. Table 2 Control Frame Structure The frame header of the control frame is 4'b0101, where b represents binary, the 4 before b represents 4 binary numbers, and 0101 represents the specific content of the binary number. The control frame timestamp is 28 bits. When the timestamp of the data frame is full, the control frame timestamp count is incremented by 1'b1. After the control frame timestamp is full, it is reset to zero and the counting starts again. The error_cnt is the data frame count. Each time a data frame is generated, error_cnt is incremented by 1. error_cnt-1 is the address offset by which the monitoring unit controls the reading of data frames from the DDR memory. The CRC check value is 8 bits, which is a cyclic redundancy check code that corrects one error and checks two errors. The check range starts from the frame header and ends at error_cnt.

7. A protection system based on single-event error state coding characterization of key signals using the method of claim 1, characterized in that: The protection system based on the key signal single-event error state coding representation includes: a PL system, a PS system, and an external DDR memory; The PL and PS systems are connected to an external DDR memory. They are responsible for designing key signal STMRs in user function circuits, marking key signal IDs, and encoding key signal IDs for error states. They are also responsible for assembling data frames and control frames and writing them to the external DDR memory. The external DDR memory is connected to the PL and PS systems and is responsible for storing data frames and control frames. The data frame storage space address range is 0 to 65535, and a maximum of 65536 frames can be stored. The control frame storage space address is 65536, and only one frame is stored. The PS system and PL system are connected to the external DDR memory. The PS system is responsible for reading control frames and data frames from the external DDR memory, extracting the error status key signal ID number information and user-defined signal status information from the data frames, performing deduplication processing, counting the number of errors, comparing it with the protection decision threshold, and finally generating a protection decision to complete the single-event soft error recovery of the PL system user function circuit. The PL system includes one or more key signal STMR modules; a key signal encoding and characterization module; a custom signal statistics module; and a framing module. One or more critical signal STMR modules perform triple mode redundancy on critical signals in user function circuits and insert error detectors. If a single-event soft error occurs in any of the three circuits, the error detector will pull up the output error status indication signal and send it to the critical signal encoding and characterization module. Up to 65,536 critical signal STMR modules can be inserted to perform triple mode redundancy and error detection on 65,536 critical signals. A key signal encoding and characterization module receives the error state indication signal output by the key signal STMR module, marks the key signals with ID numbers according to their functional weights, performs binary encoding of the ID numbers for the error state key signals, and transmits the encoding results to the framing module. Key signals in normal state are not encoded with ID numbers. A custom signal statistics module detects whether errors have occurred in user-defined signals and transmits the results to the framing module. Under the control of the data frame timestamp counter, a framing module completes one data frame assembly and writes it to the DDR memory every time the timestamp counter increments by 1. The data frame timestamp counter starts counting from 0 and is full when it reaches 65535, indicating that one critical signal error state detection cycle has been completed. At this time, the control frame timestamp counter increments by 1 and generates a control frame, which is written to the DDR memory. After both the data frame and control frame timestamp counters are full, they are reset to zero and the counting starts again.

8. A protection system utilizing the single-event error state coding characterization of key signals as described in claim 7, characterized in that: The PS system includes one or more processors; and a monitoring unit; The processor monitoring unit is configured to be executed by the one or more processors; The monitoring unit program periodically reads control frames from the DDR memory and reads data frames from the DDR memory based on the error_cnt in the control frame as the address offset. The monitoring unit program counts the number of critical signal errors in the data frames and performs deduplication processing on critical signals with the same ID number in multiple data frames to avoid counting a continuous error critical signal as multiple errors. Then, the obtained number of error status critical signals is compared with a user-preset threshold. When the number of errors exceeds the threshold, it indicates that there are too many critical signals for handling error status, and the possibility of abnormal function of the user's functional circuit is extremely high. Protective measures should be taken. At this time, the monitoring unit generates and issues protection decisions, which include resetting the PL system, reloading the PL system, and reloading both the PL and PS systems simultaneously.

Citation Information

Patent Citations

  • Single event effect protection system and method for digital signal processing platform architecture

    CN105068969A

  • SEU (single event upset)-resistant fast refreshing circuit and method applied to FPGA (field programmable gate array) and based on ECCs (error correcting codes)

    CN106293991A