Protection method and system based on key signal single-particle error state coding representation

By adopting a protection method based on the encoding and characterization of single-particle error status in the user function circuit, the problem that traditional TMR technology cannot detect single-particle soft errors and excessive resource overhead is solved, and higher system reliability and resource efficiency are achieved.

CN120074747AActive Publication Date: 2025-05-30NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202411832487.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-30
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Traditional TMR technology can only implement fault-tolerant design and cannot promptly detect and feedback single-particle soft errors in multiple redundant circuits, resulting in abnormal user functions; at the same time, the transmission bandwidth and storage overhead caused by monitoring of a large number of key signal status in complex user function circuits is too large.

Method used

The protection method based on the characterization of single-particle error status encoding of key signal is adopted, and the critical signal STMR and error key signal ID number encoding is used to realize the characterization and storage of key signal error status using data frames and control frames, and the monitoring unit program in the PS system is used to complete the error status analysis and protection decisions.

Benefits of technology

Reduces the resource overhead and data transmission bandwidth of the user function circuit, improves the reliability of the user function circuit system based on SoC type FPGA, and avoids redundant failure caused by error accumulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074747A_ABST
    Figure CN120074747A_ABST
Patent Text Reader

Abstract

The invention provides a protection method and system based on key signal single particle error state coding characterization, which realizes key signal error state characterization and storage through key signal STMR and error key signal ID number coding and through data frames and control frames, and finally completes error state analysis by using a monitoring unit program in a PS system. And forming a protection decision. According to the invention, the resource overhead and the data transmission bandwidth of the user function circuit can be reduced, the reliability of the user function circuit system realized based on the SoC type FPGA is improved, and the method has applicability and operability for the digital circuit system of the electronic equipment for space navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology for enhancing resistance to space radiation effects, and particularly to a protection method based on encoding and characterization of error states of critical signals. Background Art

[0002] SRAM-based FPGA, namely a field programmable gate array based on Static Random Access Memory, contains a large number of configuration storage units inside. It can achieve specific functions required by users through programming, and has advantages such as high integration, strong parallel computing ability, short development cycle, and repeatable programming. It is widely used in spacecraft electronic equipment. With the continuous development of microelectronics technology, FPGA has evolved from a single programmable logic resource to a heterogeneous SoC-based FPGA integrating components with different architectures and processes. The heterogeneous SoC-based FPGA developed by Xilinx Corporation integrates a Programmable Logic (PL) system and a Processing System (PS) system inside, and some models even integrate radio frequency components. The PL is the programmable logic part, and users can implement specific hardware logic functions through programming according to their own needs. The PS is the processor system part, usually referring to one or more processor cores integrated inside the FPGA, such as the ARM Cortex-A series processors. The PS part provides a software programmable environment for users, allowing the operation of operating systems and application programs.

[0003] SRAM-based FPGA belongs to a single-event upset sensitive device. A large number of storage units inside it are extremely prone to single-event upset (SEU) in the space radiation environment, which is manifested as single-event soft error coupling transmission in the sequential logic circuit system, ultimately resulting in abnormal functions of the user functional circuit. When SRAM-based FPGA is applied in space, a single-event soft error protection technology combining Tripple Module Redundancy (TMR) and refreshing is usually adopted. The schematic diagram of the traditional TMR technology implementation is shown in Figure X. However, the TMR technology will bring a resource overhead of 3 to 4 times. In the case of the continuous increase in the design complexity of the user functional circuit, it is not feasible to perform TMR on the overall design. A feasible method is to select key modules and key signals in the design for partial TMR. Since the data in the registers and block memories in the user functional circuit cannot be restored correctly through the refreshing technology after SEU occurs, there is also a problem of failure caused by error accumulation in TMR. To avoid TMR failure, the critical signals in the user functional circuit not only need TMR, but also must be restored in time after a single-event soft error occurs to ensure that the redundant critical signal decision outputs the correct result.

[0004] Current research on the key signal protection method for the sequential logic circuit of SRAM-based FPGAs focuses on how to select key signals. To avoid transmission interference between signals, key signals are often transmitted through independent signal lines and device pins. In complex design scenarios, detecting the status of key signals consumes a large amount of interface resources or data transmission bandwidth, which not only increases the data storage overhead but also increases the risk of single-event soft errors in data information, ultimately affecting the reliability of detection and protection decisions based on key signals.

[0005] Since heterogeneous SoC-based FPGAs provide two systems, namely the PL and PS systems, the processor core in the PS system can be used to detect and recover single-event flips of key signals in the user functional circuit of the PL system, improving the single-event soft error resistance of the user functional circuit without increasing the system hardware overhead. Summary of the Invention

[0006] To overcome the deficiencies of the prior art, the present invention provides a protection method and system based on the encoding representation of the single-event error state of key signals. Aiming at the problems that the prior art mainly focuses on the analysis and identification of key signals in the user functional circuit of FPGAs, without considering the excessive transmission bandwidth and storage overhead caused by the monitoring of the states of a large number of key signals in complex user functional circuits, and that the traditional TMR technology can only achieve fault-tolerant design and cannot detect and feedback single-event soft errors in multiple redundant circuits in a timely manner, ultimately leading to abnormal user functions due to the accumulation of single-event soft errors, the present invention provides a protection system and method based on the encoding representation of the single-event error state of key signals. By encoding the key signal STMR and the ID number of the error key signal, and using data frames and control frames to realize the representation and storage of the key signal error state, finally, the monitoring unit program in the PS system is used to complete the error state analysis and form a protection decision. The present invention can reduce the resource overhead and data transmission bandwidth of the user functional circuit, and at the same time improve the reliability of the user functional circuit system implemented based on SoC-based FPGAs, which has applicability and operability for the digital circuit systems of aerospace electronic equipment.

[0007] The technical problems solved by the present invention mainly include: 1) The traditional TMR can only tolerate faults and cannot represent the error state, resulting in the problem that the key signals in the user functional circuit fail due to error accumulation; 2) When a large number of key signals in complex logic circuits are in error, how to accurately detect the operating state and comprehensively represent the functional state of the FPGA circuit system; 3) By performing information fusion-based characterization encoding on a large number of key signals, the problem of excessive overhead of signal transmission interfaces and storage resources is solved.

[0008] The technical solution adopted by the present invention to solve its technical problems is:

[0009] A protection method based on the encoding and characterization of single - particle error states of key signals, comprising the following steps:

[0010] Step S10, calibrate the ID numbers of the key signals of the user function circuit in the PL system in descending order of function weights. The ID numbers correspond to the function weights from high to low in ascending order, supporting up to 65,536 key signals at most. 0 is the ID number with the highest priority, and 65,535 is the ID number with the lowest priority;

[0011] Step S20, perform a state triple - module redundancy (STMR) design on the key signals of the user function circuit in the PL system. On the basis of replicating the registers storing the key signals three times, insert an error - state detection circuit. When any one of the registers has data errors due to single - particle flips, the error - indication signal is pulled high, and the error - indication signal is sent to the key - signal encoding and characterization module;

[0012] Step S30, the key - signal encoding and characterization module in the PL system performs binary encoding on the ID numbers of the key signals that have undergone single - particle flips, and sends the encoding result to the framing module;

[0013] Step S40, the framing module in the PL system monitors single - particle flips of the key signals in the user logic circuit. When a single - particle flip occurs in a key signal, data frames and control frames are generated, and the data frames and control frames are written into an external DDR memory, serving as the basis for the protection decision - making of the monitoring program in the PS system;

[0014] Step S50, the monitoring unit program running on the processor in the PS system periodically accesses the DDR memory to read the control frames. If the control frames are updated, the data frames are read according to the control - frame information, the ID numbers of the error - key signals in the data frames are extracted, the number of key - signal errors is analyzed and calculated. When the number of errors is greater than the threshold, a protection decision is issued to achieve the single - particle soft - error recovery of the user function circuit.

[0015] The generation process of the data frames and control frames in step S40 includes the following steps:

[0016] Step S401, under the control of the internal data - frame timestamp counter in the framing module of the PL system, whenever the counter is incremented by 1, if it is detected that the error - indication signal of the key - signal STMR circuit is pulled high, a data frame is generated and written into the external DDR memory, recording the ID - number encoding of the error - key signal and the error - occurrence time information. When the data - frame timestamp counter reaches its maximum value, it indicates that a round of key - signal detection is completed. After the counter is cleared, it starts counting again and continues to perform the error - state detection of the key signals;

[0017] Step 402: Whenever the framing module in the PL system generates a data frame, it increments the error_cnt in the control frame by 1. When the data frame timestamp count reaches its maximum value, it increments the control frame timestamp counter by 1, performs framing encoding according to the control frame format, and writes the control frame into the DDR memory. When the control frame timestamp counter reaches its maximum value, it is cleared and starts counting again.

[0018] In the step S50, the process of the monitoring unit reading the control frame and the data frame, performing data parsing, and generating a protection decision includes the following steps:

[0019] Step S501: The monitoring unit program in the PS system regularly reads the control frame; first, it determines whether the timestamp in the control frame is updated. If the timestamp is updated, it indicates that a round of key signal detection has been completed, and data frame extraction and parsing are required. Otherwise, there is no need to read the data frame.

[0020] Step S502: The monitoring unit program in the PS system extracts the error_cnt value from the control frame and determines whether the error_cnt signal is 0. If the signal is 0, it means that there is no error in the current key signal, and it returns to step S501 to continue execution; if the signal is non-zero, it means that a key signal error has occurred, and it jumps to step S503 to continue execution.

[0021] Step S503: The monitoring unit program in the PS system reads error_cnt data frames from the DDR memory using error_cnt as the offset address, and extracts the time information and the error key signal ID number in the data frames.

[0022] Step S504: The monitoring unit program in the PS system analyzes the error key signal ID numbers in the data frames and eliminates duplicate ID numbers, that is, when the error key signal ID numbers in multiple data frames are the same, only one error is counted, and then the number of error key signals is summed to obtain the total number of errors.

[0023] Step S505: The monitoring unit program in the PS system determines whether the total number of errors is greater than the protection decision threshold. If not, it jumps to step S501 to continue execution; if so, it jumps to step S506 to continue execution.

[0024] Step S506: If the number of errors is greater than the protection decision trigger threshold, it means that there are many key signal errors, which are very likely to affect the normal function of the user functional circuit. Immediate single-event soft error recovery is required. The monitoring unit program in the PS system forms a protection decision and issues it, thereby realizing the single-event soft error protection of the user functional circuit in the PL system.

[0025] More preferably, in the step S10, the key signals are marked with ID numbers according to the functional weights of the key signals in the user function circuit, and different ID numbers are assigned to the key signals. Up to 65,536 key signals are supported. 0 is the ID number with the highest priority, and 65,535 is the ID number with the lowest priority.

[0026] More preferably, in the step S20, the key signal STMR improves the circuit structure on the basis of traditional triple modular redundancy by adding an error status detector, aiming to overcome the drawback that the classical TMR cannot clearly indicate whether an error has occurred. In the specific implementation process, first, the output logic units of the key signals (usually represented as LUTs or registers in the netlist logic) are copied three times, and then the outputs of each module are respectively input into the majority voter and the error status detector. The operation process of the majority voter is the same as that of the classical triple modular redundancy, that is, the minority obeys the majority principle. For the error status detector, its default output is 0, indicating that no error has occurred. When any one of the three modules is different from the other two modules, the error status detector outputs 1, indicating that this key signal has flipped, and drives the subsequent error status encoding module to perform encoding operations. The error status indication signal bit width of each key signal is 1 bit, where '1' represents the abnormal state and '0' represents the normal state. Up to 65,536 key signals are supported.

[0027] More preferably, in the step S30, the key signal ID number encoding maps the ID number from a decimal number to a binary number to reduce the data capacity of the error status representation and transmission of a large number of key signals, and is used for subsequent frame generation to generate data frames. The ID number encoding information is a 16-bit binary number, and the supported ID number range is 0 to 65,535.

[0028] More preferably, in the step S40, the frame encoding module detects and frame-encodes the error status of the key signals and the user-defined signals, and defines two types of frame formats, namely data frames and control frames.

[0029] The data frame records the ID numbers of the key signals where errors occur and the time information of the error occurrence within one round of error detection cycle. The data frame format is shown in Table 1:

[0030] Table 1 Data Frame Structure

[0031] Frame header Timestamp Key signal ID coding information Reserved bit CRC check 4’b1010 28bit 16bit 8bit 8bit

[0032] The frame header of the data frame is 4’b1010, where b represents binary, 4 in front of b represents 4 binary numbers, and 1010 is the specific content of the binary number.

[0033] The data frame timestamp counter is 28 bits in total, used to record the time when an error occurs. It is driven by the internal clock delay of the heterogeneous SoC type FPGA PL system for counting, and the starting moment of counting is synchronized with the starting time of the user functional circuit. After the data frame timestamp is full, it is cleared and starts counting again. The data frame timestamp starts counting from 0 until it is full for one round of key signal detection cycle.

[0034] The key signal ID coding information is 16 bits in total, supporting ID numbers from 0 to 65535, which is the binary coding of the DI number of the key signal where an error occurs. If there are multiple key signal status errors at the same moment, only the ID number of the key signal with the highest weight is encoded.

[0035] The reserved bits are 8 bits, all filled with 1'b0.

[0036] The CRC check value is 8 bits, which is a cyclic redundancy check code for correcting one and detecting two. The check range starts from the frame header and ends at the reserved bits.

[0037] The control frame uses the key signal detection cycle of the data frame as the time counting unit, and cyclically feeds back the program data information to the monitoring unit. This data information records the cycle count timestamp, the count value of the key signal error occurrence in the data frame (error_cnt), and the CRC check value. The control frame format is shown in Table 2:

[0038] Table 2 Control Frame Structure

[0039] Frame header Timestamp Error_cnt CRC check 4’b0101 28bit 24bit 8bit

[0040] The frame header of the control frame is 4'b0101, where b represents binary, 4 in front of b represents 4 binary numbers, and 0101 is the specific content of the binary number.

[0041] The control frame timestamp is 28 bits. When the timestamp of the data frame is full, the count value of the control frame timestamp is incremented by 1'b1. After the control frame timestamp is full, it is cleared and starts counting again.

[0042] The error_cnt is the data frame count. Each time a data frame is generated, error_cnt is incremented by 1. error_cnt - 1 is the address offset for the monitoring unit to control reading the data frame from the DDR memory.

[0043] The CRC check value is 8 bits, which is a cyclic redundancy check code for correcting one and detecting two. The check range starts from the frame header and ends at error_cnt.

[0044] The present invention also provides a protection system based on the encoding representation of the single-event error state of key signals, which includes: a PL system, a PS system, and an external DDR memory.

[0045] The PL system, the PS system are connected to an external DDR memory, and are responsible for implementing the design of the key signal STMR in the user function circuit, marking the ID number of the key signal, and encoding the ID number of the key signal of the error state. It is also responsible for framing data frames and control frames and writing them into the external DDR memory.

[0046] The external DDR memory is connected to the PL system and the PS system, and is responsible for storing data frames and control frames. The address range of the data frame storage space is 0 to 65535, and up to 65536 frames can be stored; the address of the control frame storage space is 65536, and only one frame is stored.

[0047] The PS system, the PL system are connected to an external DDR memory, and are responsible for reading control frames and data frames from the external DDR memory, extracting the ID number information of the key signal of the error state and the status information of the user-defined signal from the data frame, performing deduplication processing, counting the number of errors, comparing with the protection decision threshold, and finally generating a protection decision to complete the single-event soft error recovery of the user function circuit of the PL system.

[0048] The PL system includes one or more key signal STMR modules; a key signal encoding and characterization module; a custom signal statistics module and a framing module.

[0049] One or more key signal STMR modules perform triple modular redundancy and insert error detectors for the key signals in the user function circuit. As long as a single-event soft error occurs in any of the three circuits, the error detector will pull up the output error status indication signal and transmit it to the key signal encoding and characterization module. Up to 65536 key signal STMR modules can be inserted to perform triple modular redundancy and error detection on 65536 key signals.

[0050] A key signal encoding and characterization module receives the error status indication signal output by the key signal STMR module, marks the ID number of the key signal according to the function weight, and performs binary encoding of the ID number of the key signal of the error state, and transmits the encoding result to the framing module. The key signals in the normal state are not encoded with ID numbers.

[0051] A custom signal statistics module detects whether an error occurs in the user-defined signal and transmits the result to the framing module.

[0052] A framing module, under the control of a data frame timestamp counter, completes a data frame framing and writes it into the DDR memory every time the timestamp counter increments by 1. The data frame timestamp counter starts counting from 0 and reaches its maximum value of 65535, indicating the completion of a critical signal error status detection cycle. At this time, the control frame timestamp counter increments by 1 and generates a control frame, which is then written into the DDR memory. After both the data frame and the control frame reach their maximum values, they are cleared and start counting again.

[0053] The PS system includes one or more processors; a monitoring unit program.

[0054] One or more processors: The monitoring unit program is configured to be executed by the one or more processors.

[0055] A monitoring unit program: Periodically reads the control frame from the DDR memory, and reads the data frame from the DDR memory based on the error_cnt in the control frame as the address offset. The monitoring unit program counts the number of critical signal errors in the data frame, de-duplicates the critical signals with the same ID number in multiple data frames to avoid counting a continuously erroneous critical signal as multiple errors, and then compares the obtained number of critical signal error states with a threshold preset by the user. When the number of errors is greater than the threshold, it indicates that the number of critical signals for processing the error state is excessive, and the user function circuit is very likely to malfunction. Protective measures should be taken. At this time, the monitoring unit generates and issues a protection decision. The protection decision includes resetting the PL system, reloading the PL system, and reloading both the PL and PS systems simultaneously.

[0056] The beneficial effects of the present invention are as follows:

[0057] Perform STMR on the critical signals in the user function circuit. On the basis of traditional triple modular redundancy, an error status detector is added. When a single-event soft error occurs in one of the modules, it can indicate the error status, which serves as the basis for the monitoring unit to make a protection decision, and can avoid the accumulation of errors and ultimately lead to the failure of the redundancy measures;

[0058] Encode the ID numbers of the critical signals according to the functional weights, and reflect the influence degree of the critical signals on the function of the user function circuit through the ID numbers. Convert the 65536 critical signal error status information into 16-bit binary encoded data, realizing the compression of big data information and reducing the logic resource overhead of the user function circuit; at the same time, utilize the on-chip high-speed shared bus interface of the heterogeneous SoC type FPGA to achieve high-speed data transmission, reduce the bandwidth occupancy rate, and improve the real-time performance of critical signal error status detection.

[0059] Identify whether the control frame is updated by using the control frame timestamp, and reduce the running time of the monitoring unit. Only when the control frame timestamp is different from the previously read one, it is necessary to extract the error_cnt information. In addition, error_cnt serves as the offset address for the monitoring unit to read the data frame, which can avoid the large time overhead of reading 65,536 data frames. The timestamp in the control frame and error_cnt reduce the time overhead for the monitoring unit to read the data frame and the control frame, and improve the real-time performance of the monitoring unit to obtain effective information and make decisions.

[0060] Utilize the architecture characteristics of the internal integration of the PL system and the PS system in the heterogeneous SoC type FPGA, and run the monitoring unit program on one or more processor cores in the PS system. Without adding additional devices, realize the detection of the working state of the user function circuit running in the PL system and make a protection decision in a timely manner.

[0061] The application object of the present invention is the design of single-event soft error detection and protection for user function circuits based on heterogeneous SoC type FPGAs. The design is highly targeted, its system architecture covers the design of mainstream aerospace digital circuit systems, and the method has a certain degree of generality and can be extended to the design of single-event soft error protection for FPGAs with traditional architectures. Description of the Drawings

[0062] Figure 1 It is a flowchart of the encoding representation and protection decision of the key signal error state provided by the present invention.

[0063] Figure 2 It is a schematic structural diagram of a single-event soft error protection system with the encoding representation of key signals provided by the present invention.

[0064] Figure 3 It is a schematic circuit diagram of the key signal STMR provided by the present invention.

[0065] Figure 4 It is a schematic diagram of the key signal ID number encoding provided by the present invention. Detailed Embodiment

[0066] The present invention will be further described below in conjunction with the drawings and embodiments.

[0067] The present invention provides a protection method based on the encoding representation of the single-event error state of key signals, as Figure 1 shown, the method includes the following steps:

[0068] Step S10, calibrate the ID numbers of the key signals of the user function circuit in the PL system in the order of decreasing function weights from high to low. The ID numbers correspond to decreasing function weights from small to large. Up to 65,536 key signals are supported. 0 is the ID number with the highest priority, and 65,535 is the ID number with the lowest priority.

[0069] Step S20, perform a state triple module redundancy (STMR) design on the key signals of the user function circuit in the PL system. On the basis of replicating the registers storing the key signals three times, insert an error status detection circuit. When any one of the registers has data errors due to single-event upsets, the error indication signal is pulled high, and the error indication signal is sent to the key signal encoding and characterization module.

[0070] Step S30, the key signal encoding and characterization module in the PL system performs binary encoding on the ID numbers of the key signals that have undergone single-event upsets, and sends the encoding result to the framing module.

[0071] Step S40, the framing module in the PL system monitors single-event upsets of the key signals in the user logic circuit. When a single-event upset occurs to a key signal, a data frame and a control frame are generated, and the data frame and the control frame are written into the external DDR memory, serving as the basis for the protection decision-making of the monitoring program in the PS system.

[0072] Step S50, the monitoring unit program running on the processor in the PS system periodically accesses the DDR memory to read the control frame. If the control frame has been updated, the data frame is read according to the control frame information, the ID numbers of the error key signals in the data frame are extracted, the number of key signal errors is analyzed and calculated. When the number of errors is greater than the threshold, a protection decision is issued to achieve single-event soft error recovery of the user function circuit.

[0073] The generation process of the data frame and the control frame in step S40 includes the following steps:

[0074] Step S401, under the control of the internal data frame timestamp counter in the framing module of the PL system, whenever the counter is incremented by 1, if it is detected that the error indication signal of the key signal STMR circuit is pulled high, a data frame is generated and written into the external DDR memory, recording the ID number encoding of the error key signal and the information of the error occurrence time. When the data frame timestamp counter reaches its maximum value, it indicates that a round of key signal detection is completed. After the counter is cleared, it starts counting again and continues to perform the error status detection of the key signals.

[0075] Step 402: Whenever the framing module in the PL system generates a data frame, it increments the error_cnt in the control frame by 1. When the data frame timestamp count reaches its maximum value, it increments the control frame timestamp counter by 1, performs framing encoding according to the control frame format, and writes the control frame into the DDR memory. When the control frame timestamp counter reaches its maximum value, it is cleared and starts counting again.

[0076] The monitoring unit in step S50 reads the control frame and the data frame, and performs the processes of data parsing and generating a protection decision, including the following steps:

[0077] Step S501: The monitoring unit program in the PS system periodically reads the control frame. First, it determines whether the timestamp in the control frame is updated. If the timestamp is updated, it indicates that a round of key signal detection has been completed, and data frame extraction and parsing are required; otherwise, there is no need to read the data frame.

[0078] Step S502: The monitoring unit program in the PS system extracts the error_cnt value from the control frame and determines whether the error_cnt signal is 0. If it is 0, it means that there is no error in the current key signal, and it returns to step S501 to continue execution; if it is non-zero, it means that a key signal error has occurred, and it jumps to step S503 to continue execution.

[0079] Step S503: The monitoring unit program in the PS system reads error_cnt data frames from the DDR memory using error_cnt as the offset address, and extracts the time information and the error key signal ID number in the data frames.

[0080] Step S504: The monitoring unit program in the PS system analyzes the error key signal ID numbers in the data frames and eliminates duplicate ID numbers, that is, when the error key signal ID numbers in multiple data frames are the same, the error is only counted once, and then the number of error key signals is summed to obtain the total number of errors.

[0081] Step S505: The monitoring unit program in the PS system determines whether the total number of errors is greater than the protection decision threshold. If not, it jumps to step S501 to continue execution; if so, it jumps to step S506 to continue execution.

[0082] Step S506: If the number of errors is greater than the protection decision trigger threshold, it means that there are many key signal errors, which are very likely to affect the normal function of the user functional circuit. Immediate single-event soft error recovery is required. The monitoring unit program in the PS system forms a protection decision and issues it, so as to achieve single-event soft error protection for the user functional circuit in the PL system.

[0083] In this embodiment, in step S20, the key signal STMR improves the circuit structure on the basis of traditional triple modular redundancy by adding an error status detector, aiming to overcome the drawback that classical TMR cannot clearly indicate whether an error has occurred. The circuit schematic diagram of STMR is as Figure 2 shown. In the specific implementation process, first, the logic units at the output end of the key signal (usually represented as LUT or register in the netlist logic) are copied three times, and then the outputs of each module are respectively input into the majority voter and the error status detector. The operation process of the majority voter is the same as that of classical triple modular redundancy, that is, the minority obeys the majority principle. For the error status detector, its default output is 0, indicating that no error occurs. When any one of the three modules is different from the other two modules, the error status detector outputs 1, indicating that this key signal has flipped, and drives the subsequent error status encoding module to perform encoding operations. The error status indication signal bit width of each key signal is 1 bit, where '1' represents an abnormal state and '0' represents a normal state. Up to 65,536 key signals are supported.

[0084] In this embodiment, in step S20, the ID number of the key signal is encoded in binary, and the encoding schematic diagram is as Figure 3 shown.

[0085] The present invention also provides a single-event soft error protection system based on key signal encoding representation. The system supports encoding and representing the key signals of the user function circuit in the PL circuit, and performs framing encoding on the key signals and user-defined signals to generate data frames and control frames. By analyzing the data frames and control frames, a protection strategy for the PL system is generated to achieve single-event soft error protection for the PL system. The corresponding system structure of the present invention is as Figure 4 shown, including: a PL system, an external DDR memory, and a PS system.

[0086] PL system: Connected to the PS system and the external DDR memory, responsible for implementing the key signal STMR design, key signal ID number marking, and error status key signal ID number encoding in the user function circuit, and also responsible for framing the data frames and control frames and writing them into the external DDR memory.

[0087] External DDR memory: Connected to the PL system and the PS system, responsible for storing data frames and control frames. The storage space address range of the data frames is 0 to 65,535, and up to 65,536 frames can be stored; the storage space address of the control frame is 65,536, and only one frame is stored.

[0088] PS system: Connected to the PL system and an external DDR memory, it is responsible for reading control frames and data frames from the external DDR memory, extracting error status key signal ID number information and user-defined signal status information from the data frames, performing deduplication processing, counting the number of errors, comparing with the protection decision threshold, and finally generating a protection decision to complete the single-event soft error recovery of the user function circuit of the PL system.

[0089] The above PL system includes: one or more critical signal STMR modules; a critical signal encoding and characterization module; a custom signal statistics module and a framing module.

[0090] One or more critical signal STMR modules: Perform triple modular redundancy and insert error detectors for critical signals in the user function circuit. As long as a single-event soft error occurs in any of the three circuits, the error detector will pull up the output error status indication signal and transmit it to the critical signal encoding and characterization module. Up to 65,536 critical signal STMR modules can be inserted to perform triple modular redundancy and error detection on 65,536 critical signals.

[0091] A critical signal encoding and characterization module: Receives the error status indication signal output by the critical signal STMR module, marks the ID number of the critical signal according to the function weight, and performs binary encoding of the ID number of the error status critical signal, and transmits the encoding result to the framing module. The normal status critical signal is not encoded with an ID number.

[0092] A custom signal statistics module: Responsible for detecting whether there is an error in the user-defined signal and transmitting the result to the framing module.

[0093] A framing module: Under the control of the data frame timestamp counter, every time the timestamp counter increments by 1, a data frame is framed and written to the DDR memory. The data frame timestamp counter starts counting from 0 and counts up to 65,535 to indicate the completion of a critical signal error status detection cycle. At this time, the control frame timestamp counter increments by 1 and generates a control frame, which is written to the DDR memory. After the data frame and the control frame are full, they are both cleared and start counting again.

[0094] The above PS system includes one or more processors; a monitoring unit program.

[0095] One or more processors: The monitoring unit program is configured to be executed by the one or more processors.

[0096] A monitoring unit program: periodically reads control frames from the DDR memory, and reads data frames from the DDR memory according to the error_cnt in the control frame as the address offset. The monitoring unit program counts the number of critical signal errors in the data frames, performs deduplication processing on the critical signals with the same ID number in multiple data frames to avoid counting a continuously incorrect critical signal as multiple errors, and then compares the obtained number of critical signals in the error state with the threshold preset by the user. When the number of errors is greater than the threshold, it indicates that the number of critical signals for processing the error state is excessive, and the possibility of abnormal function of the user functional circuit is extremely high. Protective measures should be taken. At this time, the monitoring unit generates and issues a protection decision. The protection decision includes resetting the PL system, reloading the PL system, and reloading both the PL and PS systems simultaneously.

Claims

1. A protection method based on key signal single-particle error state coding characterization, characterized in that The steps include: Step S10, ID number calibration is performed on the key signals of the user function circuit in the PL system in the order of function weight from high to low, and the ID number from small to large corresponds to the function weight from high to low, and a maximum of 65536 key signals are supported, 0 is the highest priority ID number, and 65535 is the lowest priority ID number; Step S20, a triple-module redundancy design is performed on the key signals of the user function circuit in the PL system. On the basis of duplicating three copies of the register storing the key signals, an error state detection circuit is inserted. When any of the registers causes a data error due to a single-particle upset, the error indication signal is pulled high, and the error indication signal is sent to the key signal encoding characterization module; Step S30, the key signal encoding characterization module in the PL system performs binary encoding on the key signal ID number where the single event upset occurs, and sends the encoding result to the framing module; Step S40, the framing module in the PL system performs single-particle upset monitoring on the key signals in the user logic circuit, generates data frames and control frames when a single-particle upset occurs in the key signal, and writes the data frames and control frames into the external DDR memory as the basis for the monitoring program in the PS system to make protection decisions; Step S50, the monitoring unit program run by the processor in the PS system periodically accesses the DDR memory to read the control frame. If the control frame is updated, the data frame is read according to the control frame information, the error key signal ID number in the data frame is extracted, and the number of key signal errors is analyzed and calculated. When the number of errors is greater than the threshold, a protection decision is issued to realize single-particle soft error recovery of the user function circuit.

2. The protection method based on key signal single event error state coding characterization according to claim 1 is characterized in that: The process of generating the data frame and the control frame in step S40 includes the following steps: Step S401, the framing module in the PL system is under the control of the internal data frame timestamp counter. Whenever the counter is incremented by 1, if the error indication signal of the key signal STMR circuit is detected to be high, a data frame is generated and written into the external DDR memory, and the error key signal ID number code and the error occurrence time information are recorded. When the data frame timestamp counter is full, it indicates that a round of key signal detection is completed. After the counter is cleared, it starts counting again and continues to perform error state detection of the key signal. Step 402, whenever a data frame is generated by the framing module in the PL system, the error_cnt in the control frame is incremented by 1. When the data frame timestamp counter is full, the control frame timestamp counter is incremented by 1, the framing encoding is performed according to the control frame format, and the control frame is written to the DDR memory. When the control frame timestamp counter is full, it is cleared and restarted.

3. The protection method based on key signal single event error state coding characterization according to claim 1 is characterized in that: In step S50, the monitoring unit reads the control frame and the data frame, performs data analysis and generates a protection decision, including the following steps: Step S501, the monitoring unit program in the PS system periodically reads the control frame; first, it is determined whether the timestamp in the control frame is updated. The updated timestamp indicates that a round of key signal detection has been completed, and data frame extraction and parsing are required. Otherwise, there is no need to read the data frame; Step S502, the monitoring unit program in the PS system extracts the error_cnt value from the control frame and determines whether the error_cnt signal is 0. If the signal is 0, it means that there is no error in the key signal at present, and the program returns to step S501 to continue the execution; if the signal is not 0, it means that an error occurs in the key signal, and the program jumps to step S503 to continue the execution; Step S503: the monitoring unit program in the PS system uses error_cnt as the offset address to read error_cnt data frames from the DDR memory, and extracts the time information and error key signal ID number in the data frame; Step S504: the monitoring unit program in the PS system analyzes the error critical signal ID number in the data frame and removes the repeated ID number, that is, when the error critical signal ID numbers in multiple data frames are the same, only one error is counted, and then the number of error critical signals is summed to obtain the total number of errors; Step S505: The monitoring unit program in the PS system determines whether the total number of errors is greater than the protection decision threshold. If not, the program jumps to step S501 to continue execution; if yes, the program jumps to step S506 to continue execution; Step S506: If the number of errors is greater than the protection decision triggering threshold, it means that there are many key signal errors, which is very likely to affect the normal function of the user function circuit, and single-particle soft error recovery is required immediately. The monitoring unit program in the PS system forms a protection decision and sends it down, thereby realizing single-particle soft error protection for the user function circuit in the PL system.

4. The protection method based on key signal single event error state coding characterization according to claim 1 is characterized in that: In step S10, the key signals are marked with ID numbers according to their functional weights in the user function circuits. Different ID numbers are assigned to the key signals, supporting up to 65536 key signals, with 0 being the highest priority ID number and 65535 being the lowest priority ID number.

5. The protection method based on key signal single event error state coding characterization according to claim 1 is characterized in that: In the step S20, firstly, the output end logic unit of the key signal is copied three times, and then the output of each mode is input into the majority voter and the error state detector respectively, wherein the operation process of the majority voter is the same as the classic three-mode redundancy, that is, the minority obeys the majority principle. For the error state detector, its default output is 0, indicating that no error occurs. When any one of the three modes is different from the other two modes, the error state detector output is 1, indicating that the key signal is flipped, and drives the subsequent error state encoding module to perform the encoding operation. The error state indication signal bit width of each key signal is 1 bit, '1' represents an abnormal state, and '0' represents a normal state. Up to 65536 key signals are supported.

6. The protection method based on key signal single event error state coding characterization according to claim 1 is characterized in that: In step S30, the key signal ID number encoding is to map the ID number from a decimal number to a binary number to reduce the data capacity for characterizing and transmitting a large number of key signal error states, which is used for subsequent framing to generate data frames. The ID number encoding information is a 16-bit binary number, and the supported ID number range is 0 to 65535.

7. The protection method based on key signal single event error state coding characterization according to claim 1 is characterized in that: In step S40, the framing and encoding module detects and frames the error status of the key signal and the user-defined signal, and defines two types of frame formats, namely, data frame and control frame; The data frame records the key signal ID number and time information of the error in a round of error detection cycle. The data frame format is shown in Table 1: Table 1 Data frame structure The frame header of the data frame is 4'b1010, where b represents binary, the 4 before b represents 4 binary numbers, and 1010 is the specific content of the binary number; The data frame timestamp counter has a total of 28 bits, which is used to record the time when the error occurs. The counter is driven by the internal clock of the heterogeneous SoC FPGA PL system. The time when the counting starts is synchronized with the start time of the user function circuit. When the data frame timestamp is full, it is cleared and restarted. The data frame timestamp starts counting from 0 until it is full, which is a round of key signal detection cycle. The key signal ID coding information is 16 bits in total, supporting ID numbers from 0 to 65535, which is the binary code of the key signal DI number where an error occurs. If there are multiple key signal state errors at the same time, only the key signal ID number with the highest weight is encoded; The reserved bits are 8 bits, all filled with 1'b0; The CRC check value is 8 bits, which is a cyclic redundancy check code with one-check-two correction, and the check range starts from the frame header and ends at the reserved bit; The control frame uses the key signal detection cycle of the data frame as the time counting unit, and periodically feeds back the monitoring unit program data information. The data information records the cycle counting timestamp, the error count value of the key signal in the data frame, and the CRC check value. The control frame format is shown in Table 2: Table 2 Control frame structure The control frame header 4'b0101, where b represents binary, the 4 before b represents 4 binary numbers, and 0101 is the specific content of the binary number; The control frame timestamp is 28 bits. When the timestamp of the data frame is full, the control frame timestamp count value is increased by 1'b1. After the control frame timestamp is full, it is reset to zero and starts counting again. The error_cnt is a data frame count. Each time a data frame is generated, error_cnt is incremented by 1. error_cnt-1 is the address offset of the monitoring unit controlling the reading of the data frame from the DDR memory. The CRC check value is 8 bits, which is a one-check-two cyclic redundancy check code, and the check range starts from the frame header and ends at error_cnt.

8. A protection system based on key signal single event error state coding characterization using the method of claim 1, characterized in that: The protection system based on the key signal single particle error state coding characterization includes: a PL system, a PS system and an external DDR memory; The PL system and the PS system are connected to the external DDR memory, which is responsible for implementing the key signal STMR design, key signal ID number marking and error state key signal ID number coding in the user function circuit, and is also responsible for framing the data frame and control frame and writing them into the external DDR memory. The external DDR memory is connected to the PL system and the PS system, and is responsible for storing the data frame and the control frame. The data frame storage space address range is 0 to 65535, and a maximum of 65536 frames can be stored; the control frame storage space address is 65536, and only one frame is stored; The PS system and the PL system are connected to the external DDR memory and are responsible for reading the control frame and data frame from the external DDR memory, extracting the error status key signal ID number information and user-defined signal status information from the data frame, performing de-duplication processing, counting the number of errors, and comparing it with the protection decision threshold, and finally generating a protection decision to complete the single-particle soft error recovery of the user function circuit of the PL system; The PL system includes one or more key signal STMR modules; a key signal coding characterization module; a custom signal statistics module and a framing module; One or more key signal STMR modules perform triple-mode redundancy and insert error detectors on key signals in user function circuits. As long as a single-particle soft error occurs in any of the three circuits, the error detector will pull up the output error status indication signal and transmit it to the key signal coding characterization module. A maximum of 65,536 key signal STMR modules can be inserted to perform triple-mode redundancy and error detection on 65,536 key signals. A key signal encoding characterization module receives the error state indication signal output by the key signal STMR module, marks the key signal with an ID number according to the functional weight, performs ID number binary encoding on the error state key signal, and transmits the encoding result to the framing module. The normal state key signal does not perform ID number encoding; A custom signal statistics module detects whether errors occur in the user-defined signal and transmits the results to the framing module; A framing module is under the control of the data frame timestamp counter. Whenever the timestamp counter is increased by 1, a data frame is framed and written into the DDR memory. The data frame timestamp counter starts counting from 0 and is full when it reaches 65535, indicating that a critical signal error state detection cycle is completed. At this time, the control frame timestamp counter counts up by 1, and a control frame is generated and written into the DDR memory. When the data frame and control frame are full, they are cleared to zero and counted again.

9. A protection system using the key signal single event error state coding characterization according to claim 8, characterized in that: The PS system includes one or more processors; a monitoring unit program; A processor monitoring unit program is configured to be executed by the one or more processors; The monitoring unit program periodically reads control frames from the DDR memory, and reads data frames from the DDR memory based on the error_cnt in the control frame as the address offset. The monitoring unit program counts the number of key signal errors in the data frame, and deduplicates key signals with the same ID number in multiple data frames to avoid counting a key signal with a continuous error as multiple errors. The number of key signals in the error state obtained is then compared with the threshold value set in advance by the user. When the number of errors is greater than the threshold value, it means that the number of key signals for processing the error state is too large, and the possibility of malfunction of the user function circuit is very high, and protective measures should be taken. At this time, the monitoring unit generates and issues protective decisions, which include resetting the PL system, reloading the PL system, and reloading the PL and PS systems at the same time.

Citation Information

Patent Citations

  • Single event effect protection system and method for digital signal processing platform architecture

    CN105068969A

  • SEU (single event upset)-resistant fast refreshing circuit and method applied to FPGA (field programmable gate array) and based on ECCs (error correcting codes)

    CN106293991A

  • SRAM type FPGA single event upset error correction method and single event upset error correction circuit

    CN111459712A

  • Single event upset effect-oriented reliability method based on SoC chip

    CN114416436A

Cited By

  • Navigation enhancement signal transmitting method and device, equipment and storage medium

    CN121254301A