Low-cost anti-SEU high-reliability circuit design automation method

Through algorithm-level integrated EDA tools combined with spatial and temporal redundancy design, dynamically allocate ECC, DMR or TMR logic to generate self-detection and self-repair circuits, solving the problem of high-reliability and low-cost circuit design, and achieving low-cost, high-reliability and efficient SEU error repair.

CN120387405APending Publication Date: 2025-07-29ENLING INTELLIGENT TECHNOLOGY (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326702.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In existing integrated circuit designs, the cost of reinforcement technology against single-particle flip (SEU) is high, resulting in increased chip area and equipment costs, making it difficult to achieve high reliability under low-cost conditions.

Method used

The algorithm-level integrated EDA tool is used to dynamically allocate ECC, DMR or TMR logic by combining spatial redundancy and temporal redundancy detection and repair design to generate RTL-level circuit design codes with self-detection and self-repair functions.

Benefits of technology

The circuit resource consumption was significantly reduced, the SEU error rate was reduced to below 1%, and the resource consumption was only below 40% of the traditional TMR design, while 100% CRAM error repair and the FF error rate was reduced from 54.5% to 11.1%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387405A_ABST
    Figure CN120387405A_ABST
Patent Text Reader

Abstract

The invention discloses a low-cost anti-SEU (single event upset) high-reliability circuit design automation method. The method comprises the following steps that a circuit design program described by a high-level language is compiled into an intermediate data flow diagram through an algorithm-level comprehensive EDA tool; in nodes of the CDFG, detection and repair design combining space redundancy and time redundancy is adopted for logic nodes and data nodes; eCC (error correction code), DMR (dual modular redundancy) or TMR (triple modular redundancy) detection logic is dynamically allocated according to node types, and RTL-level circuit design codes with self-detection and self-repair functions are generated. The method has the beneficial effects that the algorithm-level reinforcement EDA technology in the technical scheme can reduce the area cost of the TMR reinforcement technology by 2.3 times, and meanwhile, higher anti-SEU reliability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to EDA technology for high-reliability circuit design, and particularly to a low-cost automated method for high-reliability circuit design against SEU. Background Art

[0002] In the future, space activities featuring informatization, networking, and mission diversification have high requirements for the working ability, reliability, and security of space electronic systems. Promoting the space application of nanoscale high-performance very large-scale integrated circuits is an inevitable choice. High-performance chips such as FPGAs, ASICs, and SoCs have outstanding computing capabilities and play a crucial role in promoting the processes of informatization and diversified space activities. However, the reliability of complex circuit systems in a radiation environment is affected by the characteristics of the process itself. With the development of space technology and the requirements of key technologies in various application fields, the reliability expectation of the core chips in circuit systems is continuously increasing. The demand for single-event protection research on such devices is driven by the need for soft-error protection capabilities in core technology fields such as space missions, high-energy physics experiments, and satellite communications in China. This makes the ability of circuit systems to resist soft errors an important link in the chip design process in addition to core parameters such as area and performance. (Here, soft errors refer to logical or data errors in the chip but do not involve the chip itself. Once the data is repaired, the chip can work normally.) Optimizing multi-dimensional characteristic quantities such as area, speed, and soft-error resistance ability to quickly and accurately construct a circuit system has important theoretical and practical application values and is also an important core technology in the field of very large-scale integrated circuit design. An important indicator of the soft-error resistance ability is the hardening ability against single-event upset (SEU). Single-event upset SEU refers to the situation where there is only one logical error at any given moment. SEU is the main circuit error form in space and ground irradiation environments. The anti-SEU reliability is a major performance indicator for high-reliability circuit design. What the present invention proposes is a low-cost automated method for high-reliability circuit design against SEU.

[0003] Compared with ordinary chip design, the hardening technologies at the traditional circuit design level and layout technology level have a design and production cost more than 10 times higher, and the hardening effect decreases as the working voltage of the integrated circuit decreases. For the circuit-level triple modular redundancy (TMR) hardening design, the cost also reaches 3 to 5 times. With the development of satellite communications, autonomous driving, and high-energy radiation devices, high-reliability and radiation-hardened design tools with excellent performance have become a key technology. The development of the theory and practical technology of this technology can effectively promote the application of high-end chips in key technology fields such as space, military, and nuclear industry in China, meet the major national needs, and bring significant economic benefits.

[0004] Currently, there is an urgent need for highly reliable radiation-resistant designs in aviation, aerospace, automotive, and other fields. Combining algorithm-level synthesis EDA tools to implement radiation-resistant self-detection and self-repair designs during circuit synthesis allows for rapid and accurate low-cost, high-reliability designs, which has important theoretical and practical application value.

[0005] Existing integrated circuit layout hardening technologies and circuit-level TMR hardening technologies both utilize spatial redundancy. These technologies employ redundant transistors or logic gates and arbitration circuits to detect SEU flips in logic gates and ensure output accuracy. However, their drawback is a significant increase in chip area costs, which in turn increases the cost of radiation-hardened electronic devices. Summary of the invention

[0006] The main technical problem solved by the present invention is to provide a low-cost SEU-resistant and high-reliability circuit design automation method to solve one or more of the above-mentioned existing technical problems.

[0007] To solve the above technical problems, the present invention adopts a technical solution: a low-cost SEU-resistant high-reliability circuit design automation method, the innovation of which is that it includes the following steps:

[0008] (1) Compile the circuit design program described in a high-level language into an intermediate data flow graph (CDFG) through an algorithm-level synthesis EDA tool;

[0009] (2) In the nodes of the CDFG, a detection and repair design combining spatial redundancy and temporal redundancy is adopted for logical nodes and data nodes respectively;

[0010] (3) Dynamically allocate ECC, DMR or TMR detection logic according to the node type and generate RTL-level circuit design code with self-detection and self-repair functions.

[0011] In some embodiments, the spatial redundancy design includes:

[0012] Use ECC coding detection on data nodes, including parity check code or Hamming code;

[0013] DMR or TMR redundancy design is adopted for logical nodes, and error detection and correction are achieved through majority signal selectors or XOR logic gates.

[0014] In some embodiments, the temporal redundancy design includes:

[0015] Insert redundant calculation steps during the scheduling phase to fix transient logic errors through repeated calculations;

[0016] When a data error in a storage unit is detected, a system restart signal is triggered to refresh the external storage data.

[0017] In some embodiments, the process of the algorithm-level synthesis EDA tool includes: inputting an algorithm-level program and generating a CDFG; scheduling and resource binding the CDFG nodes, and inserting redundant detection logic; outputting RTL-level Verilog or VHDL code with timing constraints.

[0018] In some embodiments, the specific TMR design is as follows:

[0019] Duplicate the logic nodes three times, and output to access the majority signal selector;

[0020] When a single-event upset causes an error in one of the logic gates, repair is performed by selecting the logic gate signals of two correct outputs.

[0021] In some embodiments, the detection and repair design can handle SEU errors and MBU errors simultaneously, specifically including:

[0022] Perform DMR detection on the logic nodes, and trigger error marking by comparing the output differences of two redundant computing units;

[0023] Implement multi-bit error detection and repair for the memory nodes by using ECC encoding.

[0024] In some embodiments, the technology is verified through ground irradiation experiments, specifically including:

[0025] Define a target area in the FPGA or ASIC, and inject CRAM (programmable memory) or FF faults;

[0026] Trigger the error repair mechanism through self-detection logic, and compare the failure rate and resource consumption before and after reinforcement;

[0027] The results show that the circuit resource consumption after reinforcement is less than 40% of the traditional TMR design, and the SEU error rate is reduced to less than 1%.

[0028] The beneficial effects of the present invention are: the repair rate of the EDA reinforcement technology for the CRAM error of this technical solution can reach 100%, and the FF error rate is reduced from 54.5% to 11.1%; and through the algorithm-level dynamic repair mechanism, even if MBU (multiple-bit flip) occurs, it can still be repaired through the ECC and redundant computing parts. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:

[0030] Figure 1 It is a schematic diagram of the connection between the triple modular redundancy logic unit and the majority selector of a low-cost anti-SEU high-reliability circuit design automation method of the present invention.

[0031] Figure 2 It is the dual modular redundancy logic unit and the exclusive OR detection circuit of a low-cost anti-SEU high-reliability circuit design automation method of the present invention.

[0032] Figure 3 It is the comprehensive flowchart from algorithm input to RTL generation of a low-cost anti-SEU high-reliability circuit design automation method of the present invention.

[0033] Figure 4 It is the comparison data of traditional RTL and HLS designs of a low-cost anti-SEU high-reliability circuit design automation method of the present invention in terms of resources and error rate. Specific implementation manners

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] As Figures 1 to 4 shown, the embodiments of the present invention include:

[0036] The implementation process of this technology is divided into the following core steps:

[0037] 1.1 Algorithm-level synthesis EDA tool compilation and intermediate data flow graph (CDFG) generation

[0038] Input processing: The user writes a circuit function algorithm program in a high-level language (such as C / C++), such as matrix multiplication (GEMM) or vector adder.

[0039] Compilation and CDFG generation: The algorithm program is compiled into an intermediate data flow graph (CDFG) through an algorithm-level synthesis EDA tool (such as a high-level synthesis tool HLS). The CDFG consists of nodes and edges. The nodes represent computing instructions (such as addition, multiplication, logical operations), and the edges represent data and control dependencies.

[0040] Node classification: The CDFG nodes are divided into logic nodes (performing Boolean operations or arithmetic operations) and data nodes (storage units or registers).

[0041] 1.2 Detection and repair design combining spatial redundancy and temporal redundancy

[0042] Spatial redundancy design:

[0043] Data node: Adopt ECC (Error Correction Code) for detection and repair.

[0044] Parity check code: Add one parity bit to the data bit width, and generate a parity value through exclusive OR operation (for example: where represents exclusive OR calculation), and detect errors through the parity bit during reading (if chk = 1, trigger repair). Hamming code: Support single-bit error detection and repair, and double-bit error detection.

[0045] Logic node: Adopt ECC, DMR (Dual Modular Redundancy) or TMR (Triple Modular Redundancy) design.

[0046] ECC: For a logic node f = {f0, f1,.., f

[0047] } with a bit width of k, the ECC scheme is to generate a new detection logic function k-1 and add one detection output bit: and add one detection output bit:

[0048]

[0049] Similarly, when chk is 1, it indicates that the node has a logic error.

[0050] TMR: Duplicate the logic node three times, and connect the output to a majority signal selector (see Figure 1 ). When a single logic gate has an error due to SEU, the selector repairs the result based on two correct outputs.

[0051] DMR: Duplicate the logic node twice and add an exclusive OR logic gate:

[0052]

[0053] When chk is 1, the logic node has a logic error.

[0054] For the calculation node, we adopt DMR detection logic. Assume u = f(x) is a calculation node for an input x, and f(x) is the calculation function. We duplicate u twice and add a comparator to generate the detection logic:

[0055] chk = f0(x) ≠ f1(x).

[0056] When chk is 1, the calculation node has a logic error. (See the DMR circuit design in Figure 2 )

[0057] The data node detection DMR scheme can not only detect SEU errors, but also detect most MBU (multiple-bit logic flip) errors, thus achieving higher reliability than existing circuit-level or layout-level TMR.

[0058] Time redundancy design:

[0059] Redundant calculation insertion: Insert duplicate calculation steps during the scheduling phase. For example, repeatedly execute critical logic nodes over multiple clock cycles and eliminate transient errors by comparing the result consistency. Storage unit repair: When the ECC detects an error in the storage unit, generate a system restart signal and reload the correct data through an external FPGA refresh circuit.

[0060] 1.3 Dynamic detection logic allocation and RTL code generation

[0061] Resource binding: Dynamically allocate detection logic resources according to node types. For example: Bind the corresponding hardware units (such as majority selectors, exclusive-OR gates) of TMR or DMR to logic nodes; Bind ECC encoders / decoders to data nodes.

[0062] RTL generation: Output Verilog / VHDL code with self-detection and self-repair functions, including timing constraints, state machines (STG), and error handling modules (such as restart control logic).

[0063] 2. Process of algorithm-level synthesis EDA tools

[0064] 2.1 Input and CDFG generation

[0065] The user provides a C / C++ algorithm program, and the EDA tool generates a CDFG through syntax parsing.

[0066] Example: The GEMM matrix multiplication algorithm is generated by 10 lines of high-level code. Traditional RTL design requires 1000 lines of Verilog code, while the HLS tool can automatically optimize parallelism and resource allocation.

[0067] 2.2 Scheduling and redundant calculation insertion

[0068] Scheduling strategy: Optimize the instruction execution order according to data dependencies and insert redundant calculation steps. For example, insert duplicate calculations after critical path nodes and verify the results through a comparator.

[0069] Resource allocation: Allocate logic resources (such as adders, multipliers) to each node and allocate additional resources for detection logic (such as redundant units of TMR).

[0070] 2.2 RTL Output and Subsequent Processes: After generating Verilog / VHDL code, logic synthesis, placement and routing, and physical verification are performed, ultimately forming a programmable bitstream or an ASIC layout.

[0071] 3 Specific Implementations for Detection and Repair of Designs

[0072] 3.1 TMR Design

[0073] Implementation Steps: Duplicate the logic nodes three times (such as three identical adders); connect the outputs to a majority voter (see Figure 1 ); when a single event upset causes an error in one adder, the voter selects the two correct outputs as the final result.

[0074] 3.2 Combination of DMR and ECC

[0075] DMR for Logic Nodes: The outputs of two redundant computing units are compared through an XOR gate. If chk = 1, an error flag is triggered and time redundancy repair (such as recalculation) is initiated.

[0076] ECC for Memory: Use Hamming code to encode the storage cells, which can detect and repair single-bit errors and detect double-bit errors.

[0077] 4. Ground Irradiation Experiment Verification

[0078] 4.1 Experiment Setup and Data Tables

[0079] The experiment verifies the SEU resistance effect by comparing the performance of the standard design, TMR-reinforced design, and the algorithm-level reinforced design of this technology. The specific data is as follows:

[0080] Table 2: Fault Injection Analysis Results of Different Design Methods for Large-Scale Matrix Operation Circuits

[0081]

[0082] Table 3: Particle Irradiation Experiment Results of Large-Scale Matrix Operation Circuits

[0083]

[0084] 4.2 Experiment Analysis and Table Application

[0085] Failure Rate Comparison (Table 2): The algorithm-level reinforced design of this technology consumes only 1 - 2 times the LUT and FF resources of the standard design, far lower than the 3 - 5 times of the TMR design. At the same time, the failure rates are all controlled below 1%.

[0086] Key Conclusion: By organically combining spatial redundancy and time redundancy designs, this technology significantly reduces the hardware resource overhead while maintaining high reliability.

[0087] Irradiation experiment results (Table 3): In the standard design, 3 system-level faults (SEFIs) were triggered among 38,810 SEU events, while in this technology, there were only 440 SEUs and no SEFI occurred.

[0088] Technical advantages: Combining the repair mechanisms of time redundancy and space redundancy, this technology has better fault tolerance for transient errors than traditional TMR designs.

[0089] 4.3 Experimental implementation details

[0090] Fault injection method:

[0091] CRAM error: Bit flips were injected into the configuration memory (CRAM) of the FPGA through the underlying ICAP interface to simulate the SEU effect.

[0092] FF error: Key flip-flops (FFs) were extracted from the synthesized netlist, and logic errors were simulated by dynamically tampering with signals.

[0093] Repair mechanism verification: When a CRAM error is detected, the system refresh signal is triggered to reload the correct code stream from the Flash.

[0094] FF errors are automatically repaired through redundant repeated calculations of time redundancy without interrupting the system operation.

[0095] 5. Verification of the economy and reliability of the technical solution

[0096] 5.1 Comparison of resource cost and error rate (combining Table 2 and Table 3)

[0097] Economy: The resource consumption of this technology (×2.6) is only 43% of that of the TMR design (×6), significantly reducing the chip area and manufacturing cost.

[0098] Reliability: The SEU error rate is reduced from 2,418 times in the TMR design to 440 times, and system-level faults (SEFIs) are completely eliminated.

[0099] 5.2 Implementation case: Strengthening effect of the vector adder (VA)

[0100] Table 4: Fault injection results of the vector adder

[0101] Resource type Number of annotation errors Number of errors Error rate / % VA_CRAM 100 22 22 VA_EDA_CRAM 100 0 0 VA_FF 22 12 54.5 VA_EDA_FF 27 3 11.1

[0102] Key conclusions:

[0103] The repair rate of the algorithm-level EDA strengthening technology for CRAM errors can reach 100%, and the FF error rate is reduced from 54.5% to 11.1%.

[0104] Through the algorithm-level dynamic repair mechanism, even if MBU (Multi-Bit Upset) occurs, it can still be repaired through ECC and redundant calculation parts.

[0105] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the content of the specification of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A low-cost anti-SEU highly reliable circuit design automation method, characterized in that: It includes the following steps: (1) Compile the circuit design program described in high-level language into an intermediate data flow graph (CDFG) through an algorithm-level synthesis EDA tool; (2) In the nodes of the CDFG, adopt a combined detection and repair design of spatial redundancy and temporal redundancy for logic nodes and data nodes respectively; (3) Dynamically allocate ECC, DMR or TMR detection logic according to the node type, and generate RTL-level circuit design code with self-detection and self-repair functions.

2. A low-cost anti-SEU high-reliability circuit design automation method according to claim 1, characterized in that: The spatial redundancy design includes: Adopt ECC encoding detection for data nodes, including parity check code or Hamming code; Adopt DMR or TMR redundancy design for logic nodes, and implement error detection and correction through a majority signal selector or an exclusive OR logic gate.

3. A low-cost anti-SEU highly reliable circuit design automation method according to claim 1, characterized in that: The temporal redundancy design includes: Insert redundant calculation steps in the scheduling stage to repair transient logic errors through repeated calculations; When a data error in the storage unit is detected, trigger a system restart signal to refresh the external stored data.

4. A low-cost anti-SEU highly reliable circuit design automation method according to claim 1, characterized in that: The process of the algorithm-level synthesis EDA tool includes: input the algorithm-level program and generate CDFG; schedule and resource bind the CDFG nodes, and insert redundant detection logic; output RTL-level Verilog or VHDL code with timing constraints.

5. A low-cost anti-SEU highly reliable circuit design automation method according to claim 1, characterized in that: The specific TMR design is: Copy the logic node three times, and output the access to the majority signal selector; When a single event upset causes an error in one of the logic gates, repair it by selecting the logic gate signals of the two correct outputs.

6. A low-cost anti-SEU highly reliable circuit design automation method according to claim 1, characterized in that: The detection and repair design can handle SEU errors and MBU errors simultaneously, specifically including: Adopt DMR detection for logic nodes, and trigger error marking by comparing the output differences of two redundant calculation units; Adopt ECC encoding for memory nodes to achieve multi-bit error detection and repair.

7. A low-cost anti-SEU highly reliable circuit design automation method according to claim 1, characterized in that: The technology is verified through ground irradiation experiments, specifically including: Define the target area in the FPGA or ASIC, and inject CRAM or FF faults; Trigger the error repair mechanism through self-detection logic, and compare the failure rate and resource consumption before and after reinforcement; The results show that the circuit resource consumption after reinforcement is less than 40% of the traditional TMR design, and the SEU error rate is reduced to less than 1%.