RTL-based fault injection and detection method
Through the RTL-based fault injection and detection method, the problems of low fault injection efficiency and slow detection speed in the existing technology are solved, flexible configuration and rapid detection of fault types are achieved, and customization of multiple fault types is supported, which improves the fault tolerance performance evaluation efficiency of spatial FPGA.
Patent Information
- Application Number
- CN202510815181.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing technology, fault injection tools for space FPGAs have low fault injection efficiency and slow detection speed, lack a universal system, and lack the ability to flexibly configure fault types and uniformly evaluate fault tolerance performance.
This paper provides an RTL-based fault injection and detection method. By configuring the fault injection tool, selecting the fault type and parameters, using extended Hamming code for encoding and decoding, generating syndromes and comparing them with preset codewords, outputting a fault detection report, and evaluating the fault-tolerance performance, the tool is packaged as a configurable IP core and integrated into an FPGA development board, supporting the customization of multiple fault types and parameters.
It significantly improves the efficiency of fault injection and detection, reduces resource overhead, enables flexible configuration and rapid detection of fault types, and supports fault tolerance performance evaluation on different FPGA platforms.
Smart Images

Figure CN120704928A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aerospace fault-tolerant computing, and in particular to an RTL-based fault injection and detection method. Background Art
[0002] With the rapid advancement of global informatization and space technology, space-based computing has become a vital component of the information infrastructure. FPGAs, with their highly parallel computing power and flexible, customizable architecture, are widely used in critical tasks such as data transmission and reception, protocol conversion, and image signal processing in spaceborne computers. However, the constant bombardment of spacecraft surfaces by a large number of high-energy particles and cosmic rays in the space environment can easily cause internal logic errors in devices. To realistically replicate the effects of this radiation, fault injection strategies are often employed in ground-based experiments. Single-event upsets, transient faults, or latch-up faults are simulated and injected into spacecraft chips to evaluate their radiation resistance and the effectiveness of their fault-tolerance schemes. Once a fault occurs, it must be detected and cleared as quickly as possible to ensure the stable operation of the spaceborne system.
[0003] Fault injection can be categorized into five levels: exposure testing, gate-level testing, register transfer level (RTL), microarchitecture, and software. While exposure testing offers the most realistic results, it is costly and time-consuming. The gate-level approach is based on physical models and lacks open-source gate circuit descriptions, making it difficult to scale. The microarchitecture level targets key modules, accurately locating them but making it difficult to trace the source of the fault. The software level offers low cost and high efficiency, but lacks realism. The RTL level strikes a balance between accuracy and controllability, offering both high injection efficiency and low cost, provided a complete RTL description is available. Currently, there are a limited number of RTL-level fault injection tools available on the market. Most support only a single fault type, such as SEU, and do not allow for user-defined key parameters such as injection location, number of bits, and duration. Existing fault detection methods often rely on system-level mechanisms (watchdog timers, heartbeats, and level monitoring), which are resource-intensive and can easily disrupt system operation. Furthermore, the lack of a universal system for evaluating fault tolerance across diverse FPGA platforms and application scenarios significantly hinders tool promotion and application. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] The present invention provides an RTL-based fault injection and detection method to solve the problems of low fault injection efficiency, slow fault detection speed and lack of a universal system in the prior art for fault tolerance performance evaluation of space FPGAs.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] The embodiment of the present invention provides a fault injection and detection method based on RTL, which includes:
[0008] Step S1: Configure the fault injection tool, select the fault type according to the RTL description and data flow form of the target model, and configure the data accuracy, injection bit number and duration parameters;
[0009] Step S2: creating a fault injection bypass in the RTL data flow of the target model, generating a fault injection location based on a random number file, and injecting a fault of a specified type and parameters into the intermediate data flow;
[0010] Step S3: Use extended Hamming code to encode and decode the data stream after the fault is injected, generate a syndrome and compare it with the preset codeword, and output a fault detection report;
[0011] Step S4: Based on the fault injection configuration and the detection report, the fault tolerance performance of the target model or FPGA is evaluated by error rate calculation.
[0012] As a preferred solution of the RTL-based fault injection and detection method described in the present invention, the fault types described in step S1 include single event upset SEU, single event transient SET and single event latch SEL, wherein:
[0013] The default duration of SEU is 2 clock cycles, corresponding to a bit flip;
[0014] The default duration of SET is 10 clock cycles, corresponding to bit reset;
[0015] The default duration of SEL is 100 clock cycles, corresponding to a set or reset operation.
[0016] As a preferred solution of the RTL-based fault injection and detection method of the present invention, the random number file is generated by:
[0017] Generate random numbers that meet the basic bit width and extended bit width constraints through Python scripts on the software side;
[0018] Generate non-repeating binary random numbers within each basic bit width and store them in text file format;
[0019] The random number file is input to the fault injection module through the parallel interface.
[0020] As a preferred solution of the RTL-based fault injection and detection method of the present invention, the extended Hamming code in step S3 adopts the (13,8) encoding rule, including:
[0021] The generation logic of data bits and check bits is based on XOR operation;
[0022] The receiving end determines the fault location and type by looking up the syndrome table;
[0023] A pipelined design is used to achieve parallel processing of checksum generation, fault injection, syndrome generation and fault output.
[0024] As a preferred solution of the RTL-based fault injection and detection method described in the present invention, the fault injection tool is encapsulated as a configurable IP core, and its interface includes:
[0025] Input port: clock signal, reset signal, fault type selection, data stream to be injected and random number file;
[0026] Output port: data flow after fault injection;
[0027] Configurable parameters include basic bit width, extended bit width, fault duration, and number of injected bits.
[0028] As a preferred solution of the RTL-based fault injection and detection method of the present invention, the fault tolerance performance evaluation in step S4 includes:
[0029] Compare the golden output without faults with the application output after fault injection and calculate the error rate;
[0030] Generate a fault tolerance performance evaluation report based on error rate, fault type, and data size.
[0031] As a preferred solution of the RTL-based fault injection and detection method described in the present invention, the fault-tolerant protection design for the convolutional layer of the convolutional neural network includes:
[0032] The weight and bias data are protected by error correction and detection code EDAC;
[0033] Adopt N-module redundancy strategy for multipliers, accumulators, adders and activation function modules;
[0034] Set up fault injection points before the redundant module and after the EDAC, and judge the fault tolerance effect through the output of the detection point.
[0035] As a preferred solution of the RTL-based fault injection and detection method described in the present invention, the method is integrated into a spatial FPGA fault tolerance performance evaluation system, including:
[0036] Host computer module: used to generate source data, random number files and analysis and evaluation reports;
[0037] FPGA module: includes fault injection tools, detection modules and the application to be tested;
[0038] Data transmission module: uses PCIe protocol to achieve high-speed data interaction between the host computer and FPGA.
[0039] As a preferred solution of the RTL-based fault injection and detection method described in the present invention, the evaluation process includes:
[0040] The first round generated trouble-free gold output;
[0041] The second round of fault injection and testing the validity of the preliminary results;
[0042] If it is invalid, perform a secondary check by comparing the golden output with the application output on the host computer.
[0043] As a preferred solution of the RTL-based fault injection and detection method of the present invention, the fault injection bypass is designed to not affect the normal execution of the trunk model, specifically by:
[0044] A timer controls a fault injection enable signal;
[0045] The bypass switches to the faulty data flow within a specified clock cycle, and maintains the original data flow in the remaining cycles.
[0046] The beneficial effects of the present invention are as follows: the present invention significantly improves the efficiency and practicality of fault tolerance performance evaluation for space FPGAs:
[0047] This invention is compatible with three single-event fault types: SEU, SET, and SEL. The number of injected bits, duration, and location can all be customized by the user. Based on random number file configuration, it solves the problems of single fault type and fixed parameters in existing tools. A timer-controlled bypass injection mechanism is used to switch to the faulty data stream within a specified clock cycle, ensuring that the main data stream is not affected and reducing the use of system resources.
[0048] A forward detection module is designed based on the (13,8) extended Hamming code. The receiving end quickly locates the fault type and location through a syndrome lookup table without the need to transmit the original data, effectively controlling the detection delay. A pipelined design is adopted in which check bit generation, fault injection, syndrome generation, and fault information output are performed in parallel, significantly optimizing hardware resource overhead.
[0049] The fault injection and detection tool is encapsulated as a configurable IP core that supports parameter adjustment such as basic bit width, extended bit width, and number of injected bits, and can be integrated into different models of FPGA development boards. It provides a PCIe high-speed interface and cooperates with the host computer to realize source data distribution, golden output comparison, error rate calculation and automatic report generation, and standardize the evaluation process and interface.
[0050] In the convolutional layer of the neural network, extended Hamming code (EDAC) is used to protect the weight and bias data, and triple modular redundancy (TMR) is deployed for the multiplier, accumulator, adder, and activation function module. The tool of the present invention is used to complete fault injection and detection, verifying the fault tolerance effect of the convolutional layer. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1 Schematic diagram of the flow of the RTL-based fault injection and detection method of the present invention.
[0053] Figure 2 This is an example diagram of the random number file content of the present invention.
[0054] Figure 3 This is a diagram illustrating the IP encapsulation and interface of the fault injection module of the present invention.
[0055] Figure 4 This is a simulation waveform diagram of the SEL of the present invention.
[0056] Figure 5 This is a simulation waveform diagram of the SET of the present invention.
[0057] Figure 6 This is a simulation waveform diagram of the SEU of the present invention.
[0058] Figure 7 This is a simulation waveform diagram without fault injection of the present invention.
[0059] Figure 8 This is a schematic diagram of the convolutional layer fault tolerance protection process of the present invention.
[0060] Figure 9 Schematic diagram of a general system architecture for fault-tolerance performance evaluation of the present invention.
[0061] Figure 10 Schematic diagram of the step-by-step fault-tolerance performance evaluation process of the present invention. DETAILED DESCRIPTION
[0062] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0063] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0064] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0065] Space-based computing, by deploying computing resources in space, enhances our ability to leverage space resources to solve problems. With the development of global informatization and space technology, space-based computing has gradually become a vital component of the global information infrastructure. FPGAs, commonly used chips for spaceborne computers, offer the advantages of high parallel computing and flexible customization. They can be used to implement functions such as data transmission and reception, protocol conversion, and image signal processing. High concentrations of high-energy particles and cosmic rays in space continuously bombard spacecraft surfaces, potentially causing errors in the internal operating logic of components. Fault-tolerance strategies are often employed to protect vulnerable modules. To realistically simulate the space radiation environment, fault injection strategies are often used on the ground to manually inject faults into spaceborne components. This allows for the evaluation of the components' inherent radiation resistance and the effectiveness of fault-tolerance strategies. Any failure in a spaceborne component requires rapid detection and timely correction to ensure the stable operation of the spaceborne computer.
[0066] The fault injection hierarchy, from bottom to top, includes exposure experiments, gate level, register transfer level (RTL), microarchitecture level, and software level. Although exposure experiments offer the most realistic results, they are costly and time-consuming. The gate level, based on physical models, lacks open-source gate descriptions, making it difficult to scale. The microarchitecture level targets key modules, accurately locating the source of the fault but making it difficult to trace the source. The software level is low-cost and efficient, but lacks realism. RTL-based fault injection offers relatively high injection efficiency and low cost, but requires a complete RTL description of the injection model. Existing RTL-based fault injection tools are still relatively rare, supporting limited fault types and lacking customization, making it difficult for users to adjust the tool's parameters to meet their specific needs. Existing fault detection modules primarily rely on system-level mechanisms such as watchdog timers, heartbeat detection, and level monitoring. These mechanisms are resource-intensive and can interfere with system operation. Developing a fast and efficient fault detection mechanism from a logical perspective is crucial. In addition, existing RTL fault injection and detection tools basically only propose a feasible method and lack a general system for evaluating FPGA fault tolerance performance, which limits their application scope.
[0067] When conducting RTL-based fault injection and detection research, it was found that some work only supported one fault type, single-event upset (SEU), without considering other fault types and lacking a clear definition of the fault effect. The present invention supplements the single-event latch (SEL) and single-event transient (SET) fault types, allowing users to customize the fault effect according to actual needs. At the same time, existing RTL-based fault injection and detection tools have a large resource overhead. The present invention designs a lightweight fault detection module based on error detection code to achieve rapid fault detection. Finally, existing work lacks a general system for FPGA fault tolerance performance evaluation and does not demonstrate the practical application of fault injection and detection tools. The present invention uses the convolutional layer of a neural network as an example to demonstrate the use effect of the tool and integrates the above tool into the FPGA to form a general system for fault tolerance performance evaluation.
[0068] Example 1, with reference to Figure 1 and Figure 2 , which is the first embodiment of the present invention. This embodiment provides an efficient fault injection and detection tool based on RTL. Users can flexibly configure the fault injection tool according to their own needs and use the fault detection module to obtain a detection report. On this basis, the fault tolerance performance of the target model or FPGA development board is evaluated. The evaluation process is as follows: Figure 1 As shown, the process includes:
[0069] Step S1: Configure the fault injection tool. The fault injection tool of the present invention is used to inject faults into the data flow between modules. The fault injection tool needs to be configured according to the data flow format. After selecting the fault type, configurable parameters include data accuracy, number of injection bits, duration, etc. Users can flexibly configure the fault injection tool according to their needs.
[0070] Step S2, performing random fault injection: given the RTL description of the target model, a fault injection bypass is created at the target point. After starting the target model, the intermediate data stream is extracted and randomly injected with faults without affecting the normal execution of the main line target model;
[0071] Step S3, fault detection; input the fault data stream generated by S2 to the fault detection module, and the detection module generates a fault detection report according to the agreed codeword;
[0072] Step S4, fault tolerance performance evaluation: compare the fault injection configuration of S1 with the fault detection report of S3. Since specific target applications (such as neural networks with high fault tolerance) or FPGA development boards (such as aerospace-grade FPGAs) have strong fault shielding capabilities, the fault tolerance performance can be preliminarily evaluated based on the comparison results.
[0073] In the RTL-based efficient fault injection tool:
[0074] The original data format includes two indicators: data accuracy and parallelism. Serial-to-parallel conversion is required before inputting into the fault injection module.
[0075] The fault injection tool of this invention supports input data formats with one base bit width (identical to the data precision) and up to three extended bit widths (determined by the degree of parallelism). It independently and randomly injects faults of any bit within each base bit width. It is compatible with three fault types: SEL, SET, and SEU. The specific impacts of these three fault types are shown in Table 1, and the fault duration is customizable. To enable multiple, continuous fault injections, a timer is started for each fault injection. The timer duration remains the same as the fault duration. Upon expiration, the timer pulls up the enable signal and samples the original data, maintaining a fault-free state pending the next fault injection.
[0076] Table 1: Specific impacts of three fault types
[0077] Fault type Specific impact Default duration (clock cycles) Single Event Latchup (SEL) Set or reset 100 Single Event Transient (SET) Bit Flip 10 Single Event Upset (SEU) Bit Flip 2
[0078] The present invention performs random fault injection, and the injection position is determined by a random number. The generation of random numbers needs to meet the requirements of three independent extended bit widths, and a specified number of non-repeating random numbers are generated within each basic bit width. Although there are system functions such as $random in Verilog for generating random numbers, this function cannot be synthesized and can only be used in testbench simulation scripts, and the multiple random numbers generated may be repeated. Therefore, consider writing a Python script on the software side to generate a random number file that meets the requirements. The random number file uses a txt text format, and each line stores the random number required for a single basic bit width. Each random number is saved in binary form. The basic bit width is 5, the extended bit widths are 1, 2, and 4 respectively, and the binary random number file for injecting two-bit faults is as follows: Figure 2 shown.
[0079] Embodiment 2 is the second embodiment of the present invention, which provides a forward fault detection module based on extended Hamming code:
[0080] Error-detection codes use a special algorithm to add redundant information to the original data. The receiver uses changes in the redundant bits during data transmission to determine whether a fault has occurred, eliminating the need for the sender to participate in the detection process again, significantly improving fault detection efficiency. Common error-detection codes include parity check codes, Hamming codes, and Reed-Solomon codes. Different codes can detect or correct different fault types and have different coding overheads. To maximize the effectiveness of fault detection and the feasibility of hardware deployment, this paper uses an extended Hamming code as the fault detection code, which can detect up to two faults and correct one fault at a time. The encoding and decoding principles are briefly explained using the (13,8) extended Hamming code as an example.
[0081] D=[p4,d7,d6,d5,d4,p3,d3,d2,d1,p2,d0,p1,p0] (1)
[0082]
[0083] The (13,8) extended Hamming code is suitable for input data with a bit width of 8, generating 5 check bits. The encoded data is shown in formula (1), where d i (i=0,1,…7) is the data bit, p i (i=0,1,…4) is the check bit. The check bit is completely generated by the data bit, and the generation rule is shown in formula (2). As shown in formula (3), the receiver generates the syndrome s according to the established rules. i (i=0,1,…4), the fault condition can be determined and corrected by looking up the syndrome in Table 2. The hardware adopts a pipelined design, which is divided into four stages: check bit generation, fault injection, syndrome generation, and fault information output. Each stage is carried out simultaneously to improve fault detection efficiency.
[0084] Table 2: Relationship between the correction factor and the fault type
[0085] <![CDATA[S4S3S2S1S0]]> Error location <![CDATA[S4S3S2S1S0]]> Error location 10001 p0 11000 p3 10010 p1 11001 d4 10011 d0 11010 d5 10100 p2 11011 d6 10101 d1 11100 d7 10110 d2 10000 p4 10111 d3 00000 No error
[0086] Example 3, reference Figure 3 、 Figure 4 、 Figure 5 and Figure 6 , which is the third embodiment of the present invention, provides IP core packaging and simulation waveforms:
[0087] like Figure 3 As shown in Figure 3, the fault injection module is encapsulated as a separate IP core, with input and output ports on the left and right sides, respectively. Table 3 shows the external interface and IP core parameter configuration information. Users can configure the IP core parameters according to their actual needs.
[0088] Table 3: External interface and IP core parameter configuration information
[0089]
[0090] Simulation waveform:
[0091] To demonstrate the fault injection effectiveness of the IP core, the following simulation waveforms are provided for three fault types and a fault-free state. The simulation stimulus is configured in a testbench, which reads a random number file and transfers it to the module under test instantiated in the testbench. Vivado automatically compiles the design under test, applies the stimulus based on the testbench, and generates simulation waveforms for all signals. The fault injection IP core uses a uniform configuration with an 8-bit base bit width and three extended bit widths of 8, 1, and 9 bits. A 2-bit error is injected into each base bit width.
[0092] The simulation waveform of SEL is as follows Figure 4 As shown in the figure, 16 bits of output data are randomly selected for illustration. The clock cycle T of a single SEL is 100, as shown by the blue line interval in the figure. The clock cycle after the blue line is the data to be injected with the fault, and the remaining 99 cycles are the data after the fault injection. SEL sets random bits, and some bits are pulled high in the second clock cycle, which is consistent with the random number result.
[0093] The simulation waveform of SET is as follows Figure 5 As shown in the figure, 16 bits of output data are randomly selected for illustration. The clock cycle for a single SET is T = 10, as shown by the blue line interval in the figure. The clock cycle after the blue line is the data to be injected with the fault, and the remaining 9 cycles are the data after the fault injection. SET resets the random bits, and some bits are pulled low in the second clock cycle, which is correct when compared with the random number result.
[0094] The simulation waveform of SEU is as follows Figure 6 As shown in the figure, 16 bits of output data are randomly selected for illustration. The clock cycle for a single SEU is T = 2, as indicated by the interval of the blue line in the figure. The clock cycle after the blue line is the data to be injected with the fault, and the remaining clock cycle is the data after the fault injection. The SEU flips random bits, and some bits flip in the second clock cycle, which is consistent with the random number result.
[0095] The simulation waveform without fault injection is as follows Figure 7 As shown in the figure, the data before and after the test remain unchanged and no fault is injected.
[0096] Example 4, reference Figure 8 , which is the fourth embodiment of the present invention, provides a convolutional layer fault-tolerant application design:
[0097] Convolutional neural networks (CNNs), as one of the basic algorithms of artificial intelligence, are widely used in the aerospace field in target detection, object classification, image reconstruction and other fields. Compared with high-computing platforms such as GPUs and NPUs, FPGAs have stronger reconfigurability. In recent years, deploying lightweight CNN applications on FPGAs has become a major trend. However, CNNs have higher computational complexity than other applications, and running in a space environment may cause operational failures due to radiation interference. Therefore, the present invention carries out fault-tolerant application design for the convolutional layer, the basic component of CNN, and uses the above-mentioned fault injection and detection tools to evaluate the fault-tolerant effect.
[0098] The RTL description of the convolution layer is provided by my previous work, including the steps of convolution window data acquisition, weight and input feature multiplication, multiplication result accumulation, bias addition, activation function, etc. The convolution layer fault tolerance protection process is as follows Figure 8 As shown, the model parameters include weights and biases, which are stored in on-chip RAM or off-chip DDR. The data path of the model parameters is long and easily affected by fault interference, resulting in erroneous data. Error correction and detection code (EDAC) is used to provide fault tolerance protection for weights and bias data. EDAC can be freely selected according to fault tolerance requirements and deployment feasibility. At the same time, N-module redundancy protection is used for the four computationally intensive IPs of multipliers, accumulators, adders, and activation functions. The specific module N can be selected according to resource conditions. After completing the fault tolerance protection, the fault injection and detection tool of the present invention is used to evaluate the fault tolerance performance of the convolutional layer. In order to evaluate the fault tolerance performance of each module in the convolutional layer separately, faults are injected before the N-module redundancy and after the EDAC, and fault detection points are set at the module output end or the data receiving end to determine whether the injected fault affects the correct output result.
[0099] Example 5, with reference to Figure 9 and Figure 10 , which is the fifth embodiment of the present invention, provides a general system for fault-tolerance performance evaluation. The system architecture includes:
[0100] Integrate the fault injection and detection tools into the FPGA development board, and use the high-speed bus and host computer to form a complete spatial FPGA fault tolerance performance evaluation system. The system architecture is as follows: Figure 9As shown. The host computer is used for data generation and analysis processing, and the source data provides input for the FPGA-side application (such as the input feature map of CNN); the random number file is generated in real time, and the fault injection point is updated; the golden output and the application output are the final outputs generated by the application without fault and after fault injection respectively; the evaluation report analyzes the fault tolerance performance of the spatial FPGA. FPGA is used for fault injection and program running, and the conversion of data transmission protocols is completed in the format conversion module, specifically including PCIe and AXI, AXI and application interface, etc.; the design to be tested includes the application and FPGA device model; the fault injection and detection tool injects faults into the design to be tested and preliminarily determines the fault condition. Due to the frequent large-scale data exchange between the host computer and the FPGA, PCIe with higher transmission efficiency is selected as the data transmission protocol. By adopting the evaluation system proposed in the present invention, it is possible to realize the rapid fault tolerance performance evaluation of different application and device model combinations on the basis of keeping the system structure unchanged, and significantly improve the evaluation efficiency.
[0101] Conduct an assessment, assessment process:
[0102] In order to give full play to the advantages of FPGA parallel processing and speed up the fault detection process, the present invention proposes a step-by-step fault tolerance performance evaluation process based on the above system architecture. Figure 10 As shown;
[0103] a. System configuration and data generation: After selecting an application and device model combination, the fault injection and detection tool is closed. The host computer generates source data and random number files and sends them to the FPGA.
[0104] b. Generate golden output: The FPGA runs the design under test without fault-tolerance protection, obtains golden output under fault-free conditions for subsequent fault detection, and feeds the golden output back to the host computer;
[0105] c. Preliminary fault detection: Start the fault injection and detection tool, inject a fault into the design under test, and then rerun the design under test to obtain the application output. The fault detection module detects faults based on the application output. The error detection capability is related to the codeword used. If all faults are successfully detected, it is considered "preliminary detection results are valid" and the process goes directly to step e; otherwise, the process goes to step d.
[0106] d. Host computer fault troubleshooting: The host computer compares the golden output with the application output again to determine whether the fault point is consistent with the random number file and obtain the complete fault situation;
[0107] e. Generate an evaluation report. The host computer summarizes all error outputs of the design under test based on the fault detection results, divides it by the total number of fault injections to obtain the error rate, and generates a final fault-tolerance performance evaluation report based on indicators such as the error rate, fault injection configuration, and source data size. This completes the fault-tolerance performance evaluation of one design combination under test, and you can return to step a to continue with other combinations.
[0108] In summary, the present invention provides an efficient RTL-based fault injection tool: the fault injection tool proposed in the present invention is compatible with multiple fault types, the injection bits and duration can be freely defined, the test process supports flexible switching between different fault types, and significantly improves fault injection efficiency;
[0109] Extended Hamming Code-based forward fault detection module: This module uses extended Hamming code to detect faults after fault injection. The receiver directly determines the fault type based on the agreed codeword without sending the original data.
[0110] The aforementioned fault injection and detection tools are added to the convolutional layer of the neural network implemented in RTL, and a triple-module redundancy strategy is used to protect key operators and enhance the fault tolerance of the convolutional layer.
[0111] Universal system for evaluating spatial FPGA fault-tolerance performance: Adopting a modular design concept, the fault injection and detection tools are encapsulated as IP cores and integrated into the FPGA development board. A data transmission interface is reserved for the host computer. Together with a high-speed bus, this system forms a complete universal system for evaluating spatial FPGA fault-tolerance performance. This system can be used to evaluate the fault-tolerance performance of different FPGA models.
[0112] Compared with the existing technology, the present invention improves the efficiency of fault injection and detection. At the same time, the application results of the fault injection and detection tool are demonstrated using the convolutional layer of a neural network as an example. Finally, a general system is proposed to evaluate the fault tolerance performance of different models of FPGA.
[0113] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A fault injection and detection method based on RTL, characterized in that: include, Step S1: Configure the fault injection tool, select the fault type according to the RTL description and data flow form of the target model, and configure the data accuracy, injection bit number and duration parameters; Step S2: creating a fault injection bypass in the RTL data flow of the target model, generating a fault injection location based on a random number file, and injecting a fault of a specified type and parameters into the intermediate data flow; Step S3: Use extended Hamming code to encode and decode the data stream after the fault is injected, generate a syndrome and compare it with the preset codeword, and output a fault detection report; Step S4: Based on the fault injection configuration and the detection report, the fault tolerance performance of the target model or FPGA is evaluated by error rate calculation.
2. The RTL-based fault injection and detection method according to claim 1, wherein: The fault types described in step S1 include single event upset SEU, single event transient SET and single event latch SEL, where: The default duration of SEU is 2 clock cycles, corresponding to a bit flip; The default duration of SET is 10 clock cycles, corresponding to bit reset; The default duration of SEL is 100 clock cycles, corresponding to a set or reset operation.
3. The RTL-based fault injection and detection method according to claim 1, wherein: The method for generating the random number file in step S2 is: Generate random numbers that meet the basic bit width and extended bit width constraints through Python scripts on the software side; Generate non-repeating binary random numbers within each basic bit width and store them in text file format; The random number file is input to the fault injection module through the parallel interface.
4. The RTL-based fault injection and detection method according to claim 1, wherein: The extended Hamming code in step S3 adopts the (13,8) encoding rule, including: The generation logic of data bits and check bits is based on XOR operation; The receiving end determines the fault location and type by looking up the syndrome table; A pipelined design is used to achieve parallel processing of checksum generation, fault injection, syndrome generation and fault output.
5. The RTL-based fault injection and detection method according to claim 1, wherein: The fault injection tool is encapsulated as a configurable IP core, whose interface includes: Input port: clock signal, reset signal, fault type selection, data stream to be injected and random number file; Output port: data flow after fault injection; Configurable parameters include basic bit width, extended bit width, fault duration, and number of injected bits.
6. The RTL-based fault injection and detection method according to claim 1, wherein: The fault tolerance performance evaluation in step S4 includes: Compare the golden output without faults with the application output after fault injection and calculate the error rate; Generate a fault tolerance performance evaluation report based on error rate, fault type, and data size.
7. The RTL-based fault injection and detection method according to claim 1, wherein: The fault-tolerance protection design for the convolutional layer of the convolutional neural network includes: The weight and bias data are protected by error correction and detection code EDAC; Adopt N-module redundancy strategy for multipliers, accumulators, adders and activation function modules; Set up fault injection points before the redundant module and after the EDAC, and judge the fault tolerance effect through the output of the detection point.
8. The RTL-based fault injection and detection method according to claim 1, wherein: Integrated into the space FPGA fault-tolerance performance evaluation system, including: Host computer module: used to generate source data, random number files and analysis and evaluation reports; FPGA module: includes fault injection tools, detection modules and the application to be tested; Data transmission module: uses PCIe protocol to achieve high-speed data interaction between the host computer and FPGA.
9. The RTL-based fault injection and detection method according to claim 8, wherein: The assessment process includes: The first round generated trouble-free gold output; The second round of fault injection and testing the validity of the preliminary results; If it is invalid, perform a secondary check by comparing the golden output with the application output on the host computer.
10. The RTL-based fault injection and detection method according to claim 1, wherein: Fault injection bypass is designed to not affect the normal execution of the trunk model, specifically through: A timer controls a fault injection enable signal; The bypass switches to the faulty data flow within a specified clock cycle, and maintains the original data flow in the remaining cycles.
Citation Information
Cited By
Anti-radiation soft error analysis method for AES encryption circuit based on CVPI
CN121562515A
Hierarchical anomaly recovery method and system for triple modular redundancy
CN122152596A