Collaborative Verification and Reliable Action Method and System for Protecting Control Equipment against Soft Error Maloperation
Through the three-core collaborative verification and self-repair mechanism, the problem of insufficient soft error protection and self-repair capabilities of the existing technology relay protection control devices is solved, efficient detection and automatic recovery of complex soft errors is achieved, and the stability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510404969.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-02
AI Technical Summary
When existing relay protection control devices face complex soft errors (such as single-particle effect and multi-bit flip), they lack effective data layer and program layer verification mechanisms, resulting in insufficient system stability and reliability and incomplete self-repair capabilities.
The three-core collaborative verification and self-healing mechanism are adopted to ensure the stability and reliability of the system by starting the collaborative work between the CPU, protecting the CPU and co-processing CPU.
Effectively detect and fix single-bit flip and multi-bit flip errors, realize real-time verification of program running status, enhance the system's anti-soft error capability and automatic recovery capability, prevent erroneous actions, and improve the stability and reliability of the system.
Smart Images

Figure CN119916729B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of electric power relay protection control and measurement control, and in particular to a method and system for coordinated verification and reliable action of soft error prevention of protection control equipment. Background Art
[0002] As the complexity and importance of power system relay protection control devices increase, the environment for safe operation of power systems becomes more and more harsh, and soft errors have become a significant threat in relay protection control devices. Soft errors are transient faults caused by single-element effects (SEU) or multiple-bit upsets (MBU), which often occur when high-energy particles hit memory or logic circuits. Although such soft errors do not directly cause physical damage to the hardware, they can profoundly affect the normal function of the equipment, and may cause data flips, calculation errors, program execution anomalies, and hidden equipment failures, which in turn cause erroneous outputs or even erroneous operations of relay protection control devices, leading to overall instability of the system, seriously threatening the safety of the power system.
[0003] In the prior art, relay protection control devices usually adopt a dual-CPU redundant architecture. Through the output results of two independently running CPUs, it is judged whether the system is in a normal state and jointly determines the protection action state behavior. This type of dual redundant design can improve the system reliability to a certain extent. However, with the frequent occurrence of soft error problems caused by the radiation effect of space particles, it is difficult to effectively deal with complex soft error scenarios. At present, there are mainly the following problems:
[0004] (1) Limited soft error protection capabilities
[0005] The dual-CPU redundant architecture mainly detects system failures by comparing the output results of the two CPUs. However, this result comparison mechanism is only suitable for capturing errors at the hardware level and cannot cope with more complex soft errors (such as MBU in the storage unit). In the case of MBU, multiple bits in the memory flip at the same time, which may affect the output of both CPUs at the same time, making their results consistent but wrong, and thus unable to detect this synchronization error. Because the architecture lacks the coordinated verification function of the data layer and the program layer, the protection capability of soft errors is obviously insufficient in these complex scenarios.
[0006] (2) Lack of program verification mechanism
[0007] Although the existing dual - CPU architecture can monitor and compare the final output results, it cannot monitor the program execution process in real time. During the program execution process, soft errors may affect the execution order or logic of instructions, resulting in program execution errors. The dual - CPU architecture fails to verify these errors in the intermediate process, making it difficult to capture the errors that occur during program operation. Especially when the program execution error is consistent with the data output result, such errors will be ignored, leading to system malfunction or incorrect processing.
[0008] (3) Imperfect self - repair mechanism
[0009] In the existing technology, the dual - CPU architecture can detect faults in some cases, but the system lacks effective self - repair ability. Once an error is detected, the system often needs to rely on external intervention or manual repair, which not only increases the response time but also may cause the system to be in an unstable or shutdown state for a long time. The lack of an automated self - repair mechanism makes the existing system unable to recover in time after a soft error occurs, thus affecting the stability and reliability of the system.
[0010] A relay protection control method and system based on a multi - core processor chip in Patent 202310912543.X proposes a solution through a multi - core processor. By using the synchronous verification between the protection core and the startup core, multiple data comparisons and verifications are carried out to ensure the reliability and safety of the relay protection control device. The specific implementation includes: 1. Using a dual - core architecture, including a startup CPU and a protection CPU, the two cores work independently and verify each other's output results; 2. Through redundant backup of programs and data, the system verifies when data errors or program anomalies occur to determine whether the system is in a healthy state.
[0011] Although this technology shows certain advantages in solving some hardware faults, it still has some limitations for complex soft errors (such as SEU, MBU). Especially when soft errors occur in multiple storage units or logic units, it is difficult to accurately locate and correct the errors simply by relying on dual - core comparison. Summary of the Invention
[0012] The purpose of the present invention is to overcome the deficiencies of the existing dual - CPU redundant architecture of relay protection control devices in aspects such as soft error protection, program verification, and self - repair ability. A three - core collaborative verification and action export execution method and system are proposed for soft error protection of relay protection control devices to improve the reliability and stability of the device system.
[0013] The technical solution for achieving the object of the present invention is as follows: On the one hand, a collaborative verification and reliable action system for protecting the soft error prevention and misoperation of a control device is provided. The system includes a startup CPU, a protection CPU, and a coprocessing CPU built in the protection control device. Through the cooperation among the three, a multi-level verification and self-repair mechanism is formed to achieve the soft error protection of the system and the execution of the action outlet.
[0014] The system performs multi-level verification through the startup CPU, the protection CPU, and the coprocessing CPU, including: performing data self-verification by the startup CPU and the protection CPU respectively, and transmitting the verification results to the coprocessing CPU; performing program co-verification by the coprocessing CPU.
[0015] Meanwhile, the coprocessing CPU makes an arbitration decision based on each verification result to determine whether the system outlet acts or performs exception handling.
[0016] The self-repair mechanism specifically includes: triggering a self-repair instruction during the exception handling, and the system performs self-repair processing until it works normally.
[0017] Further, the startup CPU is responsible for system startup, data self-check, and relay protection control outlet startup control tasks, and is built with a RAM ECC check module for performing real-time self-verification on the operating data of the startup CPU.
[0018] The protection CPU is responsible for core relay protection control logic calculation, and is built with a RAM ECC check module for performing real-time self-verification on the operating data of the protection CPU.
[0019] Both the startup CPU and the protection CPU are equipped with a data backup module and a data self-repair module, which are respectively used for backing up and repairing the operating data of the startup CPU / protection CPU.
[0020] Before performing data self-verification and transmitting the program to the coprocessing CPU, the startup CPU and the protection CPU respectively back up the operating data to the corresponding data backup modules.
[0021] Further, the coprocessing CPU includes a monitoring and verification module and an arbitration module.
[0022] The monitoring and verification module is used for verifying and monitoring the data and programs of the startup CPU and the protection CPU.
[0023] The arbitration module is used for determining whether the system outlet acts or performs exception handling according to the output of the monitoring and verification module.
[0024] Further, the monitoring and verification module includes a receiving module, a co-verification module, a self-repair module, and a monitoring module.
[0025] The receiving module is configured to receive the self-verification results of the startup CPU and the protection CPU respectively, and transmit them to the arbitration module, and at the same time transmit them to the monitoring module;
[0026] The co-verification module is configured to verify the programs transmitted by the startup CPU and the protection CPU respectively, and transmit the verification results to the arbitration module, and at the same time transmit them to the monitoring module;
[0027] The self-repair module is configured to receive self-repair instructions, perform self-repair on the program, and send the self-repair instructions to the startup CPU and the protection CPU to trigger their respective self-repair functions;
[0028] The monitoring module is configured to receive and store in real time the feedback information of the receiving module, the co-verification module, and the self-repair module to update the system state in real time; if it is monitored that the self-repair status of each module is self-repair successful, a recovery instruction is sent to the arbitration module, and the arbitration module controls the system state to return to normal, otherwise an exception instruction is sent to the arbitration module, and the arbitration module generates a locking signal to control the system to enter a locked state, and an alarm signal is sent through the monitoring system.
[0029] Further, the startup CPU and the protection CPU transfer the program to the co-verification module in a partition round-robin manner, and in each interrupt control cycle, a program of one partition is transferred in sequence, and it is ensured that all program partitions are transferred within a specified time.
[0030] Further, after receiving the program transferred by partition round-robin, the co-verification module generates a unique identifier for each partition as a snapshot, and compares it with the snapshot of the previous time of each partition one by one to output the verification result; at the same time, the co-verification module updates the snapshot of the current partition;
[0031] The verification result includes pass or fail. If the comparison is consistent, the verification result is pass, otherwise the verification result is fail;
[0032] At the same time, after receiving the program transferred by partition round-robin, the co-verification module also stores and backs up the program in a partition manner for the self-repair module to call; after receiving the self-repair instruction, the self-repair module performs self-repair by calling the program backed up in the co-verification module.
[0033] Further, the monitoring module is also configured to identify latent system faults and give an alarm while ensuring real-time update of the system state.
[0034] Further, the arbitration module includes a consistency judgment module, a decision-making module, and an exception handling module;
[0035] The consistency judgment module is used to receive the signal data output by the receiving module and the co-verification module, and respectively judge whether the self-verification result of the startup CPU is consistent with the verification result of the co-verification module for the startup CPU program, and whether the self-verification result of the protection CPU is consistent with the verification result of the co-verification module for the protection CPU program; if the verification results are all consistent, it outputs a consistency signal to the decision-making module, otherwise it generates an abnormal status signal to the decision-making module;
[0036] The decision-making module arbitrates through the multi-taking and multi-using principle. Specifically, when the startup relays and action relays corresponding to the startup CPU and the protection CPU are both in the "allowed to output" state, and the consistency judgment module outputs a consistency signal, and a "allowed to export action" instruction is generated to control each relay to continue to work normally; otherwise, it sends an activation instruction to the abnormal handling module;
[0037] The abnormal handling module is used to generate a self-repair instruction after receiving the activation instruction and transfer the self-repair instruction to the self-repair module.
[0038] Furthermore, in a new interrupt control cycle, the consistency judgment module obtains the repaired / updated system state from the monitoring module and continues to perform normal self-checking and arbitration operations to ensure the continuous and stable operation of the system.
[0039] On the other hand, based on the system, a co-verification and reliable action method for preventing soft errors of a protection control device is provided. The method includes the following steps:
[0040] Step 1, the startup CPU and the protection CPU respectively perform real-time data self-verification, and send the verification results to the co-processing CPU. At the same time, the startup CPU and the protection CPU send their respective programs to the co-processing CPU;
[0041] Step 2, the co-processing CPU processes the received program, generates a corresponding snapshot, compares the snapshot with the previous snapshot, and generates a program verification result;
[0042] Step 3, the co-processing CPU judges whether the self-verification result of the startup CPU in Step 1 is consistent with the verification result of the co-processing CPU for the startup CPU program in Step 2, and whether the self-verification result of the protection CPU in Step 1 is consistent with the verification result of the co-processing CPU for the protection CPU program in Step 2. If they are all consistent, it generates a consistency signal; otherwise, it generates an abnormal status signal;
[0043] Step 4: The coprocessor CPU makes a blanking decision based on the signal generated in Step 3. If the consistency signal is output in Step 3, and the start relays and action relays corresponding to the start CPU and the protection CPU are both in the "allowed to output" state, then the coprocessor CPU generates an "allowed to export action" instruction to control each relay to continue normal operation, including driving the export relay to execute an action; otherwise, the coprocessor CPU generates an exception handling signal and proceeds to Step 5.
[0044] Step 5: Generate a self-repair instruction to trigger the self-repair functions of the start CPU, the protection CPU, and the coprocessor CPU. At the same time, monitor the self-repair result in real time. If the self-repair result is successful, update the snapshot data and feedback to Step 1 to continue execution; otherwise, proceed to Step 6.
[0045] Step 6: Enter the locked state, issue an alarm signal, and end the process.
[0046] Compared with the prior art, the significant advantages of the present invention are as follows:
[0047] (1) Improved soft error protection ability
[0048] The present invention adopts a multi-core collaborative architecture and ECC checking, which can effectively detect and repair single-bit flip errors, and can also detect multi-bit flip errors. For multi-bit flips, the system triggers a self-repair mechanism through an exception handling module, and uses backup data to restore damaged programs or data, making up for the deficiency of difficult repair of multi-bit flips, and ensuring the stability and reliability of the system under complex soft error conditions.
[0049] (2) Real-time verification of program running status
[0050] The present invention adopts a partition polling mechanism to perform real-time verification on the programs of the start CPU and the protection CPU. It can complete the verification of all program partitions within a short time of each interrupt control cycle, effectively improving the error detection efficiency and accuracy during program execution while reducing the CPU load. When a program error is detected, the coprocessor CPU quickly restores the damaged program code through the self-repair module to ensure the continuous and stable operation of the system.
[0051] (3) Multi-level arbitration and self-repair mechanism
[0052] Through an arbitration mechanism based on the 2-out-of-2 principle of consistency judgment, the present invention can more accurately judge the system state at the relay action outlet, preventing misoperation of the outlet caused by incorrect signals. When an exception is detected, the system calls backup data through the self-repair module to repair damaged programs and data, thereby enhancing the system's anti-soft error ability and automatic recovery ability.
[0053] The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings
[0054] Figure 1 It is a schematic structural diagram of a triple-core collaborative verification system for soft errors in a relay protection control device in an embodiment.
[0055] Figure 2 It is a multi-level verification logic flowchart for soft error protection in an embodiment.
[0056] Figure 3 It is a system arbitration and self-repair flowchart under the triple-core collaborative verification architecture in an embodiment.
[0057] Figure 4 It is a schematic diagram of a relay drive and execution circuit in an embodiment.
[0058] Reference Signs: 1 - Start CPU, 2 - Protection CPU, 3 - Coprocessing CPU, 4 - Monitoring and Inspection Module, 5 - Arbitration Module, 6 - Receiving Module, 7 - Cooperative Verification Module, 8 - Self-Repair Module, 9 - Monitoring Module, 10 - Consistency Judgment Module, 11 - Decision Module, 12 - Exception Handling Module, 13 - CPU System, 14 - Relay Drive Module, 15 - Actuator, 16 - Action Relay, 17 - Start Relay, 18 - Operation Box. Detailed Embodiments
[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0060] It should be noted that if there are directional indications (such as up, down, left, right, front, back,...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0061] In addition, if the descriptions such as "first" and "second" are involved in the embodiments of the present invention, these descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0062] The present invention provides a method and system for cooperative verification and reliable action for protecting control equipment from soft errors and preventing misoperation, which will be described in detail below for a relay protection control device.
[0063] In one embodiment, a system for cooperative verification and reliable action for protecting a relay protection control device from soft errors and preventing misoperation is provided. The system includes a startup CPU, a protection CPU, and a coprocessing CPU built in the protection control equipment. Through the cooperation among the three, a multi-level verification and self-repair mechanism is formed to achieve soft error protection for the system and execution of the action output.
[0064] The system performs multi-level verification through the startup CPU, the protection CPU, and the coprocessing CPU, including:
[0065] First-level data verification: The startup CPU and the protection CPU respectively perform self-verification on the data and transfer the verification results to the coprocessing CPU.
[0066] Second-level program cooperative verification: The coprocessing CPU performs program cooperative verification.
[0067] Meanwhile, the coprocessing CPU makes an arbitration decision based on each verification result to determine whether the system output acts or performs exception handling.
[0068] The self-repair mechanism specifically includes: When exception handling is triggered, a self-repair instruction is triggered, and the system performs self-repair processing until it works normally.
[0069] Here, the multi-level verification mechanism aims to ensure the integrity and reliability of the data and programs of the relay protection control device during operation, especially having a higher protection ability when dealing with soft errors.
[0070] Further, in one of the embodiments, the startup CPU is responsible for system startup, data self-check, and startup control of the relay protection control output. It has a built-in RAM ECC verification module for performing real-time self-verification on the operation data of the startup CPU.
[0071] The protected CPU is responsible for the core relay protection control logic calculation and is built-in with a RAM ECC check module for real-time self-checking of the operation data of the protected CPU;
[0072] Both the startup CPU and the protected CPU are equipped with a data backup module and a data self-repair module, which are used to back up and repair the operation data of the startup CPU / protected CPU respectively;
[0073] Before the startup CPU and the protected CPU perform data self-checking and transfer the program to the coprocessor CPU, the operation data is backed up to the corresponding data backup module respectively.
[0074] Preferably here, the startup CPU and the protected CPU are transferred to the coprocessor CPU through an internal bus or a high-speed transmission channel.
[0075] Here, the data processed by the startup CPU includes but is not limited to initialization parameters, system status data, user configuration data, log data, etc.; the data processed by the protected CPU includes but is not limited to measurement data, protection setting parameters, historical fault records, etc.
[0076] Here, the ECC check principle: for single-bit flips, the ECC self-check can detect and repair single-bit errors; for multi-bit flips (MBU), the ECC self-check can detect multi-bit errors. Although it cannot be automatically repaired, it will generate an ECC self-repair exception signal, which is finally transmitted to the exception handling module of the coprocessor CPU for subsequent repair.
[0077] Further, in one embodiment, the coprocessor CPU includes a monitoring and inspection module and an arbitration module;
[0078] The monitoring and inspection module is used to check and monitor the data and programs of the startup CPU and the protected CPU;
[0079] The arbitration module is used to decide the system outlet action or perform exception handling according to the output of the monitoring and inspection module.
[0080] Here, the monitoring and inspection module includes a receiving module, a co-check module, a self-repair module and a monitoring module;
[0081] The receiving module is used to receive the self-check results of the startup CPU and the protected CPU respectively, and transmit them to the arbitration module and the monitoring module at the same time;
[0082] The co-check module is used to check the programs transmitted by the startup CPU and the protected CPU respectively, and transmit the check results to the arbitration module and the monitoring module at the same time;
[0083] The self - repair module is used to receive self - repair instructions and send the self - repair instructions to the startup CPU and the protection CPU to trigger their respective self - repair functions. Here, self - repair includes program repair, data repair, and system state recovery.
[0084] The monitoring module is used to receive and store in real - time the feedback information of the receiving module, the co - verification module, and the self - repair module to update the system state in real - time. If it monitors that the self - repair status is successful self - repair, it sends a recovery instruction to the arbitration module, and the arbitration module controls the system state to return to normal. Otherwise, it sends an exception instruction to the arbitration module, and the arbitration module generates a locking signal to control the system to enter the locked state, close the relay action outlet, avoid system misoperation, and send an alarm signal through the monitoring system.
[0085] Here, through the monitoring module, in the next round of self - inspection process, the arbitration module can obtain the repaired self - inspection status from the monitoring module to ensure the normal operation of the subsequent system processes and the reliability of the relay protection control device.
[0086] Preferably, in some embodiments, the startup CPU and the protection CPU transfer the programs to the co - verification module in a partition - round - robin manner, and in each interrupt control cycle, a program of one partition is transferred in sequence, and it is ensured that all program partitions are completely transferred within a specified time.
[0087] Here, the program transfer and round - robin method distinguish multiple functional areas, and the number of each functional area can be configured according to the actual application scenario. When the system verifies, a hash value or a check code is generated, or other verification means are used for comparison to ensure the integrity of the program and data. The verification means are not limited to hash values or check codes.
[0088] Exemplarily preferably, the program of the startup CPU has four - round - robin partitions: initialization code area, interrupt handling code area, system management code area, and data monitoring code area.
[0089] The program of the protection CPU has five - round - robin partitions: protection logic code area, real - time data processing area, setting parameter management area, interrupt task handling code area, and latent fault monitoring area.
[0090] Preferably, in some embodiments, after the co - verification module receives the programs transferred in partition - round - robin, it generates a unique identifier for each partition as a snapshot, and compares it with the snapshot of the previous time of each partition one - by - one to output the verification result. At the same time, the co - verification module updates the snapshot of the current partition.
[0091] The verification result includes pass or fail. If the comparison is consistent, the verification result is pass; otherwise, the verification result is fail.
[0092] Here, the snapshot data can be generated in various forms, including but not limited to hash values, check codes, encrypted signatures, or other data consistency verification methods.
[0093] Preferably, in some embodiments, after receiving the program passed by the partition round-robin, the co-verification module also stores and backs up the program in a partition manner for the self-repair module to call; after receiving the self-repair instruction, the self-repair module performs self-repair by calling the program backed up in the co-verification module.
[0094] Here, since the program is stored in the co-verification module 7 in a partition manner, it is ensured that in case of an exception, the self-repair module 8 can quickly recover the partition-level program, avoiding large-scale program damage.
[0095] Here, the self-repair process can handle various error types, including but not limited to single-bit errors, multi-bit errors, and other forms of soft errors.
[0096] Preferably, in some embodiments, the monitoring module is further configured to identify latent system faults and give an alarm while ensuring real-time update of the system state.
[0097] Further, in one of the embodiments, the arbitration module includes a consistency judgment module, a decision-making module, and an exception handling module;
[0098] The consistency judgment module is configured to receive the signal data output by the receiving module and the co-verification module, and respectively judge whether the self-verification result of the startup CPU is consistent with the verification result of the co-verification module for the startup CPU program, and whether the self-verification result of the protection CPU is consistent with the verification result of the co-verification module for the protection CPU program; if the verification results are both consistent, it outputs a consistency signal to the decision-making module, otherwise it generates an abnormal state signal to the decision-making module;
[0099] The decision-making module performs arbitration based on the 2-out-of-2 principle of consistency judgment. Specifically, when the startup relays and action relays corresponding to the startup CPU and the protection CPU are both in the "allowed to output" state, and the consistency judgment module outputs a consistency signal and generates an "allowed to export action" instruction, it controls each relay to continue to work normally; otherwise, it sends an activation instruction to the exception handling module;
[0100] The exception handling module is configured to generate a self-repair instruction after receiving the activation instruction and transmit the self-repair instruction to the self-repair module.
[0101] Preferably, in some embodiments, during the new interruption control cycle, the consistency judgment module obtains the repaired / updated system state from the monitoring module 9 and continues with normal self-checking and arbitration operations to ensure the continuous and stable operation of the system.
[0102] In one embodiment, a collaborative verification and reliable operation method for preventing soft errors in a relay protection control device is provided. The method includes the following steps:
[0103] Step 1, the startup CPU 1 and the protection CPU 2 respectively perform real-time data self-verification, and send the verification results (including error detection flags, error types, and repair status) to the coprocessor CPU 3. At the same time, the startup CPU 1 and the protection CPU 2 send their respective programs to the coprocessor CPU 3;
[0104] Step 2, the coprocessor CPU 3 processes the received programs, generates corresponding snapshots, and compares the snapshots with the previous snapshots to generate program verification results;
[0105] Step 3, the coprocessor CPU 3 determines whether the self-verification result of the startup CPU 1 is consistent with the verification result of the program of the startup CPU 1 by the coprocessor CPU 3, and whether the self-verification result of the protection CPU 2 is consistent with the verification result of the program of the protection CPU 2 by the coprocessor CPU 3. If both are consistent, a consistency signal is generated; otherwise, an abnormal state signal is generated;
[0106] Step 4, the coprocessor CPU 3 makes a punching decision according to the signal generated in Step 3: If the consistency signal is output in Step 3, and the startup relays and action relays corresponding to the startup CPU 1 and the protection CPU 2 are both in the "allowed output" state, then the coprocessor CPU 3 generates an "allowed export action" instruction to control each relay to continue normal operation, including driving the export relay to perform an action; otherwise, the coprocessor CPU 3 generates an abnormal processing signal and proceeds to Step 5;
[0107] Step 5, generate a self-repair instruction to trigger the self-repair functions of the startup CPU 1, the protection CPU 2, and the coprocessor CPU 3; at the same time, monitor the self-repair results in real time. If the self-repair results are successful, update the snapshot data and feedback to Step 1 to continue execution; otherwise, proceed to Step 6;
[0108] Step 6, enter the locked state, issue an alarm signal, and end the process.
[0109] Further, in one of the embodiments, the method further includes:
[0110] Before Step 1: The startup CPU and the protection CPU back up the running data.
[0111] Preferably, in some embodiments, in step 1, the startup CPU and the protection CPU respectively perform real-time self-check on the operation data required during the operation of the CPU through their built-in RAM ECC self-check modules.
[0112] Preferably, in some embodiments, in step 1, the startup CPU and the protection CPU transfer the program code to the co-check module of the co-processing CPU in a partition-by-round-robin manner. In each interrupt control cycle, the program code of one partition is transferred in sequence to ensure that all program partitions are co-checked within the specified time.
[0113] Preferably, in some embodiments, in step 2, the snapshot data can be generated in various forms, including but not limited to hash values, check codes, encrypted signatures, or other data consistency verification methods.
[0114] Preferably, in some embodiments, in step 5, the self-repair includes:
[0115] Program self-repair: The self-repair module of the co-processing CPU extracts the corresponding program partition from the backup program and replaces the damaged program code.
[0116] Data self-repair: The startup CPU and the protection CPU respectively extract the correct backup data from the backup data in the RAM storage unit and overwrite the damaged data. The backup data is updated before each run to ensure data integrity.
[0117] Further exemplarily, in combination with the specific modules of the above-mentioned collaborative check and reliable action system for soft error prevention and misoperation of relay protection control devices, the method includes:
[0118] Step 1: Data self-check
[0119] a. Startup and protection data backup
[0120] Used to back up the operation data to the internal RAM storage units of the startup CPU and the protection CPU before data self-check.
[0121] b. Data check of the startup CPU and the protection CPU
[0122] The startup CPU and the protection CPU respectively perform real-time data self-check. Each CPU performs real-time self-check on the operation data required during the operation of the CPU through its built-in RAM ECC self-check module. The data processed by the startup CPU includes initialization parameters, system status data, user configuration data, log data, etc.; the data processed by the protection CPU includes measurement data, protection setting parameters, historical fault records, etc.
[0123] Principle of ECC check: For single-bit flips, ECC self-checking can detect and correct single-bit errors; for multi-bit flips (MBU), ECC self-checking can detect multi-bit errors. Although it cannot be automatically corrected, it will generate an ECC self-repair exception signal, which is finally transmitted to the exception handling module of the coprocessor CPU for subsequent repair.
[0124] c. Transmission of data self-check result
[0125] While the startup CPU and the protection CPU are performing data self-checking, they will send the ECC check results (including error detection flags, error types, and repair status) to the receiving module of the coprocessor CPU for subsequent judgment and arbitration.
[0126] Step 2: Run program co-verification
[0127] a. Program round-robin transmission of startup CPU and protection CPU
[0128] The startup CPU and the protection CPU transmit the program code to the co-verification module of the coprocessor CPU in a partition round-robin manner. In each interrupt control cycle, the program code of one partition is transmitted in sequence to ensure that all program partitions are co-verified within the specified time.
[0129] Four-round-robin partitions of the startup CPU's program: initialization code area, interrupt handling code area, system management code area, data monitoring code area;
[0130] Five-round-robin partitions of the protection CPU's program: protection logic code area, real-time data processing area, setting parameter management area, interrupt task handling code area, latent fault monitoring area.
[0131] b. Program snapshot generation and comparison
[0132] Snapshot generation: After receiving the program code, the co-verification module generates corresponding snapshot data for the received program code. The snapshot data can be generated in various forms, including but not limited to hash values, check codes, encrypted signatures, or other data consistency verification methods.
[0133] Snapshot comparison: The co-verification module compares the generated snapshot with the previous snapshot to ensure the integrity and correctness of the program. If the snapshots are consistent, it is considered that the program has not made an error; if they are inconsistent, an output exception signal is generated.
[0134] c. Transmission of program co-verification result
[0135] The program co-verification results (including whether the comparison is consistent and the verification completion status) are transmitted to the consistency judgment module for further analysis, and at the same time, they are transmitted to the monitoring module for real-time status monitoring.
[0136] Step 3: Consistency Judgment and Arbitration Decision
[0137] a. Consistency Judgment
[0138] The consistency judgment module obtains the self - verification status of the startup CPU and the protection CPU passed from the receiving module, and at the same time receives the program co - verification status passed from the co - verification module, and judges whether the results of the output verification status of the protection and startup CPUs are consistent.
[0139] If the status results are consistent, an output consistency signal is generated;
[0140] If the status results are inconsistent, an output abnormal status is generated and passed to the decision - making module.
[0141] b. 2 - out - of - 2 Arbitration Decision Based on Consistency Judgment
[0142] Based on the consistency judgment result, the decision - making module arbitrates the signals of the startup CPU, the protection CPU, and the co - verification module, adopting the 2 - out - of - 2 principle: when the three signals are consistent and all are allowed to output, the system generates an "allowed export action" instruction to control the relay to continue normal operation; if the signals are inconsistent or an abnormality is detected, the abnormal handling process is entered.
[0143] Step 4: Abnormal Handling and Self - Repair
[0144] a. The Abnormal Handling Module Triggers the Self - Repair Instruction
[0145] After receiving the abnormal status, the abnormal handling module generates a self - repair instruction. This instruction is passed to the self - repair module and triggers the self - repair module to perform program self - repair, and at the same time sends a data self - repair instruction to the startup CPU and the protection CPU.
[0146] b. Program Self - Repair and Data Self - Repair
[0147] Program self - repair: The self - repair module of the co - processing CPU extracts the corresponding program partition from the backup program and replaces the damaged program code.
[0148] Data self - repair: The startup CPU and the protection CPU respectively extract the correct backup data from the backup data in the RAM storage unit to overwrite the damaged data. The backup data is updated before each run to ensure data integrity.
[0149] c. Status Monitoring after Repair Completion
[0150] The monitoring module monitors the repair result in real - time:
[0151] If the self - repair is successful, the system status will return to normal and the snapshot data will be updated.
[0152] If the self - repair fails, the system will enter the locked state and send an alarm signal through the monitoring system.
[0153] The present invention will be described in detail below with reference to the accompanying drawings through embodiments of technical decomposition.
[0154] Embodiment 1: A triple - core collaborative verification and action export execution method and system for soft - error protection of relay protection control devices
[0155] This embodiment discloses a triple - core collaborative verification and action export execution method and system for soft - error protection of relay protection control devices, which are divided into three modules as follows Figure 1 The first is the protection CPU2, which is responsible for core protection logic calculation, data self - inspection and verification, co - processing program transfer, and relay protection control export; the second is the startup CPU1, which is responsible for system startup, data self - inspection and verification, co - processing program transfer, and relay protection control export; the third is the co - processing CPU3, which is also the key module in the "triple - core collaborative architecture". The co - processing CPU3 is independent of the startup and protection CPUs and is responsible for receiving the ECC self - inspection results and status, and co - processing verification programs transferred from the protection CPU2 and the startup CPU1. At the same time, it is responsible for the consistency verification of the programs and data of the entire system, export arbitration decision - making, and part of the self - repair work.
[0156] In terms of architecture composition, the co - processing CPU3 includes a monitoring and inspection module 4 and an arbitration module 5. Further, the monitoring and inspection module 4 includes a receiving module 6, a co - verification module 7, a self - repair module 8, and a monitoring module 9; the arbitration module 5 includes a consistency judgment module 10, a decision - making module 11, and an exception handling module 12.
[0157] Its working process mainly includes the following steps:
[0158] Step 1: Self - inspection and program transfer of the startup CPU1 and the protection CPU2
[0159] During operation, the startup CPU1 and the protection CPU2 will perform real - time self - inspection on the operation data in their internal storage units and complete data integrity detection using the built - in ECC verification module.
[0160] Within each interrupt control cycle, the startup CPU1 and the protection CPU2 transfer programs to the co - processing CPU3 in a partition - round - robin manner. The program partitions of the startup CPU1 include an initialization code area, an interrupt - handling code area, a system management and scheduling code area, and a data management and monitoring code area. The program partitions of the protection CPU2 include a protection logic code area, a real - time data processing area, a parameter and setting management area, an interrupt - handling code area, and an invisible fault monitoring area.
[0161] Before starting CPU1 and protecting CPU2, the current data is backed up to the internal storage unit before transferring the program of each partition, so as to perform data self-repair in subsequent steps.
[0162] Step 2: Cooperative work of the receiving module and the co-verification module
[0163] The co-processing CPU3 consists of a monitoring and inspection module 4 and an arbitration module 5. Further, the monitoring and inspection module 4 includes a receiving module 6, a co-verification module 7, a self-repair module 8, and a monitoring module 9.
[0164] The receiving module 6 of the co-processing CPU3 receives the self-check status results transmitted by the starting CPU1 and the protecting CPU2. Part of these status results are transmitted to the consistency judgment module 10 in the arbitration module 5 for analysis, and the other part is directly transmitted to the monitoring module 9 in the monitoring and inspection module 4 for monitoring.
[0165] The co-verification module 7 receives the program code transmitted in a round-robin manner and compares it with the snapshot of the partition round-robin program during the previous round-robin. The co-verification module 7 generates a unique identifier (such as a hash value, etc.) for verification as the snapshot, and determines whether the program has been tampered with or has an error by comparing the current partition code and the unique identifier of the previous snapshot.
[0166] If the verification is consistent, the co-verification module 7 updates the snapshot and sends the status of passing the verification to the consistency judgment module 10; if the verification is inconsistent, an abnormal status is generated and transmitted to the consistency judgment module 10 and the monitoring module 9.
[0167] Step 3: Arbitration of the consistency judgment and decision-making module
[0168] The consistency judgment module 10 receives the status results from the receiving module 6 and the co-verification module 7, and analyzes whether the data and programs transmitted by the starting CPU1 and the protecting CPU2 are consistent during operation.
[0169] In the decision-making module 11, according to the 2-out-of-2 arbitration principle based on the consistency judgment, the output signals of the starting CPU1, the protecting CPU2, and the co-verification module 7 are compared. If the three signals are all consistent, an "allow export action" instruction is generated, and the system continues to run normally; if the signals are inconsistent, the system enters an abnormal state and is transmitted to the abnormal handling module 12.
[0170] Step 4: Abnormal handling and self-repair operations
[0171] After detecting the abnormal state, the abnormal handling module 12 generates a self-repair instruction, and the instruction is transmitted to the self-repair module 8.
[0172] After receiving the self - repair instruction, the self - repair module 8, on the one hand, starts to recover the damaged program from the backup program of the co - verification module 7; at the same time, it forwards the self - repair command to the startup CPU 1 and the protection CPU 2, so that the startup CPU 1 and the protection CPU 2 can recover the damaged data from their respective backup data.
[0173] After the self - repair is completed, the repair status is passed to the monitoring module 9 for updating the system status and ensuring the stable operation of the system.
[0174] Step 5: System arbitration and reliable relay output
[0175] When the 2 - out - of - 2 arbitration principle based on consistency judgment in the decision - making module 11 is completed, if all signals are consistent, the system will allow the relay to operate normally and drive the outlet relay to execute the action.
[0176] When the self - repair module 8 completes the repair task, it passes the repair status to the monitoring module. The monitoring module records this repair status so that in the next round of self - inspection process, the consistency judgment module can obtain the repaired self - inspection status from the monitoring module to ensure the normal operation of the subsequent system processes and the reliability of the relay protection control device.
[0177] Embodiment 2: Soft - error multi - level verification mechanism
[0178] This embodiment discloses a soft - error multi - level verification mechanism for a relay protection control device, as Figure 2 shown, mainly including two verification levels: the first level is data self - verification, and the second level is program co - verification. This multi - level verification mechanism aims to ensure the integrity and reliability of data and programs in the operation process of the relay protection control device, especially having higher protection capabilities when dealing with soft errors (such as single - bit flips and multi - bit flips).
[0179] Step 1: The first - layer data verification
[0180] The first - layer verification is completed by the startup CPU 1 and the protection CPU 2, mainly for real - time verification of the key data used by the system during operation.
[0181] Data verification process: The startup CPU 1 and the protection CPU 2 respectively perform self - inspection on the key data during operation through their built - in ECC verification modules. The self - inspected data includes initialization parameters, real - time measurement data, setting parameters, and log information, etc.
[0182] ECC verification: The ECC verification module generates a verification code (such as an ECC code, etc.) for each data block and compares it with the previously stored verification code to detect whether there are soft errors such as single - bit flips or multi - bit flips in the data block.
[0183] Data verification result transfer: After completing the self - inspection of data, start CPU1 and protection CPU2 to transfer the self - inspection results to the receiving module 6 of coprocessor CPU3 as basic information for further analysis. The receiving module records the results of data verification to provide reference for subsequent program verification.
[0184] Step 2: Second - layer program verification
[0185] The second - layer verification is carried out for the program code, adopting a partition - round - robin mechanism to ensure the integrity of the program during operation.
[0186] Program partition transfer and verification: The program address spaces of CPU1 and protection CPU2 are divided into multiple partitions, and the program code of each partition is regularly transferred to the co - verification module 7 of coprocessor CPU3 through a round - robin mechanism. Within each interrupt control cycle, CPU1 and protection CPU2 each transfer a program partition to the co - verification module for verification.
[0187] Snapshot generation and comparison: After the co - verification module 7 receives the transferred program partition, it generates a unique identifier (such as a hash value, etc.) of the partition as a snapshot and compares it with the snapshot of the same partition transferred last time.
[0188] Feedback of verification result: Verification passed: If the snapshot comparison is consistent, it means that the program of this partition has no error. The co - verification module 7 will update the snapshot of the current partition and feedback the verification - passed status to the monitoring module 9 and the consistency judgment module 10. Verification failed: If the snapshot comparison is inconsistent, the co - verification module will record the abnormal status of this partition and transfer this information to the monitoring module 9 and the consistency judgment module 10.
[0189] Step 3: Result evaluation of multi - level verification
[0190] In the multi - level verification mechanism, the verification results of the first layer and the second layer are independent of each other, but jointly provide the system with comprehensive error - detection capabilities. The receiving module 6 and the co - verification module 7 respectively summarize the results of data verification and program verification to ensure that the verification results at different levels of the system are consistent.
[0191] Linkage between data and program verification: The self - inspection result of data in the first layer provides basic health - status information for the system operation, while the program verification in the second layer ensures that the program code has no error during execution. The combination of the two can enhance the soft - error protection ability at both the data and program levels.
[0192] Embodiment 3: A system arbitration and self - repair method for a relay protection control device
[0193] This embodiment discloses a system arbitration and self - repair method for a relay protection control device, mainly realizing the abnormal handling and self - repair process of programs and data. This method performs arbitration and self - repair operations through a coprocessor CPU, ensuring that the system can automatically recover when soft errors or exceptions occur, and guaranteeing the normal operation of the relay protection control device. The specific process is as follows Figure 3 shown.
[0194] Step 1: Data and program backup
[0195] Data backup: Before starting the self - check of CPU1 and protection CPU2, the running data will be backed up. The backed - up data includes initialization parameters, real - time measurement data, setting parameters, etc. These data are saved in their respective internal storage units to ensure that in case of data anomalies, the backed - up data can be called for recovery.
[0196] Program backup: After CPU1 and protection CPU2 start to transfer the program to the co - verification module 7 in a partition - by - partition polling manner, the co - verification module 7 will back up the transferred program. The program is stored in the co - verification module 7 in a partitioned manner to ensure that in case of an exception, the self - repair module 8 can quickly recover the program at the partition level, avoiding large - scale program damage.
[0197] Step 2: Verification of data and programs
[0198] Data verification: CPU1 and protection CPU2 start to verify the running data respectively through the built - in ECC verification module. The verification results after self - check will be transferred to the receiving module 6 of the coprocessor CPU. This module aggregates the data verification results to provide basic data for the export arbitration decision.
[0199] Program verification: CPU1 and protection CPU2 transfer the program to the co - verification module 7 of the coprocessor CPU in a partition - by - partition polling manner. The co - verification module 7 verifies the transferred program partitions, generates a unique identifier (such as a hash value) of the program partition as a snapshot, and compares it with the snapshot of the same partition transferred in the previous polling to check for anomalies. The verification results (pass or fail) are transferred to the consistency judgment module 10 for the arbitration module to judge.
[0200] Step 3: Consistency judgment and arbitration
[0201] Consistency judgment module 10: The receiving module 6 and the co - verification module 7 transfer the data verification results and program verification results to the consistency judgment module 10. The consistency judgment module 10 combines the post - self - repair status to compare the verification status of the receiving module 6 and the co - verification module 7, and judges whether the three are consistent.
[0202] Consistent signal output: If the verification results of the three are consistent, the consistency judgment module generates a consistent signal and transmits this signal to Decision Module 11, indicating that the system can operate normally. Inconsistent signal output: If the verification results are inconsistent, an inconsistent signal is generated, an alarm is output, and the device is locked out.
[0203] Decision Module 11 (based on the 2-out-of-2 principle of consistency judgment): Decision Module 11 arbitrates based on the signal from the consistency judgment module. According to this principle, the decision module evaluates the relay permission export status in Startup CPU 1, Protection CPU 2, and Consistency Judgment Module 10. If the three are consistent, the relay is allowed to operate, and a "permission to export action" signal is output to Actuating Relay 16, and the system continues to operate normally. If an abnormality occurs, the abnormal handling process is entered.
[0204] Step 4: Abnormal handling and self-repair
[0205] Abnormal handling module: When an inconsistent signal is detected, Abnormal Handling Module 12 is activated. The abnormal handling module generates a self-repair instruction and transmits the instruction to Self-Repair Module 8. This module is responsible for directing subsequent self-repair operations.
[0206] Operations of the self-repair module: Repair Startup CPU and Protection CPU: After receiving the self-repair instruction from Abnormal Handling Module 12, Self-Repair Module 8 calls the backup data or program to overwrite the damaged program. At the same time, a self-repair instruction is sent to Startup CPU 1 and Protection CPU 2, enabling them to recover the corresponding abnormal partitions through the backup data.
[0207] Self-repair status feedback: After self-repair is completed, the repair status is fed back to Monitoring Module 9 for providing the status of the repaired data or program during subsequent verification.
[0208] Step 5: Monitoring of self-repair status and continuous operation
[0209] Monitoring Module 9: After self-repair is completed, Monitoring Module 9 continuously monitors the health status of the system. This module receives feedback information from Receiving Module 6, Co-Verification Module 7, and Self-Repair Module 8 in real time, and can identify latent system faults while ensuring real-time update of the system status.
[0210] Next round of self-check: In the next interrupt control cycle, Consistency Judgment Module 10 can obtain the repaired self-check status from Monitoring Module 9 and continue with normal self-check and arbitration operations to ensure continuous and stable operation of the system.
[0211] Embodiment 4: Driving and execution circuit of the relay
[0212] This embodiment discloses a driving and executing circuit for a relay. The relay system includes parts such as a CPU system 13, a relay driving module 14, an actuator 15, an action relay 16, a start relay 17, and an operation box 18, which will be described in detail in combination with Figure 4 for elaboration.
[0213] Step 1: Provision of the starting power supply for the relay
[0214] The starting power supply provides voltage support for the entire driving circuit through the start relay 17. When the system determines that an operation needs to be executed, the starting power supply is powered through the start relay 17 to ensure that the circuit can execute the operation smoothly.
[0215] Step 2: Logic control of relay driving
[0216] The CPU system receives the allowed outlet action signal from the decision-making module 11 through the relay driving module and generates a control signal according to the result of logical judgment. This control signal is used to control the action states of the start relay 17 and the action relay 16 to ensure that the relay driving logic is executed according to the design.
[0217] Step 3: Driving and operation of the relay
[0218] When the relay driving module controls the start relay 17 to first connect the start circuit, the action relay 16 obtains power support. Subsequently, the action relay 16 acts according to the set logic to control the operation box to execute the subsequent operation process. This process ensures that the driving and operation can be effectively linked according to the system design.
[0219] Step 4: Action control of the actuator
[0220] After receiving the action signal of the relay, the actuator 15 performs corresponding mechanical operations.
[0221] In summary, through multi-level data and program verification and combined with an automatic self-repair mechanism, the present invention significantly improves the anti-soft error ability of the relay protection control device and ensures the high reliability and security of the system.
[0222] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A collaborative verification and reliable action system for protecting control device soft error anti-misoperation, characterized in that, The system includes a startup CPU, a protection CPU, and a coprocessor CPU built into the protection control device. Through the mutual cooperation among the three, a multi-level verification and self-repair mechanism is formed to achieve soft error protection for the system and execution of the action outlet. The system conducts multi-level verification through the startup CPU, the protection CPU, and the coprocessor CPU, including: performing data self-verification by the startup CPU and the protection CPU respectively, and transmitting the verification results to the coprocessor CPU; performing program co-verification by the coprocessor CPU. Meanwhile, the coprocessor CPU makes an arbitration decision based on the verification results to determine the system outlet action or perform exception handling. The self-repair mechanism specifically includes: triggering a self-repair instruction during the exception handling, and the system conducts self-repair processing until it works normally. The startup CPU is responsible for system startup, data self-check, and startup control of the relay protection control outlet. It has a built-in RAM ECC verification module for real-time self-verification of the operating data of the startup CPU. The protection CPU is responsible for core relay protection control logic calculation. It has a built-in RAM ECC verification module for real-time self-verification of the operating data of the protection CPU. Both the startup CPU and the protection CPU are equipped with a data backup module and a data self-repair module, which are used to back up and repair the operating data of the startup CPU / protection CPU respectively. Before performing data self-verification and transmitting the program to the coprocessor CPU, the startup CPU and the protection CPU back up the operating data to the corresponding data backup modules respectively. The coprocessor CPU includes a monitoring and verification module and an arbitration module. The monitoring and verification module is used to verify and monitor the data and programs of the startup CPU and the protection CPU. The arbitration module is used to determine the system outlet action or perform exception handling according to the output of the monitoring and verification module.
2. The collaborative verification and reliable action system for protecting control device soft error anti-misoperation according to claim 1, characterized in that, The monitoring and verification module includes a receiving module, a co-verification module, a self-repair module, and a monitoring module. The receiving module is used to receive the self-verification results of the startup CPU and the protection CPU respectively, and transmit them to the arbitration module and the monitoring module simultaneously. The co-verification module is used to verify the programs transmitted by the startup CPU and the protection CPU respectively, and transmit the verification results to the arbitration module and the monitoring module simultaneously. The self-repair module is used to receive the self-repair instruction, perform self-repair on the program, and send the self-repair instruction to the startup CPU and the protection CPU to trigger their respective self-repair functions. The monitoring module is used to receive and store the feedback information of the receiving module, the co-verification module, and the self-repair module in real time to update the system status in real time. If it monitors that the self-repair status of each module is self-repair successful, it sends a recovery instruction to the arbitration module, and the arbitration module controls the system status to return to normal. Otherwise, it sends an exception instruction to the arbitration module, and the arbitration module generates a blocking signal to control the system to enter the blocking state and issues an alarm signal through the monitoring system.
3. The collaborative verification and reliable action system for protecting against soft errors of a control device according to claim 2, characterized in that, The startup CPU and the protection CPU transfer the program to the co-verification module in a partition-by-round-robin manner, and in each interrupt control cycle, transfer the program of one partition in sequence, and ensure that all program partitions are transmitted within the specified time.
4. The collaborative verification and reliable action system for protecting against soft errors of a control device according to claim 3, characterized in that, After receiving the program transferred by partition-by-round-robin, the co-verification module generates a unique identifier for each partition as a snapshot, and compares it with the snapshot of the previous time for each partition one by one, and outputs the verification result; at the same time, the co-verification module updates the snapshot of the current partition; The verification result includes passing or not passing. If the comparison is consistent, the verification result is passing, otherwise the verification result is not passing; At the same time, after receiving the program transferred by partition-by-round-robin, the co-verification module also stores and backs up the program in a partition manner for the self-repair module to call; after receiving the self-repair instruction, the self-repair module performs self-repair by calling the program backed up in the co-verification module.
5. The collaborative verification and reliable action system for protecting against soft errors of a control device according to claim 2, characterized in that, The monitoring module is also used to identify latent system faults and give an alarm while ensuring real-time update of the system state.
6. The collaborative verification and reliable action system for protecting control device soft error prevention and misoperation according to claim 2, characterized in that The arbitration module includes a consistency judgment module, a decision-making module and an exception handling module; The consistency judgment module is used to receive the signal data output by the receiving module and the co-verification module, and respectively judge whether the self-verification result of the startup CPU is consistent with the verification result of the co-verification module for the startup CPU program, and whether the self-verification result of the protection CPU is consistent with the verification result of the co-verification module for the protection CPU program; If the verification results are all consistent, output a consistency signal to the decision-making module, otherwise generate an abnormal state signal to the decision-making module; The decision-making module arbitrates according to the principle of taking multiple and getting multiple. Specifically, when the startup relays and action relays corresponding to the startup CPU and the protection CPU are both in the "allowed to output" state, and the consistency judgment module outputs a consistency signal and generates an "allowed to export action" instruction, control each relay to continue to work normally; otherwise, send an activation instruction to the exception handling module; The exception handling module is used to generate a self-repair instruction after receiving the activation instruction and transfer the self-repair instruction to the self-repair module.
7. The collaborative verification and reliable action system for protecting against soft errors in a control device according to claim 6, characterized in that In the new interrupt control cycle, the consistency judgment module obtains the repaired / updated system state from the monitoring module and continues to perform normal self-check and arbitration operations to ensure the continuous and stable operation of the system.
8. A collaborative verification and reliable action method for preventing soft errors in a protection control device based on the system according to any one of claims 1 to 7, characterized in that, The method includes the following steps: Step 1, the startup CPU and the protection CPU respectively perform real-time data self-verification, and send the verification results to the co-processing CPU. At the same time, the startup CPU and the protection CPU send their respective programs to the co-processing CPU; Step 2, the co-processing CPU processes the received program, generates a corresponding snapshot, and compares the snapshot with the previous snapshot to generate a program verification result; Step 3: The coprocessor CPU determines whether the self-check result of the startup CPU in Step 1 is consistent with the verification result of the startup CPU program by the coprocessor CPU in Step 2, and whether the self-check result of the protection CPU in Step 1 is consistent with the verification result of the protection CPU program by the coprocessor CPU in Step 2. If both are consistent, a consistency signal is generated; otherwise, an abnormal status signal is generated. Step 4: The coprocessor CPU makes a blanking decision based on the signal generated in Step 3. If the consistency signal is output in Step 3, and the startup relay and the action relay corresponding to the startup CPU and the protection CPU are both in the "allowed output" state, then the coprocessor CPU generates an "allowed outlet action" instruction to control each relay to continue normal operation, including driving the outlet relay to execute an action; otherwise, the coprocessor CPU generates an exception handling signal and proceeds to Step 5. Step 5: A self-repair instruction is generated to trigger the self-repair functions of the startup CPU, the protection CPU, and the coprocessor CPU. At the same time, the self-repair result is monitored in real time. If the self-repair result is successful, the snapshot data is updated, and Step 1 is fed back and continued to be executed; otherwise, it proceeds to Step 6. Step 6: Enter the locked state, issue an alarm signal, and end the process at the same time.
Citation Information
Patent Citations
A relay protection method and system based on a multi-core processor chip
CN116631492B
Unified, adaptive RAS for hybrid systems
CN103582874A
Method and system for detecting and recovering memory bit flipping in secondary power equipment
WO2021208341A1