Heterogeneous redundancy control method suitable for multi-core embedded processor

By adopting the heterogeneous redundancy control method of AMP and checkpoint rollback recovery technology in multi-core processors, the problems of resource waste and poor isolation of traditional redundancy methods are solved, and effective fault tolerance and rapid recovery of transient and permanent faults are achieved.

CN120780518APending Publication Date: 2025-10-14HARBIN INST OF TECH +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510859509.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Traditional hardware redundancy methods waste resources seriously and are susceptible to common cause failures, while software redundancy methods have poor isolation and cannot effectively deal with transient and permanent failures of multi-core processors.

Method used

A heterogeneous redundant control solution based on AMP and checkpoint rollback recovery technology is adopted. The self-check module, synchronization processing module and synchronization comparison module in the heterogeneous redundant control system are utilized. Through task-level synchronization and checkpoint rollback recovery technology, core-level redundant control of multi-core processors is realized, reducing hardware resource waste and improving isolation.

Benefits of technology

It achieves effective fault tolerance for transient and permanent faults in multi-core processors, reduces fault recovery time and hardware resource waste, and improves the real-time performance and isolation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780518A_ABST
    Figure CN120780518A_ABST
Patent Text Reader

Abstract

The invention relates to a heterogeneous redundancy control method suitable for a multi-core embedded processor. The invention relates to the technical field of fault tolerance of embedded multi-core processors, in particular to a core-level redundancy method of a multi-core processor based on an asymmetric multi-core technology, the AMP technology can provide good isolation for core-level redundancy, and the problem that software redundancy cannot deal with permanent faults is solved. Compared with module-level redundancy, the method has the advantages that a check point rollback recovery mode is adopted in response to transient faults, switching and check-in are not needed, and the probability of faults possibly caused in the switching process is reduced. According to the invention, multiple cores in the processor module are used as redundant bodies, so that the waste of hardware space resources is reduced to a great extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of embedded multi-core processor fault tolerance, and is a heterogeneous redundancy control method applicable to multi-core embedded processors. Background Art

[0002] An embedded microprocessor is a computer chip used within a variety of devices and equipment to provide additional functionality. A microprocessor is a digital electronic component whose transistors are integrated on a single semiconductor IC, resulting in a compact and low-power design. Embedded systems have a wide range of applications and are an integral part of modern technology and intelligent advancements.

[0003] The continuous advancement of integrated circuit manufacturing processes has posed a serious threat to the reliability of microprocessor computing due to transient faults. This is particularly true at the ultra-deep submicron level, where the probability of transient faults is significantly increased. Furthermore, the rapid development of computer architecture has ushered in the multi-core era for microprocessors. Transient faults can be caused by factors such as radiation and signal interference. These non-physical faults, which occur within the chip, can affect the correctness of program execution. Therefore, in-depth research on fault-tolerance technologies for multi-core computing platforms is necessary.

[0004] Redundancy methods are categorized into hardware redundancy and software redundancy. Traditional hardware redundancy implements redundancy at the module level, requiring significant space resources. Each redundant element results in significant resource waste, and homogeneous redundant systems are susceptible to common cause failures. Software redundancy utilizes replicated processes as redundant elements, reducing the need for hardware resources but also reducing the isolation of the redundant system. While it offers better fault tolerance for transient failures, software redundancy becomes ineffective if a permanent hardware failure occurs. Summary of the Invention

[0005] Aiming to overcome the deficiencies of traditional hardware and software redundancy, the present invention proposes a heterogeneous redundancy control scheme for multi-core embedded processors based on AMP and checkpoint rollback recovery technology, which reduces hardware costs and has better fault tolerance for transient and permanent faults.

[0006] The present invention provides the following technical solutions: Step 1: The embedded processor and FPGA start up. The FPGA completes self-test and processor detection. If no faults are detected, it enters the working mode. The embedded processor core starts the Linux, RT-thread, and FreeRTOS operating systems respectively and executes the same task program and AMPcrr. The other core directly executes the bare metal program. Step 2: The four processor cores use task-level synchronization to synchronize tasks. When the program reaches a checkpoint, each core broadcasts a task synchronization frame. After receiving the synchronization frames of the remaining modules, the synchronization ends and the task continues. If no module synchronization frame is received within a timeout, the synchronization frame is recorded and sent to the voting module, which saves the result. Step 3: The FPGA selects three outputs and marks them as active outputs, and one output as backup output. Based on the comparison results of the four outputs, they are marked as Class A operation state, Class A reconstruction state, Class B operation state, Class B reconstruction state, Class C operation state, and Class C reconstruction state. After power-on, the system defaults to Class A operation state. Step 4: Each core program runs to the output comparison node. AMPcrr enables the checkpointsave() function to save the running status of each core program as a checkpoint file. The FPGA collects four outputs and synchronously compares the output results of the three outputs marked as current outputs.

[0007] Preferably, a two-out-of-three vote is performed on the calculation results output on duty, and a unique calculation result is output, and the output unique result is compared with the backup output result: When the output of the backup core is consistent with the voting result, its status is considered normal and it continues to operate as a backup core in Class A status. The FPGA's working status indicator remains unchanged. When the backup core is inconsistent with the voting result, and the degree of inconsistency reaches the system-set threshold, the FPGA sends a reconstruction command to the backup core. The backup core executes the restore() function, parses the checkpoint saved in the secure memory, and restores the working state of the backup core to the last normal state. The FPGA identifies the processor working state as a Class A reconstruction state. When the backup core can be successfully restored, the processor working state returns to the A-level operation state. Otherwise, the FPGA blocks the output of the backup core and sends a shutdown command to shut down the system running the backup core; The FPGA selects two of the remaining three outputs as on-duty outputs and the other as a backup output, marking the processor operating state as level B.

[0008] Preferably, the on-duty core exception handling strategy is determined. In the Class A operating state, inconsistent calculation results may occur between on-duty outputs. In this case, the system performs the following processing according to the specific abnormal situation: When the result of any of the three is inconsistent with the other two, and the degree reaches the set threshold, the FPGA sends the reconstruction command of the core, the core executes the restore() function, and the working state is restored to the last normal state. The FPGA marks the processor working state as A-level reconstruction state. When the backup core result is inconsistent with the voting result at this time, the backup core also performs reconstruction operation, and the FPGA marks the system as B-level reconstruction state; When the results of the three are all inconsistent, if the backup output result is consistent with the result of any of the three, the other two abnormal decision cores are reconstructed, and the system enters C-level running state; When the result of the backup core is inconsistent with the results of all the three, only one core continues to work, and the other three cores participate in reconstruction, and the system enters C-level running state.

[0009] Preferably, in the B-level running state, when the core performing the reconstruction operation can successfully restore to the last working state, the FPGA re-sets the processor state to the A-level running state; when the core performing the reconstruction operation cannot successfully restore to the last working state, the FPGA masks the core output and sends a shutdown instruction, and the system enters the two-out-of-three working mode; After the system is degraded to the two-out-of-three working mode, the FPGA synchronously performs two-out-of-three voting on the outputs of the three cores, outputs the correct voting result, and when any core is inconsistent with the other cores, the core is reconstructed and restored, and the system is marked as B-level reconstruction state; when the outputs of the three cores are all inconsistent, two cores are selected for reconstruction operation, and the system is marked as C-level reconstruction state; In the B-level reconstruction state, when the core performing the reconstruction operation can successfully restore to the last working state, the FPGA re-sets the processor state to the B-level running state; when the core performing the reconstruction operation cannot successfully restore to the last working state, the FPGA masks the core output and sends a shutdown instruction, and the system enters the two-machine working mode and is converted to C-level working state; In the C-level working state, the FPGA starts the single-machine fault detection module to monitor the health state of the active core in real time.

[0010] Preferably, when the cleaning and reconstruction are not completed, and any of the decision CPUs is inconsistent with the other two, the cleaning and reconstruction of the abnormal CPU are triggered, and the system enters the B-level running state; When the results of the three decision CPUs are all inconsistent, one decision CPU is reserved, and the other two decision CPUs perform cleaning and reconstruction, and the system enters C-level running state.

[0011] Preferably, in the B-level running state: If the cleaning CPU successfully completes the reconstruction, it is added to the decision group, and the system returns to the A-level running state; If the cleaning CPU fails to reconstruct multiple times, the backup CPU joins the judgment group, and the cleaning CPU continues to try to reconstruct, and the system enters the A-level reconstruction state; If the backup CPU and the voting result are inconsistent, the backup CPU performs cleaning and reconstruction; If the results of the two judgment CPUs are inconsistent, and the backup CPU agrees with one of them, the abnormal judgment CPU is cleaned and reconstructed; If the backup CPU is inconsistent with both decision CPUs, only one decision CPU is kept working, and the remaining CPUs perform cleaning and reconstruction, and the system enters the C-level operation state.

[0012] Preferably, in the Class C operating state: If the CPU is cleaned and the reconstruction is completed, the system status is adjusted according to the number of successful reconstructions: Reconstruct a CPU and the system enters the B-level operation state; After reconstructing the two CPUs, the system enters the A-level operation state.

[0013] If the clean CPU fails to reconstruct multiple times, the backup CPU is added to the judgment group and the system status is adjusted.

[0014] In the C-level refactoring state: The system checks the status of all current CPUs. If the status of three or four CPUs is consistent, the system is adjusted to Class B or Class A operating status based on the consistency. If there are less than two consistent CPUs, the system remains in the C-level reconstruction state and pauses outputting calculation results.

[0015] A heterogeneous redundant control system applicable to a multi-core embedded processor, the system comprising: A self-test module is started by the embedded processor and FPGA. The FPGA completes self-test and processor detection and enters the working mode after no fault is found. The embedded processor core starts the Linux, RT-thread, and FreeRTOS operating systems respectively and executes the same task program and AMPcrr. Another core directly executes the bare metal program; A synchronization processing module, which synchronizes the four processor cores using task-level synchronization. When the program reaches a checkpoint, each core broadcasts a task synchronization frame. After receiving the synchronization frames of the remaining modules, the synchronization ends and the task continues. If no module synchronization frame is received within a timeout, the synchronization frame is recorded and sent to the voting module, which saves the result. An output module, wherein the output module selects three outputs of the FPGA and marks them as on-duty outputs and another output as a backup output, and marks them as level A operation state, level A reconstruction state, level B operation state, level B reconstruction state, level C operation state, and level C reconstruction state based on the comparison results of the four outputs. The system is in level A operation state by default after power-on; A synchronous comparison module runs each core program to the output comparison node. AMPcrr enables the checkpointsave() function to save the running status of each core program as a checkpoint file. FPGA collects four outputs and synchronously compares the output results of three of them marked as on-duty outputs.

[0016] A computer-readable storage medium stores a computer program, which is executed by a processor to implement a heterogeneous redundancy control method applicable to a multi-core embedded processor.

[0017] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, a heterogeneous redundancy control method applicable to a multi-core embedded processor is implemented.

[0018] The present invention has the following beneficial effects: This invention proposes a core-level redundancy method for multi-core processors based on asymmetric multicore processor (AMP) technology. AMP technology provides superior isolation for core-level redundancy, resolving the issue of software redundancy's inability to cope with permanent failures. Compared to module-level redundancy, this invention utilizes a checkpoint rollback recovery method to address transient failures, eliminating the need for a switchover and reducing the potential for failures during the switchover process. This invention utilizes the multiple cores within the processor module for redundancy, significantly reducing the waste of hardware space resources.

[0019] The present invention divides system resources into heterogeneous multi-module redundant partitions. Each core occupies a certain memory space and does not interfere with each other. A local failure occurring on a core cannot affect other parts of the system, and the system has strong isolation.

[0020] The present invention runs a time-sharing operating system (Linux) and a real-time operating system (RT-Thread) on different cores respectively, greatly improving the real-time performance of the system while ensuring interactivity.

[0021] The present invention designs a stable, scalable and fast checkpoint recovery (C / R) method. When a core operation fails, there is no need to restart the entire system for repair. Only the core program needs to be rolled back to the previous checkpoint, which greatly saves the failure recovery time.

[0022] The present invention adopts a multi-machine parallel redundant voting output method to output the correct system calculation results at the same time as fault detection. It will not affect the system output during the entire process of fault detection, positioning and single-core recovery, and will minimize the increase in time cost compared to other redundancy methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 Shown is a schematic diagram of the four-module synchronization technology based on synchronization frames of the present invention; Figure 2 Shown is a schematic diagram of the operating state of the embedded quad-core processor of the present invention; Figure 3 Shown is a schematic diagram of the checkpoint rollback recovery mechanism of the present invention. DETAILED DESCRIPTION

[0025] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] The present invention is described in detail below with reference to specific embodiments. Specific embodiment one: according to Figures 1 to 3 As shown, the specific optimization technical solution adopted by the present invention to solve the above technical problems is: the present invention relates to a heterogeneous redundancy control method suitable for multi-core embedded processors.

[0028] The present invention provides a heterogeneous redundancy control method applicable to a multi-core embedded processor, the method comprising the following steps: Step 1: The embedded processor and FPGA start up. The FPGA completes self-test and processor detection. If no faults are detected, it enters the working mode. The embedded processor core starts the Linux, RT-thread, and FreeRTOS operating systems respectively and executes the same task program and AMPcrr. The other core directly executes the bare metal program. Step 2: The four processor cores use task-level synchronization to synchronize tasks. When the program reaches a checkpoint, each core broadcasts a task synchronization frame. After receiving the synchronization frames of the remaining modules, the synchronization ends and the task continues. If no module synchronization frame is received within a timeout, the synchronization frame is recorded and sent to the voting module, which saves the result. Step 3: The FPGA selects three outputs and marks them as active outputs, and one output as backup output. Based on the comparison results of the four outputs, they are marked as Class A operation state, Class A reconstruction state, Class B operation state, Class B reconstruction state, Class C operation state, and Class C reconstruction state. After power-on, the system defaults to Class A operation state. Step 4: Each core program runs to the output comparison node. AMPcrr enables the checkpointsave() function to save the running status of each core program as a checkpoint file. The FPGA collects four outputs and synchronously compares the output results of the three outputs marked as current outputs. Specific embodiment two: The difference between the second embodiment of the present application and the first embodiment is that: Perform a two-out-of-three vote on the current output operation results, output a unique operation result, and compare the output unique result with the backup output result: When the output of the backup core is consistent with the voting result, its status is considered normal and it continues to operate as a backup core in Class A status. The FPGA's working status indicator remains unchanged. When the backup core is inconsistent with the voting result, and the degree of inconsistency reaches the system-set threshold, the FPGA sends a reconstruction command to the backup core. The backup core executes the restore() function, parses the checkpoint saved in the secure memory, and restores the working state of the backup core to the last normal state. The FPGA identifies the processor working state as a Class A reconstruction state. When the backup core can be successfully restored, the processor working state returns to the A-level operation state. Otherwise, the FPGA blocks the output of the backup core and sends a shutdown command to shut down the system running the backup core; The FPGA selects two of the remaining three outputs as on-duty outputs and the other as a backup output, marking the processor operating state as level B. Specific embodiment three: The only difference between the third embodiment of the present application and the second embodiment is that: Determine the on-duty core exception handling strategy. In the A-level operating state, inconsistent calculation results may occur between on-duty outputs. At this time, the system performs the following processing based on the specific exception situation: When any of the calculation results in the value output is inconsistent with the other two, and the degree reaches the set threshold, the FPGA sends a reconstruction command to the core. The core executes the restore() function to restore the working state to the previous normal state. The FPGA marks the processor working state as the A-level reconstruction state. When the result of the backup core is inconsistent with the voting result at this time, the backup core also performs the reconstruction operation, and the FPGA marks the system as the B-level reconstruction state. If the calculation results of the three on-duty outputs are inconsistent, if the calculation result of the backup output is consistent with the result of any on-duty output, the other two abnormal judgment cores will be reconstructed and the system will enter the C-level operation state; When the calculation results of the backup core are inconsistent with the results of all the on-duty cores, only one core is retained to continue working, and the other three cores participate in reconstruction, and the system enters the C-level operation state. Specific embodiment four: The only difference between the fourth embodiment of the present application and the third embodiment is that: In the B-level operating state, if the core undergoing the reconstruction operation can successfully recover to the previous working state, the FPGA will restore the processor state to the A-level operating state. If the core undergoing the reconstruction operation cannot successfully recover to the previous working state, the FPGA will block the core output and send a shutdown command, and the system will switch to the three-out-of-two operating mode. After the system is downgraded to the 2-out-of-3 operating mode, the FPGA performs a synchronous 2-out-of-3 vote on the outputs of the three cores and outputs the correct voting result. If the output of any core is inconsistent with the other cores, it is reconstructed and recovered, and the system is marked as a Class B reconstruct state. If the outputs of all three cores are inconsistent, two cores are selected for reconstructing, and the system is marked as a Class C reconstruct state. In the B-level reconfiguration state, if the core undergoing the reconfiguration operation can successfully recover to the previous working state, the FPGA restores the processor state to the B-level operation state. If the core undergoing the reconfiguration operation cannot successfully recover to the previous working state, the FPGA blocks the core output and sends a shutdown command, and the system switches to dual-machine working mode and transforms to the C-level working state. In the C-level working state, the FPGA starts the single-machine fault detection module to monitor the health status of the on-duty core in real time. Specific embodiment five: The only difference between the fifth embodiment of the present invention and the fourth embodiment is that: When cleaning and reconstruction are not completed, and the results of any CPU in the judgment are inconsistent with the other two CPUs, the cleaning and reconstruction of the abnormal CPU is triggered, and the system enters the B-level operation state; When the results of the three judgment CPUs are inconsistent, one judgment CPU is retained, and the other two judgment CPUs perform cleaning and reconstruction, and the system enters the C-level operation state. Specific embodiment six: The only difference between the sixth embodiment of the present invention and the fifth embodiment is that: In Class B operating state: If the cleaning CPU successfully completes the reconstruction, it is added to the judgment group and the system returns to the Class A operating state; If the cleaning CPU fails to reconstruct multiple times, the backup CPU joins the judgment group, and the cleaning CPU continues to try to reconstruct, and the system enters the A-level reconstruction state; If the backup CPU and the voting result are inconsistent, the backup CPU performs cleaning and reconstruction; If the results of the two judgment CPUs are inconsistent, and the backup CPU agrees with one of them, the abnormal judgment CPU is cleaned and reconstructed; If the backup CPU is inconsistent with both decision CPUs, only one decision CPU is kept working, and the remaining CPUs perform cleaning and reconstruction, and the system enters the C-level operation state. Specific embodiment seven: The only difference between the seventh embodiment of the present invention and the sixth embodiment is that: In C-level operating state: If the CPU is cleaned and the reconstruction is completed, the system status is adjusted according to the number of successful reconstructions: Reconstruct a CPU and the system enters the B-level operation state; After reconstructing the two CPUs, the system enters the A-level operation state.

[0035] If the clean CPU fails to reconstruct multiple times, the backup CPU is added to the judgment group and the system status is adjusted.

[0036] In the C-level refactoring state: The system checks the status of all current CPUs. If the status of three or four CPUs is consistent, the system is adjusted to Class B or Class A operating status based on the consistency. If there are less than two consistent CPUs, the system remains in the C-level reconstruction state and pauses outputting calculation results. Specific embodiment eight: The only difference between the eighth embodiment of the present invention and the seventh embodiment is that: The present invention provides a heterogeneous redundant control system applicable to a multi-core embedded processor, the system comprising: A self-test module is started by the embedded processor and FPGA. The FPGA completes self-test and processor detection and enters the working mode after no fault is found. The embedded processor core starts the Linux, RT-thread, and FreeRTOS operating systems respectively and executes the same task program and AMPcrr. Another core directly executes the bare metal program; A synchronization processing module, which synchronizes the four processor cores using task-level synchronization. When the program reaches a checkpoint, each core broadcasts a task synchronization frame. After receiving the synchronization frames of the remaining modules, the synchronization ends and the task continues. If no module synchronization frame is received within a timeout, the synchronization frame is recorded and sent to the voting module, which saves the result. An output module, wherein the output module selects three outputs of the FPGA and marks them as on-duty outputs and another output as a backup output, and marks them as level A operation state, level A reconstruction state, level B operation state, level B reconstruction state, level C operation state, and level C reconstruction state based on the comparison results of the four outputs. The system is in level A operation state by default after power-on; A synchronous comparison module runs each core program to the output comparison node. AMPcrr enables the checkpointsave() function to save the running status of each core program as a checkpoint file. FPGA collects four outputs and synchronously compares the output results of three of them marked as on-duty outputs. Specific embodiment nine: The only difference between the ninth embodiment of the present invention and the eighth embodiment is that: The present invention provides a computer-readable storage medium on which a computer program is stored. The program is executed by a processor to implement a heterogeneous redundancy control method applicable to a multi-core embedded processor. Specific embodiment ten: The only difference between the tenth embodiment of the present invention and the ninth embodiment is that: The present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements a heterogeneous redundancy control method applicable to a multi-core embedded processor when executing the computer program. Specific embodiment eleven: The only difference between the eleventh embodiment of the present invention and the tenth embodiment is that: The purpose of the present invention is to overcome the deficiencies of traditional hardware and software redundancy and propose a heterogeneous redundancy design suitable for multi-core embedded processors based on AMP and checkpoint rollback recovery technology, which reduces hardware costs and has better fault tolerance for transient faults and permanent faults.

[0041] To achieve the above object, the present invention provides the following technical solutions: First, a quad-core embedded processor is used, leveraging asymmetric multi-core processor (AMP) technology. This multi-core processor can run different operating systems simultaneously on different cores. Heterogeneous operating systems can be run in parallel on the quad-core embedded processor, with cores 0 through 3 running Linux, RT-Thread, FreeRTOS, and bare-metal programs, respectively. Each operating system has its own memory, preventing interference and providing good isolation. The same task program runs on different operating systems (bare-metal), utilizing FPGAs for synchronous control and output comparison, enabling effective fault detection, location, and fault tolerance.

[0042] Secondly, a four-module dynamic redundant fault-tolerant control algorithm is deployed on the FPGA. This algorithm synchronously compares the four outputs transmitted by the processor, selects the most identical results as the correct output, and performs operations such as cleaning and recovery on inconsistent outputs. Unlike traditional two-out-of-three and two-out-of-four methods, this algorithm selects three outputs as the active outputs for synchronous comparison and one output as a backup. Based on the comparison results of the four core outputs, the operating state is defined and fault detection and reconstruction and recovery operations are performed. Each operation causes the operating state to shift to one of the defined states.

[0043] Thirdly, a checkpoint rollback recovery software AMPcrr is developed for embedded operating systems. Checkpoint() and restore() functions are added to key nodes of the program running inside the embedded processor. The checkpoint() function freezes the process and collects information about the process, including the stack information of the registers, and saves it as an image file. After the FPGA detects inconsistency in the output of a single core, it controls the core to start the restore() operation, reads the checkpoint image file containing the stack information of the previous state stored in the memory, replaces the current stack with the stack in the file, and thus rolls back to the checkpoint where the previous running state was normal to achieve fault tolerance. Specific implementation method 12: The present invention provides a heterogeneous redundancy control method applicable to a multi-core embedded processor, the method comprising the following steps: Step 1: The embedded processor and FPGA are started. The FPGA completes self-test and processor detection. If no faults are detected, it enters the working mode. The embedded processor core starts the Linux, RT-thread, and FreeRTOS operating systems respectively and executes the same task program and AMPcrr software. The other core directly executes the bare metal program. Step 2: The FPGA sends millisecond-level synchronization signals and receives the four outputs of the processor to synchronize the quad-core processors and ensure that the four heterogeneous operating systems execute tasks in a synchronized manner. Step 3: FPGA selects three outputs and marks them as on-duty outputs, and marks another output as backup output. Based on the comparison results of the four outputs, they are marked as level A operation state, level A reconstruction state, level B operation state, level B reconstruction state, level C operation state, and level C reconstruction state. After the system is powered on, it defaults to level A operation state.

[0045] Step 4: Each core program runs to the output comparison node. The AMPcrr software enables the checkpointsave() function to save the running status of each core program as a checkpoint file. The FPGA collects the four outputs and synchronously compares the output results of the three outputs marked as on-duty outputs. Step 5: Take two out of three votes on the current output operation results, output a unique operation result, and compare the output unique result with the backup output result: If the output of the backup core is consistent with the voting result, its status is considered normal and it continues to operate as a backup core in Class A status. The FPGA's working status indicator remains unchanged. If the backup core disagrees with the voting result, and the degree of inconsistency reaches a system-defined threshold, the FPGA sends a rebuild command to the backup core. The backup core executes the restore() function, parsing the checkpoint stored in secure memory and restoring the backup core's operating state to its last normal state. The FPGA then marks the processor's operating state as a Class A rebuild state. If the backup core successfully recovers, the processor's operating state returns to Class A operation. Otherwise, the FPGA disables the backup core's outputs and sends a shutdown command to shut down the system running the backup core. The FPGA selects two of the remaining three outputs as active outputs and one as a backup output, marking the processor's operating state as Class B.

[0046] Step 6: Determine the on-duty core exception handling strategy. In the A-level operating state, inconsistent calculation results may occur between on-duty outputs. In this case, the system performs the following processing based on the specific exception situation: If any calculation result in the current output is inconsistent with the other two, and the degree reaches the set threshold, the FPGA sends a reconstruction command to the core, and the core executes the restore() function to restore the working state to the previous normal state. The FPGA marks the processor working state as the A-level reconstruction state. If the result of the backup core is inconsistent with the voting result at this time, the backup core also performs a reconstruction operation, and the FPGA marks the system as the B-level reconstruction state.

[0047] If the calculation results of the three on-duty outputs are inconsistent, if the calculation result of the backup output is consistent with the result of any on-duty output, the other two abnormal judgment cores are reconstructed and the system enters the C-level operation state; If the calculation results of the backup core are inconsistent with the results of all the on-duty cores, only one core will be retained to continue working, and the other three cores will participate in reconstruction, and the system will enter the C-level operation state.

[0048] Step 7: In the B-level operating state, if the core undergoing the reconfiguration operation can successfully recover to its previous operating state, the FPGA restores the processor state to the A-level operating state. If the core undergoing the reconfiguration operation cannot successfully recover to its previous operating state, the FPGA blocks the core's output and sends a shutdown command, switching the system to a two-out-of-three operating mode.

[0049] Step 8: After the system is downgraded to 2-out-of-3 mode, the FPGA performs a simultaneous 2-out-of-3 vote on the outputs of the three cores and outputs the correct result. If any core output is inconsistent with the others, a reconstruction recovery operation is performed on it, marking the system as a Class B reconstruction state. If the outputs of all three cores are inconsistent, two cores are selected for reconstruction, marking the system as a Class C reconstruction state.

[0050] Step 9: In the B-level reconfiguration state, if the core undergoing the reconfiguration operation can successfully recover to its previous operating state, the FPGA restores the processor state to the B-level operating state. If the core undergoing the reconfiguration operation cannot successfully recover to its previous operating state, the FPGA blocks the core's output and sends a shutdown command, switching the system to dual-machine operation mode and transitioning to the C-level operating state.

[0051] Step 10: In the C-level working state, FPGA starts the single-machine fault detection module to monitor the on-duty core in real time. If the cleaning and reconstruction are not completed, and the results of any CPU in the judgment are inconsistent with the other two CPUs, the cleaning and reconstruction of the abnormal CPU will be triggered, and the system will enter the B-level operation state; If the results of the three judgment CPUs are inconsistent, one judgment CPU is retained, and the other two judgment CPUs perform cleaning and reconstruction, and the system enters the C-level operation state.

[0052] In Class B operating state: If the cleaning CPU successfully completes the reconstruction, it is added to the judgment group and the system returns to the Class A operating state; If the cleaning CPU fails to reconstruct multiple times, the backup CPU joins the judgment group, and the cleaning CPU continues to try to reconstruct, and the system enters the A-level reconstruction state; If the backup CPU and the voting result are inconsistent, the backup CPU performs cleaning and reconstruction; If the results of the two judgment CPUs are inconsistent, and the backup CPU agrees with one of them, the abnormal judgment CPU is cleaned and reconstructed; If the backup CPU is inconsistent with both decision CPUs, only one decision CPU is kept working, and the remaining CPUs perform cleaning and reconstruction, and the system enters the C-level operation state.

[0053] Step 4: Processing strategy in C-level status In C-level operating state: If the CPU is cleaned and the reconstruction is completed, the system status is adjusted according to the number of successful reconstructions: Reconstruct a CPU and the system enters the B-level operation state; After reconstructing the two CPUs, the system enters the A-level operation state.

[0054] If the clean CPU fails to reconstruct multiple times, the backup CPU is added to the judgment group and the system status is adjusted.

[0055] In the C-level refactoring state: The system checks the status of all current CPUs. If the status of three or four CPUs is consistent, the system is adjusted to Class B or Class A operating status based on the consistency. If there are less than two consistent CPUs, the system remains in the C-level reconstruction state and pauses outputting calculation results.

[0056] The above description is merely a preferred embodiment of a heterogeneous redundancy control method for a multi-core embedded processor. The scope of protection of a heterogeneous redundancy control method for a multi-core embedded processor is not limited to the above embodiment. All technical solutions based on this concept fall within the scope of protection of the present invention. It should be noted that improvements and variations that do not depart from the principles of the present invention are within the scope of protection of the present invention.

Claims

1. A heterogeneous redundancy control method applicable to a multi-core embedded processor, characterized by: The method comprises the following steps: Step 1: The embedded processor and FPGA start up. The FPGA completes self-test and processor detection. If no faults are detected, it enters the working mode. The embedded processor core starts the Linux, RT-thread, and FreeRTOS operating systems respectively and executes the same task program and AMPcrr. The other core directly executes the bare metal program. Step 2: The four processor cores use task-level synchronization to synchronize tasks. When the program reaches a checkpoint, each core broadcasts a task synchronization frame. After receiving the synchronization frames of the remaining modules, the synchronization ends and the task continues. If no module synchronization frame is received within a timeout, the synchronization frame is recorded and sent to the voting module, which saves the result. Step 3: The FPGA selects three outputs and marks them as active outputs, and one output as backup output. Based on the comparison results of the four outputs, they are marked as Class A operation state, Class A reconstruction state, Class B operation state, Class B reconstruction state, Class C operation state, and Class C reconstruction state. After power-on, the system defaults to Class A operation state. Step 4: Each core program runs to the output comparison node. AMPcrr enables the checkpointsave() function to save the running status of each core program as a checkpoint file. The FPGA collects four outputs and synchronously compares the output results of the three outputs marked as current outputs.

2. The method according to claim 1, wherein: Perform a two-out-of-three vote on the current output operation results, output a unique operation result, and compare the output unique result with the backup output result: When the output of the backup core is consistent with the voting result, its status is considered normal and it continues to operate as a backup core in Class A status. The FPGA's working status indicator remains unchanged. When the backup core is inconsistent with the voting result, and the degree of inconsistency reaches the system-set threshold, the FPGA sends a reconstruction command to the backup core. The backup core executes the restore() function, parses the checkpoint saved in the secure memory, and restores the working state of the backup core to the last normal state. The FPGA identifies the processor working state as a Class A reconstruction state. When the backup core can be successfully restored, the processor working state returns to the A-level operation state. Otherwise, the FPGA blocks the output of the backup core and sends a shutdown command to shut down the system running the backup core; The FPGA selects two of the remaining three outputs as on-duty outputs and the other as a backup output, marking the processor operating state as level B.

3. The method according to claim 2, wherein: Determine the on-duty core exception handling strategy. In the A-level operating state, inconsistent calculation results may occur between on-duty outputs. At this time, the system performs the following processing based on the specific exception situation: When any of the calculation results in the value output is inconsistent with the other two, and the degree reaches the set threshold, the FPGA sends a reconstruction command to the core. The core executes the restore() function to restore the working state to the previous normal state. The FPGA marks the processor working state as the A-level reconstruction state. When the result of the backup core is inconsistent with the voting result at this time, the backup core also performs the reconstruction operation, and the FPGA marks the system as the B-level reconstruction state. If the calculation results of the three on-duty outputs are inconsistent, if the calculation result of the backup output is consistent with the result of any on-duty output, the other two abnormal judgment cores will be reconstructed and the system will enter the C-level operation state; When the calculation results of the backup core are inconsistent with the results of all the on-duty cores, only one core is retained to continue working, and the other three cores participate in reconstruction, and the system enters the C-level operation state.

4. The method according to claim 3, wherein: In the B-level operating state, if the core undergoing the reconstruction operation can successfully recover to the previous working state, the FPGA will restore the processor state to the A-level operating state. If the core undergoing the reconstruction operation cannot successfully recover to the previous working state, the FPGA will block the core output and send a shutdown command, and the system will switch to the three-out-of-two operating mode. After the system is downgraded to the 2-out-of-3 operating mode, the FPGA performs a synchronous 2-out-of-3 vote on the outputs of the three cores and outputs the correct voting result. If the output of any core is inconsistent with the other cores, it is reconstructed and recovered, and the system is marked as a Class B reconstruct state. If the outputs of all three cores are inconsistent, two cores are selected for reconstructing, and the system is marked as a Class C reconstruct state. In the B-level reconfiguration state, when the core undergoing the reconfiguration operation can be successfully restored to the previous working state, the FPGA restores the processor state to the B-level operating state; If the core undergoing the reconstruction operation cannot be successfully restored to the previous working state, the FPGA blocks the core output and sends a shutdown command, and the system switches to dual-machine working mode and transforms to Class C working state; In the C-level working state, the FPGA starts the single-machine fault detection module to monitor the health status of the on-duty core in real time.

5. The method according to claim 4, wherein: When cleaning and reconstruction are not completed, and the results of any CPU in the judgment are inconsistent with the other two CPUs, the cleaning and reconstruction of the abnormal CPU is triggered, and the system enters the B-level operation state; When the results of the three judgment CPUs are inconsistent, one judgment CPU is retained, and the other two judgment CPUs perform cleaning and reconstruction, and the system enters the C-level operation state.

6. The method according to claim 5, wherein: In Class B operating state: If the cleaning CPU successfully completes the reconstruction, it is added to the judgment group and the system returns to the Class A operating state; If the cleaning CPU fails to reconstruct multiple times, the backup CPU joins the judgment group, and the cleaning CPU continues to try to reconstruct, and the system enters the A-level reconstruction state; If the backup CPU and the voting result are inconsistent, the backup CPU performs cleaning and reconstruction; If the results of the two judgment CPUs are inconsistent, and the backup CPU agrees with one of them, the abnormal judgment CPU is cleaned and reconstructed; If the backup CPU is inconsistent with both decision CPUs, only one decision CPU is kept working, and the remaining CPUs perform cleaning and reconstruction, and the system enters the C-level operation state.

7. The method according to claim 6, wherein: In C-level operating state: If the CPU is cleaned and the reconstruction is completed, the system status is adjusted according to the number of successful reconstructions: Reconstruct a CPU and the system enters the B-level operation state; Reconstruct the two CPUs and the system enters the A-level operation state; If the CPU cleaning fails to be reconstructed multiple times, the backup CPU is added to the judgment group and the system status is adjusted; In the C-level refactoring state: The system checks the status of all current CPUs. If the status of three or four CPUs is consistent, the system is adjusted to Class B or Class A operating status based on the consistency. If there are less than two consistent CPUs, the system remains in the C-level reconstruction state and pauses outputting calculation results.

8. A heterogeneous redundant control system suitable for multi-core embedded processors, characterized by: The system comprises: A self-test module is started by the embedded processor and FPGA. The FPGA completes self-test and processor detection and enters the working mode after no fault is found. The embedded processor core starts the Linux, RT-thread, and FreeRTOS operating systems respectively and executes the same task program and AMPcrr. Another core directly executes the bare metal program; A synchronization processing module, which synchronizes the four processor cores using task-level synchronization. When the program reaches a checkpoint, each core broadcasts a task synchronization frame. After receiving the synchronization frames of the remaining modules, the synchronization ends and the task continues. If no module synchronization frame is received within a timeout, the synchronization frame is recorded and sent to the voting module, which saves the result. An output module, wherein the output module selects three outputs of the FPGA and marks them as on-duty outputs and another output as a backup output, and marks them as level A operation state, level A reconstruction state, level B operation state, level B reconstruction state, level C operation state, and level C reconstruction state based on the comparison results of the four outputs. The system is in level A operation state by default after power-on; A synchronous comparison module runs each core program to the output comparison node. AMPcrr enables the checkpointsave() function to save the running status of each core program as a checkpoint file. FPGA collects four outputs and synchronously compares the output results of three of them marked as on-duty outputs.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method according to claims 1 to 7.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method of claims 1-7 is implemented.

Citation Information

Cited By

  • Process memory allocation method and electronic equipment

    CN121210335A

  • A satellite-borne high-performance computing platform based on a multi-core of Hongmeng and a star flash wireless bus and a plug-and-play method thereof

    CN122411721A