Fault processing and recovery method based on lock-step architecture

By analyzing the processor's lock loss status and formulating targeted fault handling measures, the problem of reduced reliability in traditional lock step fault handling methods is solved, and the reliability of the flight control system is improved.

CN119938347APending Publication Date: 2025-05-06XIAN FLIGHT SELF CONTROL INST OF AVIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411983958.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The traditional lockstep fault handling method is too silent when detecting faults, resulting in a reduction in the overall reliability of the flight control computer.

Method used

By analyzing the causes of lock loss in different lock loss states of the processor, specific fault handling and recovery measures are formulated for various lock loss causes, including reloading the target code, forcibly reset the processor, designing the fault monitor, etc.

Benefits of technology

It improves the reliability of the flight control system in the loading and operation states, and enhances the system's ability to detect and recover faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938347A_ABST
    Figure CN119938347A_ABST
Patent Text Reader

Abstract

The invention provides a fault processing and recovering method based on a lockstep architecture, which comprises the following steps that: 1) software inquires comparison results of address buses, control buses or data buses of two same processors by reading a lockstep state register; 2) when the flight control computer is in a loading state, judging a lock step state, and formulating a lock step fault processing and recovery process of the flight control computer in the loading state; (3) when the flight control computer is in the running state, lock step state judgment is carried out, and classification is carried out according to different lock step states; 4) formulating different fault processing and recovery methods according to different lock losing states of the RAM module, the FLASH module, the CPBUS module, the network module and the NVM module; according to the fault analysis and processing method of the lockstep architecture computer, provided by the invention, the fault mode of lockstep lock losing is analyzed in a detailed manner, different fault processing measures are formulated, and the reliability of the flight control computer is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fault processing and recovery, and in particular to a fault processing and recovery method based on a lockstep architecture. Background Art

[0002] In order to improve the availability and safety of flight control computers, hardware redundancy and software redundancy technologies are often introduced when designing flight control computers. When a redundant component in the flight control computer fails, the system will reorganize the remaining valid components to complete redundant reorganization and monitoring voting to achieve system functions. Lockstep technology is a method of hardware redundancy and an effective method of organizing redundant processor components to achieve high-integrity computing.

[0003] In a lockstep processor system, the hardware logic cycle compares the instructions of two identical processors. When the address bus, control bus or data bus of the two processors are different, the lock loss result and reason are fed back to the software processor through the register value.

[0004] The traditional lock-step fault handling method adopts a silent handling method. When a lock-step fault occurs, the flight control computer cuts off the external output to ensure that no wrong flight control instructions are sent. This method has a high fault detection rate, but reduces the overall reliability of the flight control computer. Summary of the invention

[0005] Purpose of the invention: The present invention proposes a lock-step fault handling method to analyze the lock-step fault causes of different processor lock states. Different fault handling measures and recovery measures are formulated for various lock-step fault causes, thereby enhancing the reliability of the flight control system.

[0006] In a first aspect, the present application provides a fault handling and recovery method based on a lockstep architecture, wherein the flight control computer includes a downloader software, a FLASH module, a RAM module, a CPBUS module, an NVM module, and a network module; the method includes:

[0007] When the flight control computer is in the loading state, the host computer sends the target code to the downloader software through the network;

[0008] The downloader software fixes the target code into the FLASH module, periodically reads the lock-step flag of the flight control computer, and identifies the lock-step state of the current system based on the lock-step flag;

[0009] If the FLASH module or the network module is unlocked during the loading process, the downloader software reloads the target code. When the number of reloads exceeds the preset number, the software reports the failure of loading to the upper computer.

[0010] If the RAM module, CPBUS module or NVM module loses lock during the loading process, continue loading.

[0011] In a second aspect, the present application further provides a fault handling and recovery method based on a lockstep architecture, wherein the flight control computer includes an application, a FLASH module, a RAM module, a CPBUS module, an NVM module, a network module, and a processor; the method includes:

[0012] When the flight control computer is in operation, the application program periodically reads the lock-step flag to identify the lock-step state of the current system.

[0013] Preferably, when the flight control computer is in a running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including:

[0014] When the RAM module loses lock, it is considered that the processor has a fault when running the target code. At this time, the output instructions of the flight control computer are in an uncontrollable state and the lock failure of the flight control computer cannot be recovered;

[0015] The flight control computer forces a processor reset, and the RAM module re-enters the running state as the computer is reset. If the RAM module loses lock again within a period of time, it is considered a RAM hardware failure, and the flight control computer enters a fault silent state.

[0016] Preferably, when the flight control computer is in a running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including:

[0017] When the FLASH module is unlocked, it is considered that the two processors of the flight control computer have inconsistent comparisons of the statically stored target code or parameter data items. This fault is an irrecoverable fault, and the target code and parameter data items will directly affect the running branches and output instructions of the flight control computer software. At this time, the flight control computer directly enters a fault silent state.

[0018] Preferably, when the flight control computer is in a running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including:

[0019] When the CPBUS module is unlocked, it means that the data interaction between the flight control computer and the IO interface board is unlocked, including the external 429 bus data, AFDX data, IMB data and CLDL data. This fault is a recoverable fault and is allowed to occur occasionally during system operation.

[0020] Design the CPBUS module fault monitor module, set the monitor's threshold value, count increase ratio value, and count decrease ratio value; when a fault occurs once, the count value increases once, and when the fault does not occur, the count value decreases once; when the count value reaches the threshold value, the flight control computer directly enters the fault silent state.

[0021] Preferably, when the flight control computer is in a running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including:

[0022] When the network module loses lock, the fault is handled according to the ground status of the flight control computer.

[0023] Preferably, when the network module is unlocked, the fault processing is performed according to the ground state of the flight control computer, including:

[0024] When the flight control computer is in the ground state, it is mostly in the maintenance state and needs to be connected to the ground maintenance equipment through the network. A reliable network state is required. At this time, the flight control computer is restarted;

[0025] When the flight control computer is in the air, there is no need to connect the network and onboard devices. The loss of network module will not affect the flight status. At this time, the loss of network module will be ignored and the flight control computer will continue to perform the mission.

[0026] Preferably, when the flight control computer is in a running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including:

[0027] When the NVM module loses lock, it means that the fault value or zeroing data of the monitor stored in the flight control computer is inconsistent when the two processors compare. The fault value of the monitor is set by the application software during the cycle, which is a recoverable fault.

[0028] The zeroing data is stored in NVM before the cycle runs, which is an unrecoverable fault. The NVM module fault monitor module is designed to set the monitor threshold value, count increase ratio value, and count decrease ratio value;

[0029] When a fault occurs once, the count value increases once, and when a fault does not occur, the count value decreases once;

[0030] When the count value reaches the threshold, the flight control computer directly enters the fault silent state.

[0031] Beneficial technical effects of this application:

[0032] The present application provides a technical means for handling and recovering the lock-step faults of the flight control computer in a loaded state. According to the impact analysis of the lock-step failure of the computer module in the loaded state, a corresponding fault handling and recovery method is proposed, thereby improving the reliability of the flight control system in the loaded state.

[0033] The flight control computer's lock-step fault handling and recovery technical means in the present application are analyzed according to the different conditions of each computer module in the air and on the ground, and the impact of the computer lock-step failure in the running state is analyzed, and corresponding fault handling and recovery methods are proposed, thereby improving the reliability of the flight control system when it is in the running state.

[0034] In summary, the present invention achieves improved reliability of the flight control system in multiple scenarios, and has significant technical progress and practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a flowchart of lock-step fault processing and recovery of a flight control computer in a loaded state provided by an embodiment of the present application;

[0036] Figure 2 The present invention provides a flowchart of a lock-step fault handling and recovery process of a flight control computer in operation. DETAILED DESCRIPTION

[0037] The present application provides a method for fault synthesis and processing of a lock-step architecture computer, which realizes fault processing and recovery of a highly reliable locksetp architecture computer through software logic. The detailed implementation method includes the following steps:

[0038] When the flight control computer is in the loading state, the host computer can send the target code to the flight control computer in lock step through the network, and the downloader software of the flight control computer will solidify the target code into the FLASH module. The downloader software identifies the lock step state of the current system by periodically reading the lock step flag. If the FLASH module or network module loses lock during the loading process, the downloader reloads the target code. When the number of reloads exceeds 3 times, the host computer is fed back that the loading failed. When the RAM module, CPBUS module or NVM module loses lock, the loading continues. The lock step fault handling and recovery process of the flight control computer in the loading state is as shown in the attached figure. Figure 1 shown.

[0039] When the flight control computer is in operation, the application program periodically reads the lock-step flag to identify the current system lock-step state, and divides the lock-loss state into the following states. The lock-step fault handling and recovery process when the flight control computer is in operation is as follows: Figure 2 shown.

[0040] The different lockstep fault handling and recovery strategies are as follows:

[0041] When the RAM module loses lock, it is considered that the processor has a fault when running the target code. At this time, the output instructions of the flight control computer are in an uncontrollable state and the flight control computer's lock failure cannot be recovered. The flight control computer will force the processor to reset, and the RAM module will re-enter the running state as the computer resets. If the RAM module loses lock again within a period of time, it is considered that the RAM hardware has failed, and the flight control computer enters a fault silent state (shutdown mode).

[0042] When the FLASH module is unlocked, it is considered that the two processors of the flight control computer have an inconsistency in the comparison of the statically stored target code or parameter data items. This fault is an irrecoverable fault, and the target code and parameter data items will directly affect the running branches and output instructions of the flight control computer software. At this time, the flight control computer directly enters the fault silent state (shutdown mode).

[0043] When the CPBUS module loses lock, it means that the data interaction between the flight control computer and the IO interface board has lost lock, including external 429 bus data, AFDX data, IMB data and CLDL data. This fault is a recoverable fault and can be allowed to occur occasionally when the system is running. Design the CPBUS module fault monitor module, set the threshold value of the monitor, the count increase ratio value, and the count decrease ratio value. When the fault occurs once, the count value increases once, and when the fault does not occur, the count value decreases once. When the count value reaches the threshold value, the flight control computer directly enters the fault silent state (shutdown mode).

[0044] When the network module is unlocked, the fault is handled according to the ground status of the flight control computer. When the flight control computer is on the ground, it is mostly in maintenance status and needs to be connected to the ground maintenance equipment through the network. A reliable network status is required, and the flight control computer is restarted at this time. When the flight control computer is in the air, it does not need to be connected to the network and onboard equipment. The network module unlocking will not affect the flight status. At this time, the network module unlocking is ignored and the flight control computer continues to perform the mission.

[0045] When the NVM module loses lock, it means that the fault value or zeroing data of the monitor stored in the flight control computer is inconsistent when compared by two processors. The fault value of the monitor is set by the application software during the cycle and is a recoverable fault. The zeroing data is stored in the NVM before the cycle runs and is an unrecoverable fault. Design the NVM module fault monitor module, set the threshold value of the monitor, the count increase ratio value, and the count decrease ratio value. When the fault occurs once, the count value increases once, and when the fault does not occur, the count value decreases once. When the count value reaches the threshold value, the flight control computer directly enters the fault silent state (shutdown mode).

[0046] In other embodiments of the present application, the specific method provided by the present invention is as follows:

[0047] Step 1: The processor uses two POWEPC755 processors to implement lock-step logic. The address, data, and control signals accessed on the local bus between the processors are compared every 20us to see if they are consistent. If the comparison result is inconsistent, a 16-bit register is set, where each bit represents a loss of lock state. When it is 1, it means loss of lock, and when it is 0, it means normal. The register will be reset each time it is read.

[0048] Step 2: When the flight control computer is in the loading state, the downloader software reads the lock-step flag in a 50ms cycle to identify the lock-step state of the current system. If the FLASH module or network module unlock bit in the lock-step register is 1 during the loading process, the downloader reloads the target code. When the number of reloads exceeds 3 times, the loading failure is fed back to the host computer.

[0049] Step 3: When the flight control computer is in operation, the application reads the lockstep register in a 12.5ms cycle to determine whether each valid bit is 1 and identify the lockstep status of the current system.

[0050] Step 4: When the RAM module unlock bit is 1, read the last reset time of the computer stored in NVM and calculate the time difference with the current time. If the time difference is greater than 30 minutes, reset the computer. If the time difference is less than 30 minutes, the flight control computer enters the fault silent state (shutdown mode).

[0051] Step 5: When the RAM module unlock bit is 1, the flight control computer enters a fault silent state (shutdown mode).

[0052] Step 6: When the CPBUS module unlock bit is 1, design the CPBUS module fault monitor module, set the monitor threshold value to 60, the count increase ratio value to 15, and the count decrease ratio value to 5. When a fault occurs once, the count value increases once, and when a fault does not occur, the count value decreases once. When the count value reaches the threshold value, the flight control computer directly enters the fault silent state (shutdown mode).

[0053] Step 7: When the network module unlock bit is 1, determine whether the flight control computer is on the ground. When the flight control computer is on the ground, reset the flight control computer. When the flight control computer is in the air, the flight control computer continues to execute the mission.

[0054] Step 8: When the NVM module unlock bit is 1, design the NVM module fault monitor module, set the monitor threshold value to 45, the count increase ratio value to 15, and the count decrease ratio value to 5. When a fault occurs once, the count value increases once, and when a fault does not occur, the count value decreases once. When the count value reaches the threshold value, the flight control computer directly enters the fault silent state (shutdown mode).

[0055] In other embodiments of the present application, the present application provides a fault handling and recovery method based on a lockstep architecture, including:

[0056] 1) The software queries the comparison result of the address bus, control bus or data bus of two identical processors by reading the lockstep status register.

[0057] 2) When the flight control computer is in the loading state, the lock-step state is judged, and the lock-step fault handling and recovery process of the flight control computer in the loading state is formulated.

[0058] 3) When the flight control computer is in operation, the lock-step state is determined and classified according to different lock-step states.

[0059] 4) Different fault handling and recovery methods are formulated for different lock-out states of RAM module, FLASH module, CPBUS module, network module and NVM module.

[0060] The fault analysis and processing method of the lockstep architecture computer provided by the present invention analyzes the fault mode of lockstep loss of lock in detail, formulates different fault processing measures, and increases the reliability of the flight control computer.

Claims

1. A fault handling and recovery method based on a lockstep architecture, characterized in that: The flight control computer includes downloader software, a FLASH module, a RAM module, a CPBUS module, an NVM module and a network module; the method includes: When the flight control computer is in the loading state, the host computer sends the target code to the downloader software through the network; The downloader software fixes the target code into the FLASH module, periodically reads the lock-step flag of the flight control computer, and identifies the lock-step state of the current system based on the lock-step flag; If the FLASH module or the network module is unlocked during the loading process, the downloader software reloads the target code. When the number of reloads exceeds the preset number, the software reports the failure of loading to the upper computer. If the RAM module, CPBUS module or NVM module loses lock during the loading process, continue loading.

2. A fault handling and recovery method based on a lockstep architecture, characterized in that: The flight control computer includes an application program, a FLASH module, a RAM module, a CPBUS module, an NVM module, a network module and a processor; the method includes: When the flight control computer is in operation, the application program periodically reads the lock-step flag to identify the lock-step state of the current system.

3. The method according to claim 2, characterized in that When the flight control computer is in the running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including: When the RAM module loses lock, it is considered that the processor has a fault when running the target code. At this time, the output instructions of the flight control computer are in an uncontrollable state and the lock failure of the flight control computer cannot be recovered; The flight control computer forces a processor reset, and the RAM module re-enters the running state as the computer is reset. If the RAM module loses lock again within a period of time, it is considered a RAM hardware failure, and the flight control computer enters a fault silent state.

4. The method according to claim 2, characterized in that: When the flight control computer is in the running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including: When the FLASH module is unlocked, it is considered that the two processors of the flight control computer have inconsistent comparisons of the statically stored target code or parameter data items. This fault is an irrecoverable fault, and the target code and parameter data items will directly affect the running branches and output instructions of the flight control computer software. At this time, the flight control computer directly enters a fault silent state.

5. The method according to claim 2, characterized in that: When the flight control computer is in the running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including: When the CPBUS module is unlocked, it means that the data interaction between the flight control computer and the IO interface board is unlocked, including the external 429 bus data, AFDX data, IMB data and CLDL data. This fault is a recoverable fault and is allowed to occur occasionally during system operation. Design the CPBUS module fault monitor module, set the monitor's threshold value, count increase ratio value, and count decrease ratio value; when a fault occurs once, the count value increases once, and when the fault does not occur, the count value decreases once; when the count value reaches the threshold value, the flight control computer directly enters the fault silent state.

6. The method according to claim 2, characterized in that When the flight control computer is in the running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including: When the network module loses lock, the fault is handled according to the ground status of the flight control computer.

7. The method according to claim 6, characterized in that When the network module is unlocked, the fault is processed according to the ground state of the flight control computer, including: When the flight control computer is in the ground state, it is mostly in the maintenance state and needs to be connected to the ground maintenance equipment through the network. A reliable network state is required. At this time, the flight control computer is restarted; When the flight control computer is in the air, there is no need to connect the network and onboard devices. The loss of network module will not affect the flight status. At this time, the loss of network module will be ignored and the flight control computer will continue to perform the mission.

8. The method according to claim 2, characterized in that: When the flight control computer is in the running state, the application program periodically reads the lock-step flag to identify the lock-step state of the current system, including: When the NVM module loses lock, it means that the fault value or zeroing data of the monitor stored in the flight control computer is inconsistent when the two processors compare. The fault value of the monitor is set by the application software during the cycle, which is a recoverable fault. The zeroing data is stored in NVM before the cycle runs, which is an unrecoverable fault. The NVM module fault monitor module is designed to set the monitor threshold value, count increase ratio value, and count decrease ratio value; When a fault occurs once, the count value increases once, and when a fault does not occur, the count value decreases once; When the count value reaches the threshold, the flight control computer directly enters the fault silent state.