A microprocessor fault recovery device and method, and a chip
By introducing a field backup module and a security management module into the dual-core lock step structure of the CPU, and using software and hardware collaboration to back up the field environment at the preset recovery point, the problem of achieving high-function safety failure recovery at low hardware costs in the existing technology is solved, and fast and secure CPU failure recovery is achieved.
Patent Information
- Application Number
- CN202411783595.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-06
AI Technical Summary
While the prior art meets high safety level ASIL-D, it is difficult to achieve high functional and safe fault recovery at lower hardware costs, especially in the event of CPU failure. Traditional reset methods cannot meet the requirements of fault-tolerant time interval FTTI.
The microprocessor CPU adopts a dual-core lock-step structure, combined with the field backup module and the security management module, backs up the field environment at a preset recovery point, and restores to a safe operating state when the CPU fails, and uses the collaboration of software and hardware to achieve rapid failure recovery.
It realizes the fast recovery of the CPU's safe operation state at a lower hardware cost, shortens the function recovery time, enhances functional safety, and improves the robustness of automotive-grade functional safety.
Smart Images

Figure CN119248578B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fault processing, and in particular to a microprocessor fault recovery device and method, and a chip. Background Art
[0002] Chips with embedded high-speed central processing units (CPUs) as the core are widely used in automotive electronic and electrical systems. As the complexity of the system increases, the risk of system failure and random hardware failure increases. In order to enhance automotive safety, automotive electronic and electrical systems need to meet the ISO 26262 "Road Vehicle Functional Safety" international standard. Among them, ASIL-D is the highest safety level in the ISO26262 standard.
[0003] In order to meet the safety level of ASIL-D, one solution is that the CPU adopts a dual-core lock-step structure. The CPU contains two identical processor cores, one is the main core and the other is the secondary core. They execute the same code and are strictly synchronized. The outputs are sent to the comparison logic module to check the consistency of the data, address and control lines between them. When any inconsistency is detected, it means that the CPU has an error. When a CPU error occurs, the conventional practice is to directly reset the two cores, but this is difficult to meet the requirements of the fault tolerance time interval FTTI (Fault Torelant time interval), and it is likely that serious consequences will occur before the system is restored.
[0004] Another solution is to use a three-core CPU structure and a three-choice-two voting circuit to determine which core has an error and only correct the error. This solution does not require resetting the system, can well meet the time requirements of FTTI, and has very high reliability, but the disadvantage is that the hardware overhead is large, and compared with the previous solution, one more CPU is added, which reduces the market competitiveness of the product. Summary of the invention
[0005] One of the purposes of the present invention is to overcome the deficiencies in the prior art and to provide a microprocessor fault recovery device, method and chip that can take into account the advantages of the two existing solutions and achieve a high functional safety level at a lower hardware cost.
[0006] The technical solution provided by the present invention is as follows:
[0007] A microprocessor fault recovery device comprises a microprocessor CPU with a dual-core lock-step structure, a field backup module electrically connected to the microprocessor, and a safety management module electrically connected to the microprocessor and the field backup module respectively;
[0008] When the program runs to the preset recovery point in the program code, the microprocessor notifies the on-site backup module to back up the current on-site environment;
[0009] The on-site backup module records the location of the recovery point and the values of important parameters in the current on-site environment;
[0010] When receiving a fault indication from the CPU, the safety management module notifies the CPU to suspend operation, and after receiving a successful response from the CPU to suspend operation, notifies the on-site backup module to restore the on-site operation;
[0011] After receiving the on-site recovery instruction, the on-site backup module overwrites the corresponding current data with the backup data, so that the current on-site environment rolls back to the on-site environment corresponding to the recovery point;
[0012] After receiving the on-site recovery success response from the on-site backup module, the security management module adjusts the program counter of the CPU to the position of the recovery point and then starts the operation of the CPU.
[0013] In some embodiments, after receiving a successful response from the CPU to suspend operation, the security management module resets the CPU and sets the reset reason to a second type of reset;
[0014] After receiving the reset instruction, the CPU executes the initialization program to restore the settings of the first category of important parameters, and checks the reset cause after the initialization program is completed. If the reset cause is a second category reset, the security management module is notified that the restoration of the first category of important parameters has been completed;
[0015] The security management module notifies the on-site backup module to restore the on-site to restore the second category of important parameters.
[0016] In some embodiments, after receiving the suspension instruction, the CPU stops executing new instructions and waits for the completion of the instructions being executed. After the instructions being executed are completed, the CPU sends a suspension success response to the security management module.
[0017] The present invention also provides a microprocessor fault recovery method, based on the above-mentioned microprocessor fault recovery device, comprising:
[0018] When the program runs to the preset recovery point in the program code, the microprocessor CPU notifies the on-site backup module to back up the current on-site environment;
[0019] The on-site backup module records the location of the recovery point and the values of important parameters in the current on-site environment;
[0020] When receiving a fault indication from the CPU, the safety management module notifies the CPU to suspend operation, and after receiving a successful response from the CPU to suspend operation, notifies the on-site backup module to restore the on-site operation;
[0021] After receiving the on-site recovery instruction, the on-site backup module overwrites the corresponding current data with the backup data, so that the current on-site environment rolls back to the on-site environment corresponding to the recovery point;
[0022] After receiving the on-site recovery success response from the on-site backup module, the security management module adjusts the program counter of the CPU to the position of the recovery point and then starts the operation of the CPU.
[0023] In some embodiments, after receiving a successful response from the CPU to suspend operation, the security management module notifies the on-site backup module to restore the on-site, including:
[0024] After receiving the CPU's successful response to suspending operation, the security management module resets the CPU and sets the reset reason to the second type of reset;
[0025] After receiving the reset instruction, the CPU executes the initialization program to restore the settings of the first category of important parameters, and checks the reset cause after the initialization program is completed. If the reset cause is a second category reset, the security management module is notified that the restoration of the first category of important parameters has been completed;
[0026] The security management module notifies the on-site backup module to restore the on-site to restore the second category of important parameters.
[0027] In some embodiments, it includes: after receiving the suspension instruction, the CPU stops executing new instructions and waits for the completion of the instructions being executed, and sends a suspension success response to the security management module after the instructions being executed are completed.
[0028] In some embodiments, it includes: if the security management module does not receive a successful response from the CPU to suspend operation within a preset time after notifying the CPU to suspend operation, the security management module resets the CPU and sets the reset reason to a first type of reset.
[0029] In some embodiments, including:
[0030] If the security management module does not receive a successful on-site recovery response from the on-site backup module within a preset time after notifying the on-site backup module to restore the on-site, the security management module resets the CPU and sets the reset reason to a first-class reset.
[0031] In some embodiments, including:
[0032] After receiving the reset instruction, the CPU starts to execute from the initialization program. After executing the initialization program, it checks the reset cause. If the reset cause is a first-class reset, the subsequent program code is executed in sequence.
[0033] The present invention also provides a chip, comprising the microprocessor fault recovery device of any of the aforementioned embodiments.
[0034] The microprocessor fault recovery device, method and chip provided by the present invention can at least bring the following beneficial effects:
[0035] 1. The present invention backs up the corresponding on-site environment when the CPU runs safely to the recovery point. When the CPU fails, the backup on-site environment covers the current on-site environment, so that the CPU can quickly return to the safe operation state corresponding to the recovery point, shortening the recovery time of the CPU function, enhancing the safety of the CPU function, and improving the robustness of the automotive-grade functional safety.
[0036] 2. The present invention subdivides the parameters of the on-site environment and only backs up some parameters at the recovery point, thereby reducing the cost of hardware backup. The recovery of some parameters is then achieved through software (system initialization program), and the recovery of other parameters is achieved through hardware. In this way, through the collaboration of software and hardware, in the event of a CPU error, the software can be assisted to recover the site as quickly as possible with less hardware overhead.
[0037] 3. The present invention provides a complete fault recovery solution for CPU fault recovery, including the first type of reset recovery and the second type of reset recovery. The first type of reset adopts software, allowing the software to have the opportunity to execute any recovery process to recover a small number of scenes that the second type of reset cannot recover; the second type of reset adopts the collaborative working mode of software and hardware, allowing the software to realize the recovery of a part of the scene, and then quickly recover all the remaining scenes through the hardware according to the preset recovery process, thereby shortening the recovery time.
[0038] 4. The present invention can set recovery points in the program according to actual needs, and quickly save the on-site environment of the recovery point only when needed. There is no need to back up after each instruction is executed. Compared with the solution of frequently generating CPU operation snapshots or frequently saving recovery points, it significantly saves power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The preferred implementation scheme will be described below in a clear and understandable manner in conjunction with the accompanying drawings to further illustrate the above-mentioned characteristics, technical features, advantages and implementation methods of a microprocessor fault recovery device and method, and chip.
[0040] Figure 1 It is a structural schematic diagram of an embodiment of a microprocessor fault recovery device of the present invention;
[0041] Figure 2 is a flow chart of an embodiment of a microprocessor fault recovery method of the present invention;
[0042] Figure 3 It is a schematic structural diagram of an embodiment of a chip of the present invention.
[0043] Description of Figure Numbers:
[0044] 100. Microprocessor, 200. On-site backup module, 300. Security management module, 10. Microprocessor fault recovery device. DETAILED DESCRIPTION
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the specific implementation methods of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings and other implementation methods can be obtained based on these drawings without creative work.
[0046] In order to simplify the drawings, only the parts related to the present invention are schematically shown in each figure, and they do not represent the actual structure of the product. In addition, in order to simplify the drawings and facilitate understanding, in some figures, only one of the parts with the same structure or function is schematically drawn or marked. In this article, "one" not only means "only one", but also means "more than one".
[0047] One embodiment of the present invention, as Figure 1 As shown, a microprocessor fault recovery device 10 includes a microprocessor 100 with a dual-core lockstep structure, a field backup module 200 electrically connected to the microprocessor, and a safety management module 300 electrically connected to the microprocessor and the field backup module respectively.
[0048] The microprocessor CPU contains two identical processor cores, one is the main core and the other is the secondary core. They execute the same program code and are strictly synchronized, sending the outputs to the comparison logic module respectively. The comparison logic module checks the consistency of the data, address and control lines between them. When any inconsistency is detected, a warning is triggered, such as initiating an interrupt or setting a signal to inform the CPU that an error has occurred.
[0049] A number of recovery points are pre-set in the program code. When the program runs to the recovery point preset in the program code, the microprocessor notifies the on-site backup module to back up the current on-site environment. For example, a special instruction is set at the recovery point, and the instruction is used by the CPU to notify the on-site backup module to back up the current on-site environment. When the program code runs to the recovery point, the CPU executes this special instruction. The recovery point can be set at a place where the functional correlation between the previous and next two programs is weak, which can reduce the backup amount of parameters. For example, setting a recovery point between the program that implements function A and the program that implements function B is less than setting a recovery point in the program of function A. Usually, fewer parameters need to be backed up. The number and location of the recovery points can be determined according to actual needs.
[0050] The purpose of backing up the current on-site environment is to enable the system to fall back to a certain safe operating state when a CPU failure occurs. The on-site backup module records the location of the recovery point and the values of various parameters in the current on-site environment. In order to reduce the backup volume and consumption of storage resources, only the values of important parameters can be backed up. Important parameters refer to parameters whose current values affect the subsequent operation of the CPU, such as global variables in the program. Unimportant parameters refer to parameters whose current values do not affect the subsequent operation of the CPU, such as some temporary variables.
[0051] The on-site backup module can save data from only one recovery point. In this case, the data from the new recovery point overwrites the data from the old recovery point. It can also save data from more than one recovery point. In this case, a circular overwriting method can be used. If the storage space is insufficient to store the data from the new recovery point, the data from the earliest saved recovery point will be overwritten.
[0052] The on-site backup module can be implemented with hardware circuits, which can efficiently perform data backup. In some scenarios where the state changes slowly, the on-site backup module can be implemented with a software program, such as executing an interrupt service program to complete the data backup. The CPU continues the subsequent process after the interrupt service program is completed. In vehicle driving scenarios, it is recommended that the on-site backup module use a hardware circuit.
[0053] When the CPU detects that the dual-core operation is inconsistent, it sends a fault indication to the safety management module. The safety management module notifies the CPU to suspend operation upon receiving the fault indication from the CPU.
[0054] CPU suspension means suspending the operation of two cores at the same time. By suspending the CPU operation, the CPU is put into a safe state to prevent the fault from expanding.
[0055] After receiving the pause instruction, the CPU stops executing new instructions and waits for the completion of the instructions being executed. When the instructions being executed are completed, the CPU sends a successful pause response to the security management module. After receiving the successful pause response from the CPU, the security management module notifies the on-site backup module to restore the site.
[0056] After receiving the on-site recovery instruction, if the on-site backup module only saves data from one recovery point, it will use the backup data to update the current on-site environment, so that the current on-site environment rolls back to the on-site environment corresponding to the most recently recorded recovery point, and then feedback a successful recovery response to the security management module. If the on-site backup module saves data from multiple recovery points, the security management module can instruct the recovery point to roll back, or it can roll back to the most recent recovery point first, and if unsuccessful, roll back to a more distant recovery point, that is, roll back step by step.
[0057] The on-site backup module is implemented using hardware, which can complete the backup and recovery of the on-site environment in a shorter time than software implementation.
[0058] The security management module can obtain the location of the recovery point from the on-site backup module, and after receiving the on-site recovery success response from the on-site backup module, adjust the program counter of the CPU to the location of the recovery point, and then start the operation of the CPU. For example, the location of the recovery point is carried in the on-site recovery success response. The CPU jumps to the recovery point and starts execution.
[0059] It should be noted that the CPU fault indication is not limited to the fault indication when the dual-core operation is inconsistent, but may also include fault indications in other situations. As long as the conventional solution to the fault is a conventional reset, and the CPU can still normally receive and execute instructions from the security management module under the fault, the CPU fault recovery measures provided in this embodiment can be used to shorten the fault recovery time.
[0060] In this embodiment, the field backup module backs up the corresponding field environment when the CPU runs safely to the recovery point. When the CPU fails, the backup data is used to overwrite the current field environment, so that the CPU can quickly roll back to the safe operation state corresponding to the recovery point, avoiding the traditional CPU reset that must be executed from the beginning, shortening the recovery time of the CPU function, and enhancing the safety of the CPU function, thereby achieving the ASIL-D safety level of the vehicle.
[0061] In one embodiment, it further includes:
[0062] If the security management module does not receive a successful response from the CPU to suspend operation within a preset time after notifying the CPU to suspend operation, the security management module resets the CPU.
[0063] If the security management module does not receive a successful on-site recovery response from the on-site backup module within a preset time after notifying the on-site backup module to recover the on-site, the security management module resets the CPU.
[0064] The CPU starts execution from the beginning after receiving a reset instruction.
[0065] The above two situations indicate that under the current fault, the CPU / field backup module cannot normally execute the instructions of the security management module, and the aforementioned solution of rolling back to the recovery point cannot be accurately and reliably implemented. Therefore, the traditional method of resetting the CPU is adopted, and the CPU starts to execute from the beginning.
[0066] In one embodiment, after receiving the CPU's successful response to suspending operation, the security management module notifies the on-site backup module to restore the on-site, including:
[0067] After receiving the CPU's successful response to suspending operation, the security management module resets the CPU and sets the reset reason to the second type of reset;
[0068] After receiving the reset instruction, the CPU executes the initialization program to restore the settings of the first category of important parameters. After the initialization program is completed, the reset reason is checked. If the reset reason is a second category reset, the security management module is notified that the restoration of the first category of important parameters has been completed; the security management module then notifies the on-site backup module to restore the site to restore the second category of important parameters.
[0069] Specifically, in some cases, there are many parameters that need to be backed up in the field environment. If the field backup module adopts a software method, the backup time is long; if a hardware circuit is used, the hardware overhead is large; plus the characteristics of the parameters themselves, some parameters need to restore fixed values, such as the interrupt enable switch parameter, which must always be restored to "allow interrupt enable", which is not affected by the restore point. Therefore, the parameters can be divided into two parts: the first type of important parameters (the values that need to be restored are not affected by the restore point) and the second type of important parameters (the values that need to be restored are affected by the restore point, and the values of different restore points may be different). The first type of important parameters does not need to be backed up at the restore point, and the required settings can be restored by executing the system initialization program; the second type of important parameters are placed in the field backup module for backup and recovery, which reduces the amount of data backed up by the field backup module each time, reduces hardware costs, and is conducive to realizing the function of the field backup module in hardware. For example, the field backup module can be implemented with a storage circuit, and data backup is realized by writing data into the storage circuit, and data recovery is realized by reading data from the storage circuit.
[0070] Based on the above reasons, after receiving the CPU's successful response to suspending operation, the security management module first resets the CPU and lets the CPU execute the initialization program to restore the settings of the first category of important parameters, and then notifies the on-site backup module to restore the second category of important parameters after the restoration of the first category of important parameters is completed. The on-site backup module rolls back the second category of important parameters to the values at the restoration point.
[0071] This embodiment also includes:
[0072] If the security management module does not receive a successful response from the CPU to suspend operation within a preset time after notifying the CPU to suspend operation, the security management module resets the CPU and sets the reset reason to the first type of reset.
[0073] If the security management module does not receive a successful on-site recovery response from the on-site backup module within a preset time after notifying the on-site backup module to restore the on-site, the security management module resets the CPU and sets the reset reason to a first-class reset.
[0074] After receiving the reset instruction, the CPU starts to execute from the initialization program. After executing the initialization program, it checks the reset cause. If the reset cause is a first-class reset, the subsequent program code is executed in sequence.
[0075] In order to enable the CPU to distinguish between a regular reset and a reset that rolls back to a recovery point, the reset cause is divided into a first-class reset and a second-class reset. The first-class reset corresponds to a regular reset, and the second-class reset corresponds to a reset that rolls back to a recovery point. After executing the initialization program, the CPU checks the reset cause. If it is a first-class reset, the subsequent program code is executed in sequence. If it is a second-class reset, the security management module is notified that the recovery of the first-class important parameters has been completed, so that the security management module can recover the second-class important parameters.
[0076] This embodiment provides a complete fault recovery method for CPU fault recovery, including a first type of reset and a second type of reset. The first type of reset adopts software, allowing the software to have the opportunity to execute any recovery process to recover a small number of scenes that cannot be recovered by the second type of reset; the second type of reset can adopt a software and hardware collaborative working method to shorten the recovery time, allowing the software to realize the recovery of a part of the scene, and then quickly recover all the remaining scenes through the hardware according to the preset recovery process.
[0077] One embodiment of the present invention, as Figure 2 As shown, based on the microprocessor fault recovery device of the above embodiment, a microprocessor fault recovery method includes:
[0078] Step S100: When the program code reaches a preset recovery point, the CPU notifies the on-site backup module to back up the current on-site environment.
[0079] Step S200: The on-site backup module backs up the current on-site environment.
[0080] Step S300: When it is detected that the dual-cores are running inconsistently, the CPU sends a fault indication to the security management module.
[0081] In step S400, the security management module notifies the CPU to suspend operation, and after receiving a successful response from the CPU to suspend operation, notifies the on-site backup module to restore the on-site.
[0082] Step S500: After receiving the on-site recovery instruction, the on-site backup module overwrites the corresponding current data with the backup data, so that the current on-site environment rolls back to the on-site environment corresponding to the recovery point.
[0083] Step S600: After receiving the on-site recovery success response from the on-site backup module, the security management module adjusts the program counter of the CPU to the position of the recovery point and then starts the operation of the CPU.
[0084] Specifically, several recovery points are preset in the program code. The microprocessor CPU runs the program code. When the program runs to the preset recovery point, it can notify the on-site backup module to back up the current on-site environment by executing a special instruction, or writing a certain value to a specific hardware register, or using a similar method. The purpose of backing up the current on-site environment is to enable the system to fall back to a certain safe operating state when a CPU failure occurs.
[0085] The location and number of recovery points can be determined according to actual needs, so that the on-site environment of the recovery points can be quickly saved only when needed, which significantly saves power consumption compared to frequently generating CPU operation snapshots or frequently saving recovery points.
[0086] The field backup module records the current value of the program counter (i.e., the location of the recovery point) and also records the values of various parameters in the current field environment, including at least the values of important parameters. Important parameters refer to parameters whose current values affect the subsequent CPU operation, which can be the values of certain registers or certain memories. Unimportant parameters refer to parameters whose current values do not affect the CPU operation.
[0087] The on-site backup module can save data from only one recovery point, in which case the data from the new recovery point overwrites the data from the old recovery point; it can also save data from more than one recovery point, in a cyclic overwriting manner.
[0088] The on-site backup module can be a software program or a hardware circuit. Compared with the former, the latter is more efficient and is preferred.
[0089] When the safety management module receives the CPU fault indication, it notifies the CPU to suspend operation and puts the CPU in a safe state. After receiving the suspension indication, the CPU stops executing new instructions and waits for the completion of the instructions being executed. When the instructions being executed are completed, it sends a successful suspension response to the safety management module. After receiving the successful suspension response from the CPU, the safety management module notifies the on-site backup module to restore the site.
[0090] If the on-site backup module only saves data from one recovery point, the backup data is used to update the current on-site environment, so that the current on-site environment is rolled back to the on-site environment corresponding to the most recently recorded recovery point, and then a successful recovery response is fed back to the security management module. If the on-site backup module saves data from multiple recovery points, the security management module can indicate the recovery point to roll back to, or it can try to roll back step by step in a preset order.
[0091] After receiving the recovery success response, the security management module adjusts the PC pointer to point to the recovery point, and then starts the CPU operation; the CPU jumps to the recovery point and starts execution.
[0092] In this embodiment, by backing up the on-site environment of the recovery point under safe system operation, when a CPU failure occurs, the backup data is used to overwrite the current on-site environment, so that the CPU can quickly roll back to the safe operation state corresponding to the recovery point, avoiding the CPU from resetting from the beginning, shortening the recovery time of system functions, and thus achieving a high functional safety level.
[0093] In one embodiment, step S400 includes: if the security management module does not receive a successful response from the CPU to suspend operation within a preset time after notifying the CPU to suspend operation, the security management module resets the CPU.
[0094] The process after step S500 also includes: if the security management module does not receive a successful on-site recovery response from the on-site backup module within a preset time after notifying the on-site backup module to recover the on-site, the security management module resets the CPU.
[0095] The CPU starts execution from the beginning after receiving a reset instruction.
[0096] The above two situations indicate that under the current fault, the CPU / field backup module cannot normally execute the instructions of the security management module, and the aforementioned solution of rolling back to the recovery point cannot be accurately and reliably implemented. Therefore, the traditional method of resetting the CPU is adopted, and the CPU starts to execute from the beginning.
[0097] In one embodiment, the step S400 in which the security management module notifies the CPU to suspend operation and the step S400 in which the CPU receives a successful response to suspend operation includes:
[0098] In step S410, if the security management module does not receive a successful response from the CPU to suspend operation within a preset time after notifying the CPU to suspend operation, the security management module resets the CPU and sets the reset reason to the first type of reset.
[0099] After receiving the reset instruction, the CPU starts to execute from the initialization program. After executing the initialization program, it checks the reset cause. If the reset cause is a first-class reset, the subsequent program code is executed in sequence.
[0100] In one embodiment, in step S400, after receiving the CPU's successful response to suspending operation, the security management module notifies the on-site backup module to restore the on-site, including:
[0101] Step S420: After receiving the CPU's successful response to suspending operation, the security management module resets the CPU and sets the reset reason to the second type of reset;
[0102] Step S430: After receiving the reset instruction, the CPU executes the initialization program to restore the settings of the first category of important parameters, and checks the reset reason after the initialization program is completed. If the reset reason is a second category reset, the security management module is notified that the restoration of the first category of important parameters has been completed;
[0103] In step S440, the security management module notifies the on-site backup module to restore the on-site to restore the second category of important parameters.
[0104] In one embodiment, it also includes: if the security management module does not receive a successful on-site recovery response from the on-site backup module within a preset time after notifying the on-site backup module to restore the site, the security management module resets the CPU and sets the reset reason to the first type of reset.
[0105] After receiving the reset instruction, the CPU starts to execute from the initialization program. After executing the initialization program, it checks the reset cause. If the reset cause is a first-class reset, the subsequent program code is executed in sequence.
[0106] It should be noted that the embodiment of the microprocessor fault recovery method provided by the present invention and the embodiment of the microprocessor fault recovery device provided above are based on the same inventive concept and can achieve the same technical effect. Therefore, other specific contents of the embodiment of the microprocessor fault recovery method can refer to the contents of the embodiment of the microprocessor fault recovery device provided above.
[0107] One embodiment of the present invention, as Figure 3 As shown, a chip includes the microprocessor fault recovery device 10 described in any of the above embodiments.
[0108] The microprocessor fault recovery device 10 can be integrated into a chip, and the chip can be used in various devices, such as automotive electronic and electrical systems.
[0109] It should be noted that the above embodiments can be freely combined as needed. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered as the protection scope of the present invention.
Claims
1. A microprocessor fault recovery device, characterized in that: It includes a microprocessor CPU with a dual-core lock-step structure, a field backup module electrically connected to the microprocessor, and a safety management module electrically connected to the microprocessor and the field backup module respectively; When the program code reaches a preset recovery point, the microprocessor notifies the on-site backup module to back up the current on-site environment; the parameters of the current on-site environment include the first type of important parameters and the second type of important parameters; The on-site backup module records the location of the recovery point and the value of the second type of important parameters in the current on-site environment; When receiving a fault indication from the CPU, the security management module notifies the CPU to suspend operation, and after receiving a successful response from the CPU to suspend operation, resets the CPU and sets the reset reason to a second type of reset; After receiving the reset instruction, the CPU executes an initialization program to restore the settings of the first category of important parameters, and checks the reset cause after the initialization program is completed. If the reset cause is a second category reset, the CPU notifies the security management module that the restoration of the first category of important parameters has been completed; The security management module notifies the on-site backup module to restore the second type of important parameters; After receiving the on-site restoration instruction, the on-site backup module overwrites the corresponding current data with the backup data, so that the second type of important parameters of the current on-site environment are rolled back to the on-site environment corresponding to the restoration point; After receiving the on-site recovery success response from the on-site backup module, the security management module adjusts the program counter of the CPU to the position of the recovery point and then starts the operation of the CPU.
2. The microprocessor fault recovery device according to claim 1, characterized in that: A special instruction is preset at the recovery point for the microprocessor CPU to notify the on-site backup module to back up the current on-site environment. When the program code runs to the recovery point, the microprocessor CPU executes the special instruction.
3. The microprocessor fault recovery device according to claim 1, characterized in that: After receiving the suspension instruction, the CPU stops executing new instructions and waits for the completion of the instructions being executed. After the instructions being executed are completed, the CPU sends a suspension success response to the security management module.
4. A microprocessor fault recovery method, characterized in that: Based on the microprocessor fault recovery device according to claim 1, the method comprises: When the program code reaches a preset recovery point, the microprocessor CPU notifies the on-site backup module to back up the current on-site environment; the parameters of the current on-site environment include the first type of important parameters and the second type of important parameters; The on-site backup module records the location of the recovery point and the value of the second type of important parameters in the current on-site environment; The security management module notifies the CPU to suspend operation when receiving a fault indication from the CPU, and resets the CPU after receiving a successful response to suspend operation from the CPU, and sets the reset reason to a second type of reset; After receiving the reset instruction, the CPU executes an initialization program to restore the settings of the first category of important parameters, and checks the reset cause after the initialization program is completed. If the reset cause is a second category reset, the CPU notifies the security management module that the restoration of the first category of important parameters has been completed; The security management module notifies the on-site backup module to restore the second type of important parameters; After receiving the on-site restoration instruction, the on-site backup module overwrites the corresponding current data with the backup data, so that the second type of important parameters of the current on-site environment are rolled back to the on-site environment corresponding to the restoration point; After receiving the on-site recovery success response from the on-site backup module, the security management module adjusts the program counter of the CPU to the position of the recovery point and then starts the operation of the CPU.
5. The microprocessor fault recovery method according to claim 4, characterized in that: When the program code reaches a preset recovery point, the microprocessor notifies the on-site backup module to back up the current on-site environment, including: A special instruction is preset at the recovery point for the microprocessor CPU to notify the on-site backup module to back up the current on-site environment. When the program code runs to the recovery point, the microprocessor CPU executes the special instruction.
6. The microprocessor fault recovery method according to claim 4, characterized in that: include: After receiving the suspension instruction, the CPU stops executing new instructions and waits for the completion of the instructions being executed. After the instructions being executed are completed, the CPU sends a suspension success response to the security management module.
7. The microprocessor fault recovery method according to claim 4, characterized in that: include: If the security management module does not receive a successful response from the CPU to suspend operation within a preset time after notifying the CPU to suspend operation, the security management module resets the CPU and sets the reset reason to the first type of reset.
8. The microprocessor fault recovery method according to claim 4, characterized in that: include: If the security management module does not receive a successful on-site recovery response from the on-site backup module within a preset time after notifying the on-site backup module to restore the on-site, the security management module resets the CPU and sets the reset reason to the first type of reset.
9. The microprocessor fault recovery method according to claim 7 or 8, characterized in that: include: After receiving the reset instruction, the CPU starts to execute from the initialization program, and checks the reset cause after executing the initialization program. If the reset cause is a first type of reset, the subsequent program codes are executed in sequence.
10. A chip, characterized in that: The microprocessor fault recovery device comprises the microprocessor fault recovery device according to any one of claims 1 to 3.
Citation Information
Patent Citations
Hardware rapid recovery architecture of dual-core lockstep processor, electronic equipment and method
CN118733352A
Cited By
Fault recovery device and method for dual-core lockstep processor
CN120448191A
A dual-core lockstep processor fault recovery device and method
CN120448191B