BMC Register Capture Before CPU Reset
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In computers with processing modules, fatal errors result in the loss of information stored in registers due to automatic restarts, preventing the operating system from determining the error source, as existing fault management mechanisms are ineffective in corrupted systems.
Innovation Solution
A method where the programmable logic circuit suspends the reset request and alerts the management controller to read and store information from chosen registers before allowing the reset, enabling real-time data recovery post-error.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the CPU automatically restarts following a fatal error, then the system recovers quickly, but the information in volatile registers is erased and cannot be analyzed
Solution Approach 1:
The management controller performs preliminary action by reading and storing register information before the CPU restart is executed. The BMC intercepts the restart command, captures the state of volatile registers while they still contain valid data, and preserves this information in non-volatile storage, thus preventing information loss before recovery occurs
Solution Approach 2:
The baseboard management controller (BMC) acts as an intermediary between the CPU and the restart process. It receives the restart command from the CPU, mediates by reading and storing the register information, then allows the restart to proceed. This intermediary role enables both information preservation and system recovery
2Loss of information
If the operating system attempts to read register information after a fatal error, then information can be obtained, but the CPU is already corrupted and must be restarted, erasing the data
Solution Approach 1:
The system performs preliminary action by having the BMC read and store register information before the CPU restart is executed. This preliminary capture occurs while the CPU is still functional and before volatile memory is cleared, ensuring information is preserved despite the subsequent restart
Solution Approach 2:
The management controller autonomously detects when a fatal error has occurred and automatically reads and stores the register information without requiring OS intervention. This self-service capability ensures information is captured even when the OS cannot intervene due to CPU corruption
3Loss of information
If the management controller reads register information before reset, then data is preserved for analysis, but the reset process is delayed
Solution Approach 1:
The BMC performs preliminary action by reading and storing register information in advance of the reset operation. This preliminary capture is executed quickly during the window between error detection and reset, minimizing the delay while ensuring information is preserved before volatile memory is cleared
Solution Approach 2:
The management controller rushes through the information capture process by reading only the essential register data quickly and efficiently. This rapid capture minimizes the time delay before reset while still obtaining the necessary diagnostic information
Data Source
Figure 1
Figure 2
AI summary
A method makes it possible to obtain information stored in registers (R11 -R4J) of at least one processing module (MT1 -MTJ) of a computer (CA), each processing module (MT1 -MTJ) furthermore comprising a management controller (CG1 -CGJ) able to read the information stored in the associated registers (R11 -R4J) and a programmable logic circuit (CL1 -CLJ) to trigger a reset requested following a fatal error. Accordingly, in the event of reception of a reset request by a programmable logic circuit (CL2) of a processing module (MT2), this programmable logic circuit (CL2) suspends the triggering of this reset and alerts the occurrence of a fatal error to the associated management controller (CG2) which, if it is capable thereof, reads the information stored in associated and chosen registers (R12-R42), then stores this read information in a file, then the associated programmable logic circuit (CL2) is authorized to trigger the requested reset.