Method and device for collecting hardware errors in starting stage

A technology for hardware errors in the startup phase, applied in fault hardware testing methods, generation of response errors, and error detection/correction.

Active Publication Date: 2021-01-26
SUZHOU LANGCHAO INTELLIGENT TECH CO LTD
View PDF3 Cites 0 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Problems solved by technology

When the machine crashes and automatically restarts, the operating system cannot process the fault data in time, but after the machine restarts, some hardware error information in Intel's CPU can still be retained. At this time, there is still a chance to collect error information and record it in the fault log. After a hardware error, the operating system crashes and restarts automatically, and the operating system cannot process the fault data in time. After the machine restarts, the fault data is collected and stored in the BERT (Boot Error Record Table, Boot Error Record Table) of ACPI (Advanced Configuration and Power Management Interface) after the machine restarts. table) form, which provides error information to the operating system
[0004] However, when a fatal error occurs in the system, the problem of restart failure often occurs. At this time, the code for collecting fault data has not yet been run, so the problem that the fault data cannot be collected occurs, and the cause of the fault cannot be located.
Moreover, the fault data collected in the later stage of startup is directly filled into the BERT table of ACPI and reported to the operating system. After receiving the fault data, the operating system isolates the address. In many cases, it cannot do further in-depth Hierarchical fault location

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • Method and device for collecting hardware errors in starting stage
  • Method and device for collecting hardware errors in starting stage
  • Method and device for collecting hardware errors in starting stage

Examples

Experimental program
Comparison scheme
Effect test

Embodiment Construction

[0030] The preferred embodiments of the present invention will be described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described here are only used to illustrate and explain the present invention, and are not intended to limit the present invention.

[0031] Embodiments of the present invention provide a method for collecting hardware errors during startup, such as Figure 1~2 shown, including:

[0032] S1. The first stage: Before the memory is not initialized, collect the first MCA error information of a single core and non-core in each CPU and send it to the BMC to detect whether there is an IERR error in the last startup, and if so, perform a cold restart , if not, enter the second stage;

[0033] S2, the second stage: After all the cores in the memory and the CPU are initialized, before the MCA is initialized, collect the second MCA error information of all cores and non-cores in each CPU and send it to the ...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

PUM

No PUM Login to View More

Abstract

The invention provides a method and device for collecting hardware errors in a starting stage, and the method comprises the steps: in the first stage, collecting first MCA error information of a single core and a non-core in each CPU before a memory is not initialized, transmitting the first MCA error information to a BMC, detecting whether there is an IERR error in the last starting, carrying outthe cold restarting if there is an IERR error, and entering the second stage if not; in the second stage, after initialization of the memory and all the cores in the CPUs and before MCA initialization, collecting and sending second MCA error information of all the cores and non-cores in all the CPUs to the BMC, detecting whether IERR errors exist in last starting or not, if so, conducting cold restarting once, and if not, after MCA initialization and before IO error processing initialization, collecting IO fault data of the IO and sending the IO fault data to the BMC, so that the BMC processes the first MCA error information, the second MCA error information and the IO fault data, diagnoses, locates and outputs a fault log. According to the method and device for collecting the hardware errors in the starting stage, faults can be positioned more accurately.

Description

technical field [0001] The invention relates to the technical field of server monitoring, in particular to a method and device for collecting hardware errors in the startup phase. Background technique [0002] At present, with the development of the Internet era in recent years, the demand for massive data processing capabilities is growing rapidly, thus putting forward higher requirements for servers. To a decisive role, in today's rapid development of network technology, virtualization technology, and distributed applications, the availability, reliability, and serviceability indicators required by servers are getting higher and higher. Financial services and telecommunication services have become indispensable elements of economic and social life anytime and anywhere. The normal operation of financial and telecommunication services is highly dependent on the continuous and stable operation of the information system, which also puts forward high requirements on the availab...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

Application Information

Patent Timeline
no application Login to View More
Patent Type & AuthorityApplications(China)
IPC IPC(8): G06F11/14G06F11/22
CPCG06F11/1417G06F11/2236G06F11/2273
Inventor罗鹏芳王兵陈思彤
OwnerSUZHOU LANGCHAO INTELLIGENT TECH CO LTD