Error log collection method and data processing equipment
By implementing the error log collection method in the data processing device, the operating system obtains and persists the error logs when hardware exception notifications are notified, solving the problem of high server downtime in cloud computing data centers, realizing the persistence and reporting of multi-platform error logs, helping to locate the root cause and avoid batch downtime.
Patent Information
- Application Number
- CN202311812929.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
In hyperscale cloud computing data centers, the server downtime is high, resulting in system reliability, availability and serviceability. It is difficult for the existing technology to effectively collect hardware error logs to locate the root cause.
By implementing an error log collection method in a data processing device, the operating system responds to the hardware exception notification sent by the firmware, obtains the hardware error log from the specified memory buffer, and writes the error log to the flash memory when the persistent storage conditions are met, restarts the operating system, and obtains and persists the error log from the flash memory after restart through the persistent storage file system.
It realizes the error log of the hardware platform error interface specification through the persistent storage error interface before the operating system goes down. It is universal and is suitable for the persistence of error logs on multiple platforms. It also obtains hardware error logs from the flash memory after the operating system is restarted. It triggers the user-state process to collect and persist storage through the preset notification mechanism, solving the problem of large-scale impact of batch downtime.
Smart Images

Figure CN120216228A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for collecting error logs and a data processing device. Background Art
[0002] In a large-scale cloud computing data center, the server downtime rate has always been a key indicator for measuring the reliability, availability, and serviceability (RAS, Reliability, Availability, and Serviceability) of the system, and it is also the primary issue for meeting the SLA (Service Level Agreement) of cloud computing end users.
[0003] For unexpected hardware downtime in the production environment, producers usually need to analyze the root cause of the downtime based on the error logs that trigger the hardware errors in order to find solutions to avoid large-scale batch downtime affecting projects and end users. Therefore, there is an urgent need for a general method that can collect hardware error logs when hardware downtime occurs. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a method for collecting error logs. One or more embodiments of this specification also relate to a data processing device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a method for collecting error logs is provided, which is applied to a data processing device. The data processing device includes an operating system, hardware, and firmware installed on the hardware. Among them,
[0006] The operating system, in response to a hardware exception notification sent by the firmware, obtains the error log of the hardware from a specified memory buffer.
[0007] Among them, the hardware exception notification is sent by the firmware in response to an error interrupt of the hardware after writing the error log of the hardware into the specified memory buffer according to the hardware platform error interface specification.
[0008] When the error log meets the persistent storage condition, write the error log into the flash memory and restart the operating system.
[0009] Load the persistent storage file system. When the system driver of the persistent storage file system determines that there is an error log in the flash memory, trigger a user-mode process to collect and persistently store the error log according to a preset notification mechanism.
[0010] According to a second aspect of the embodiments of the present specification, a method for collecting error logs is provided, which is applied to the operating system of a data processing device, and includes:
[0011] In response to a hardware exception notification sent by the firmware, obtain the error log of the hardware from a specified memory buffer.
[0012] Wherein, the firmware is installed on the hardware, and the hardware exception notification is sent after the firmware writes the error log of the hardware into the specified memory buffer in accordance with the hardware platform error interface specification in response to an error interrupt of the hardware.
[0013] When the error log meets the persistent storage condition, write the error log into the flash memory and restart the operating system.
[0014] Load the persistent storage file system. When the system driver of the persistent storage file system determines that there is an error log in the flash memory, trigger a user-mode process to collect and persistently store the error log according to a preset notification mechanism.
[0015] According to a third aspect of the embodiments of the present specification, a data processing device is provided. The data processing device includes an operating system, hardware, and firmware installed on the hardware. Among them,
[0016] The hardware is used to record error information corresponding to the hardware error in an error record register when a hardware error is detected, form the error log of the hardware, and trigger an error interrupt of the firmware.
[0017] The firmware is used to write the error log of the hardware into a specified memory buffer in accordance with the hardware platform error interface specification in response to the error interrupt of the hardware, and send a hardware exception notification to the operating system.
[0018] The operating system is used to obtain the error log of the hardware from the specified memory buffer in response to the hardware exception notification sent by the firmware. When the error log meets the persistent storage condition, write the error log into the flash memory and restart the operating system, load the persistent storage file system. When the system driver of the persistent storage file system determines that there is an error log in the flash memory, trigger a user-mode process to collect and persistently store the error log according to a preset notification mechanism.
[0019] According to a fourth aspect of the embodiments of the present specification, a computing device is provided, including:
[0020] A memory and a processor;
[0021] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned error log collection method are implemented.
[0022] According to the fifth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above-mentioned error log collection method are implemented.
[0023] According to the sixth aspect of the embodiments of the present specification, a computer program is provided. When the computer program is executed on a computer, the computer is made to execute the steps of the above-mentioned error log collection method.
[0024] The error log collection method provided by the embodiments of the present specification is applied to a data processing device, which includes an operating system, hardware, and firmware installed on the hardware. Among them, the operating system, in response to a hardware exception notification sent by the firmware, obtains the error log of the hardware from a specified memory buffer. The hardware exception notification is sent after the firmware writes the error log of the hardware into the specified memory buffer in accordance with the hardware platform error interface specification in response to a hardware error interrupt of the hardware. When the error log meets the persistent storage condition, the error log is written into the flash memory, and the operating system is restarted; a persistent storage file system is loaded. When the system driver of the persistent storage file system determines that there is an error log in the flash memory, according to a preset notification mechanism, a user-mode process is triggered to collect and persistently store the error log.
[0025] Based on this, the error log collection method writes the error log of the hardware platform error interface specification into the flash memory through a persistent storage error interface before the operating system crashes, which is universal and applicable to the persistence of multiple platform error logs. After the operating system restarts, the hardware error log is obtained from the flash memory, and a user-mode process is triggered through a preset notification mechanism to collect and persistently store the error log, so as to realize the collection of error logs when the hardware errors and crashes, which is applicable to the reporting of multiple platform error logs. Subsequently, the root cause can be located by analyzing the collected error logs at the time of crashing, avoiding batch crashes, and breaking through the limitation of the prior art that the error log collection is only applicable to the architecture of one hardware error detection and reporting mechanism and cannot be universal. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is an architecture diagram of a data processing device provided by an embodiment of the present specification;
[0027] Figure 2It is a flowchart of a method for collecting error logs applied to a data processing device provided by an embodiment of this specification;
[0028] Figure 3 It is a flowchart of the processing procedure of a method for collecting error logs provided by an embodiment of this specification;
[0029] Figure 4 It is a flowchart of a method for collecting error logs applied to an operating system of a data processing device provided by an embodiment of this specification;
[0030] Figure 5 It is a schematic structural diagram of a data processing device provided by an embodiment of this specification;
[0031] Figure 6 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0032] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0033] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0034] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".
[0035] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0036] First, the noun terms involved in one or more embodiments of this specification are explained.
[0037] APEI: Advanced Platform Error Interfaces, an advanced platform error interface specification that unifies the interfaces between software and hardware.
[0038] HEST: Hardware Error Source Table, which provides a way for the platform firmware to describe the system hardware error sources to the operating system, including notification methods, error data storage locations, response methods, etc.
[0039] GHES: Generic Hardware Error Source. GHES is the kernel driver of the above HEST, responsible for registering the specified notification handling functions in the HEST table, copying the error data filled by the firmware to the operating system, and handling errors.
[0040] ERST: Error Record Serialization Table. ERST is essentially an abstract interface for permanently storing errors. Software can write various error information into the physical medium available for permanent storage through the interfaces and formats specified by the ERST table.
[0041] PStore: A mechanism for persistent storage that provides a method for saving data in the file system.
[0042] Persistence: A mechanism for converting program data between persistent and transient states; generally speaking, it is to persist transient data (such as data in memory, which cannot be permanently saved) into persistent data (such as persisting to a database, which can be saved for a long time).
[0043] Unexpected system downtime not only affects the normal operation of projects but also damages the enterprise's reputation. In computing clusters and data centers, the hardware and software density deployed on a single physical machine is increasing. Hundreds of virtual machines and thousands of container instances can be deployed on each physical machine. Although hardware failures rarely occur, any server downtime can result in huge cost losses. According to a study of 63 data centers, each minute of downtime causes nearly $9,000 in losses.
[0044] Modern processors support rich RAS features, such as the hardware error detection and reporting mechanism of X86 and the RAS extensions of ARM64, which support the discovery, reporting, handling, and recovery of hardware errors. Thus, with the cooperation of the entire system stack of hardware, firmware, and operating system, hardware errors can be collected and analyzed, greatly reducing unexpected downtime in large-scale data centers. Among them, the hardware is responsible for discovering and recording errors, the firmware is responsible for error collection and reporting, and the OS is responsible for error handling and recovery.
[0045] Downtime error logs depend on two aspects: Persistence during downtime: When a fatal APEI error occurs, Linux will print the error log to the console or dump the persistent log through kdump (a tool and service used to dump memory running parameters when the system crashes, deadlocks, or freezes), and then restart. However, this highly depends on the normal operation of the operating system and the above services, and for some specific virtual machines and hosts, network communication also needs to be ensured to be normal. Persistence after hot restart: The registers that record APEI errors will retain their contents after a hot restart, so hardware errors can be recorded to disk or network after restart. However, the system may not be able to perform a hot restart, which may result in the loss of hardware error logs.
[0046] In this specification, an error log collection method is provided. This specification also relates to a data processing device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0047] See Figure 1 , Figure 1 shows an architecture diagram of a data processing device provided according to an embodiment of this specification. The data processing device includes an operating system, hardware, and firmware installed on the hardware.
[0048] Specifically, the error log collection method for the data processing device corresponding to this architecture diagram will be described in detail.
[0049] Among them, the data processing device can be understood as a computer or a server, which is not limited here.
[0050] The hardware in the data processing device, when detecting a hardware error, records the error information corresponding to the hardware error in an error record register and triggers an error interrupt of the firmware.
[0051] Among them, the hardware can be various hardware devices capable of generating interrupts, which can include a CPU (Central Processing Unit), a network card, a hard disk, etc. In the embodiments of this specification, the hardware can be understood as a processor; the error source can be understood as the hardware where the error interrupt occurs.
[0052] The error record register can be understood as a register for recording the error information corresponding to the hardware error, which can include a general-purpose register, a Processor State (PSTATE) register, a Program Counter (PC) register, an Exception Syndrome Register (ESR) at each exception level (ELn), that is, an ESR-ELn register, etc., but is not limited thereto.
[0053] Specifically, the hardware (such as a central processing unit) detects an error, such as an ECC error. ECC (Error Correcting Code) - error correction code, error correction code, ECC is an electronic method used to check the overall data stored in DRAM (Dynamic Random Access Memory); records the error information in the error record register to form an error log of the hardware error; at the same time, triggers relevant interrupts or exceptions of the hardware platform to notify the firmware that the hardware has detected an error.
[0054] The firmware in the data processing device, in response to the error interrupt of the hardware, writes the error log of the hardware recorded in the error record register into the memory buffer specified by the hardware error source table according to the hardware platform error interface specification and sends a hardware exception notification to the operating system.
[0055] Among them, the operating system can be understood as the Linux operating system. Of course, the error log collection method provided in the embodiments of this specification can also be applied to other operating systems. The hardware platform error interface specification is APEI, and the hardware error source table is HEST. The hardware error source table describes error information, which may include: notification type, reported error format, physical address of error data, etc. The notification type refers to the mechanism type by which the firmware notifies the kernel, and may include: System Error Interrupt (SEI), Synchronous External Abort (SEA), Software Delegated Exception (SDE), etc. Different notification types correspond to different system events, and different notification methods are registered with different event handlers.
[0056] Specifically, the firmware responds to the hardware interrupt, writes the error log in the error record register into the memory buffer specified by HEST in APEI format, and sends a hardware exception notification to the operating system.
[0057] The operating system in the data processing device responds to the hardware exception notification sent by the firmware, and obtains the error log of the hardware from the memory buffer specified by the hardware error source table. The hardware exception notification is sent after the firmware responds to the hardware error interrupt, writes the error log of the hardware recorded in the error record register into the memory buffer specified by the hardware error source table according to the hardware platform error interface specification.
[0058] That is, the operating system can respond to the hardware exception notification sent by the firmware, read the error data in the memory buffer specified by the hardware error source table according to the hardware platform error interface specification; when it is determined that the error log meets the persistent storage condition, write the error log into the flash through the persistent storage error interface and restart; after restarting, load the persistent storage file system, and when the system driver of the persistent storage file system determines that there is the error log in the flash through the persistent storage error interface, trigger the user-mode process to collect and persistently store the error log according to the preset notification mechanism.
[0059] Among them, the persistent storage error interface is the ERST interface; the persistent storage file system is the Pstore file system.
[0060] Specifically, in response to the hardware exception notification sent by the firmware, the operating system reads the error log in APEI format from the memory buffer specified by HEST through the GHES handler (event handling function) pre-registered in the kernel according to the notification type of the notification, and determines the error type of the error log, so as to implement the check of the severity of the hardware error; in the case that the hardware error is a fatal and uncorrectable error, the error log is written into the flash through the ERST interface, and the operating system restart is triggered.
[0061] In the case of an operating system restart, the Pstore file system is loaded by mounting the / sys / fs / postore file directory; the system driver of the Pstore file system checks whether there is an error log in the flash through the ERST interface; in the case that there is an error log in the flash, the system driver of the Pstore file system reads the error log from the flash through the ERST interface, deletes the error log in the flash, and triggers the RAS tracepoint to generate an error log trace event (RAS event), so that when the user-space process (Rasdaemon) monitors the error log trace event, it records and persists it to the database to collect and persistently store the error log.
[0062] The error log collection method provided by the embodiments of this specification is applied to a data processing device. By writing the error log that conforms to the hardware platform error interface specification into the flash through the persistent storage error interface before the operating system crashes, it has generality and is applicable to the persistence of error logs on multiple platforms. And after the operating system restarts, it obtains the hardware error log from the flash, and triggers the user-space process to collect and persistently store the error log through a preset notification mechanism, so as to realize the collection of error logs when the hardware error crashes, which is applicable to the reporting of error logs on multiple platforms. Subsequently, by analyzing the error logs collected during the crash, the root cause can be located to avoid batch crashes.
[0063] See Figure 2 , Figure 2 FIG. shows a flowchart of an error log collection method applied to a data processing device provided by an embodiment of this specification, which specifically includes the following steps.
[0064] This error log collection method is applied to a data processing device, and the data processing device includes an operating system, hardware, and firmware installed on the hardware.
[0065] Step 202: The operating system, in response to the hardware exception notification sent by the firmware, obtains the error log of the hardware from the specified memory buffer.
[0066] Among them, the hardware exception notification is sent after the firmware writes the error log of the hardware into the specified memory buffer in accordance with the hardware platform error interface specification in response to the error interrupt of the hardware.
[0067] Specifically, when the firmware responds to the error interrupt of the hardware, it writes the error log of the hardware recorded in the error record register into the memory buffer specified by the hardware error source table in accordance with the hardware platform error interface specification, and then sends a hardware exception notification to the operating system.
[0068] Among them, the hardware error source table is HEST. The hardware exception notification can be understood as a message notifying the operating system that the hardware has detected a hardware error.
[0069] Specifically, in response to the hardware exception notification sent by the firmware, the operating system obtains the error log of the faulty hardware from the memory buffer specified by HEST.
[0070] In one or more embodiments of this specification, before the operating system responds to the hardware exception notification sent by the firmware, when the hardware detects a hardware error, it first triggers an error interrupt of the firmware, so that the firmware responds to the error interrupt of the firmware and sends a hardware exception notification to the operating system. The specific implementation method is as follows:
[0071] Before responding to the hardware exception notification sent by the firmware, it further includes:
[0072] When the hardware detects a hardware error, it records the error information corresponding to the hardware error in the error record register to form the error log of the hardware, and triggers the error interrupt of the firmware.
[0073] Among them, the error information may include the notification type, the reported error format, and the physical address of the error data, etc.; the firmware is the underlying software running on the hardware.
[0074] Specifically, when the hardware detects a hardware error, it records the error information corresponding to the hardware error, such as the notification type of the reported error, the error format, and the physical address of the error data, in the error record register to form the error log corresponding to the hardware according to the above error information, and triggers the error interrupt of the firmware.
[0075] For the error log collection method provided by the embodiments of this specification, when the hardware detects a hardware error, it saves the error information in the error record register and triggers the error interrupt registered by the firmware, so that the firmware responds to the error interrupt of the hardware and sends a hardware exception notification to the operating system, which improves the efficiency of hardware error exception handling.
[0076] In one or more embodiments of this specification, after triggering an error interrupt of the firmware, the firmware can respond to the error interrupt and write the error log in the error record register to the memory buffer specified by the hardware error source table, so that the operating system can obtain the error log from the memory buffer specified by the hardware error source table. The specific implementation method is as follows:
[0077] After triggering the error interrupt of the firmware, it further includes:
[0078] The firmware, in response to the error interrupt of the hardware, writes the error log of the hardware recorded in the error record register to the memory buffer specified by the hardware error source table according to the hardware platform error interface specification, and sends a hardware exception notification to the operating system.
[0079] Among them, the hardware platform error interface specification is APEI; specifically, by writing to the memory buffer specified by the hardware error source table according to the APEI specification, the error log is written to the memory buffer specified by HEST in the APEI format.
[0080] In specific implementation, the firmware can respond to the error interrupt of the hardware, generate a CPER (Common Platform Error Record) entry, that is, a Common Error Data Entry, by calling the CPER (Common Platform Error Record) Generation Library according to the error information recorded in the hardware error record register, and write the CPER entry to HEST. CPER is the carrier for the firmware to transfer error information to the kernel, and CPER is carried by HEST; the HEST table supports multiple error source types. Among them, the error information recorded in the error record register may include: error type, address of error data, error number, etc.; the error type may include: CE (Correctable Error) and UE (Uncorrectable Error), etc.
[0081] When the firmware responds to the error interrupt of the hardware, it can send a hardware exception notification to the operating system, so that the kernel of the operating system can determine a suitable event processing function for exception handling according to the hardware exception notification.
[0082] For the error log collection method provided by the embodiments of this specification, the error handling of the hardware is in the firmware priority mode. The error interrupt of the hardware is first responded to by the firmware. The firmware collects the error log of the hardware, and then notifies the kernel and hands over the control right to the kernel for efficient processing.
[0083] In one or more embodiments of the present specification, after triggering an error interrupt of the firmware, the firmware can respond to the error interrupt and write the error log in the error record register to the memory buffer specified by the hardware error source table, so that the operating system can obtain the error log from the memory buffer specified by the hardware error source table. The specific implementation method is as follows:
[0084] The operating system includes a general hardware error source driver running in the kernel;
[0085] Responding to the hardware exception notification sent by the firmware and obtaining the error log of the hardware from the specified memory buffer includes:
[0086] The operating system, in response to the hardware exception notification sent by the firmware, triggers the general hardware error source driver and determines the corresponding target event handling function according to the notification type carried in the hardware exception notification;
[0087] Obtain the error log of the hardware from the memory buffer specified by the hardware error source table through the target event handling function.
[0088] Among them, the general hardware error source driver running in the kernel can be understood as the general hardware error source (Generic Hardware Error Source, GHES) driver of the kernel; this driver can register event handling functions (Handlers) according to the notification type carried in the hardware exception notification; the notification type refers to the mechanism type by which the firmware notifies the kernel, different notification types correspond to different system events, and different notification methods correspond to different registered event handling functions.
[0089] The event handling function can be understood as a function in the kernel that processes exceptions, including a synchronous external abort (SEA) handling function, a software delegated exception interface (SDEI) handling function, etc.
[0090] Specifically, the operating system can respond to the hardware exception notification sent by the firmware, trigger the GHES driver, determine the corresponding target event handling function from the pre-registered event handling functions according to the notification type by which the firmware notifies the kernel, and obtain the error log of the hardware from the memory buffer specified by the hardware error source table through the target event handling function.
[0091] For example, the operating system receives a hardware exception notification sent by the firmware, and this hardware exception notification indicates a memory ECC (Error Correcting Code) error; the operating system triggers the General Hardware Error Source driver GHES, which is responsible for handling various types of hardware errors; GHES determines the corresponding target event handling function, such as handle_ecc_error, according to the notification type carried in the hardware exception notification, that is, the memory ECC error.
[0092] The target event handling function handle_ecc_error obtains the error log related to the memory ECC error from the memory buffer specified in the hardware error source table; the error log can contain information such as the time when the error occurred, the affected memory address, the error type, etc., so that the operating system can perform further processing according to the obtained error log, such as recording the log, notifying the system administrator, attempting to repair the error or taking other appropriate actions.
[0093] In practical applications, the error log in the error recording register will be actively read by the firmware, and the operating system will obtain the hardware error log from the firmware in response to the hardware exception notification sent by the firmware.
[0094] The error log collection method provided by the embodiments of this specification determines the corresponding target event handling function according to the notification type of the hardware exception notification from the pre-registered event handling functions, so as to realize reasonably obtaining the hardware error log from the memory buffer specified in the hardware error source table.
[0095] Step 204: When the error log meets the persistent storage condition, write the error log into the flash memory and restart the operating system.
[0096] Among them, the persistent storage condition can be understood as the condition that can be used to persistently store the error log set in advance; for example, if the error log conforms to a certain error type, it can be determined that the error log meets the persistent storage condition.
[0097] The flash memory is a form of electronically erasable programmable read-only memory, which is a memory that allows to be erased or written multiple times during operation.
[0098] The persistent storage error interface is the ERST interface. The main function of ERST is to store various hardware or platform-related errors. Subsequently, the errors can be read out through appropriate methods for analysis, so as to quickly locate the cause of the error and solve it.
[0099] Specifically, the event handling function running in the kernel can write the error log in APEI format into the flash memory through the ERST interface and trigger the restart of the operating system when it determines that the error log meets the persistent storage conditions, so as to achieve the permanent storage of the error log. Moreover, writing the error log in APEI format into the flash memory through the ERST interface is universal and applicable to the persistence of error logs on multiple platforms.
[0100] In one or more embodiments of this specification, when it is determined that the error type of the error log is an uncorrectable error, the error log is written into the flash memory through the persistent storage error interface to persistently store the error log and ensure that the error log will not be lost. The specific implementation method is as follows:
[0101] When the error log meets the persistent storage conditions, writing the error log into the flash memory includes:
[0102] The operating system parses the error log through the target event handling function and determines the error type of the error log according to the parsing result;
[0103] When it is determined that the error type of the error log is the first error type, the error log is written into the flash memory through the persistent storage error interface.
[0104] Among them, the first error type can be understood as the UE uncorrectable error, that is, an error that the hardware cannot automatically correct or recover; the persistent storage error interface is ERST.
[0105] The operating system can parse the error log through the target event handling function. When it is determined that the error type of the error log is an uncorrectable error according to the parsing result, the error log is written into the flash memory through the ERST interface to permanently store the error log.
[0106] For example, a power management unit (PMU) error usually refers to hardware or software problems related to power management, such as abnormal voltage regulation, power state conversion failure, power consumption limit error, etc.; these errors may be caused by various reasons, including hardware failures, driver problems, or operating system errors; errors that exceed the capabilities of the error correction mechanism can be regarded as UEs.
[0107] For power management unit errors, the operating system obtains the hardware error log from the memory buffer specified in the hardware error source table through the target event handling function (such as handle_pmu_error); and uses the handle_pmu_error function to parse the error log and analyze the error information contained therein, such as the time when the error occurred, the power management unit functions involved, the error code, etc.
[0108] According to the parsing result, when it is determined that the error type of the error log is the first error type, that is, the UE uncorrectable error, the operating system writes the error log into the flash memory through ERST, so that even after the operating system restarts, the error log can still be accessed and analyzed to identify and solve power management related problems.
[0109] For example, after the operating system restarts, by obtaining the error log, specific power management error types are identified from the error log, such as abnormal voltage regulation, power state transition failure, etc., and relevant power management hardware, such as power supply units, batteries, power management chips, etc., are checked, so as to locate the cause of the error and solve it.
[0110] The error log collection method provided by the embodiments of this specification determines the error type of the error log, and for errors that the hardware cannot automatically correct or recover, writes the corresponding error log into the flash memory to permanently store the error log, ensuring that the error log will not be lost. Subsequently, the error log can be read from the flash memory and analyzed, so as to quickly locate the cause of the error and solve it.
[0111] In one or more embodiments of this specification, when it is determined that the error type of the error log is a correctable error, a user-mode process can be triggered according to a preset notification mechanism to collect and persistently store the error log. The specific implementation method is as follows:
[0112] After determining the error type of the error log according to the parsing result, it further includes:
[0113] The operating system, when it is determined that the error type of the error log is the second error type according to the parsing result, triggers the user-mode process to collect and persistently store the error log according to the preset notification mechanism.
[0114] Among them, the second error type can be understood as the CE correctable error, that is, an error that the hardware can automatically correct or recover.
[0115] Preset notification mechanism: The hardware error source driver actively triggers, triggering the RAS tracepoint to generate an error log tracking event (RAS event), so that when the user-mode process (Rasdaemon) monitors the error log tracking event, it records and persists it to the database to collect and persistently store the error log.
[0116] The operating system can parse the error log through the target event handling function. When it is determined according to the parsing result that the error type of the error log is a correctable error, according to the preset notification mechanism, it triggers the user-mode process to collect and persistently store the error log.
[0117] That is, although for correctable errors, the hardware can automatically correct or recover, it is still necessary to notify the user-mode process to collect and persistently store the corresponding error log, so that the user can understand and collect any error that has occurred in the hardware.
[0118] The error log collection method provided by the embodiments of this specification, by determining the error type of the error log, for errors that the hardware can automatically correct or recover, according to the preset notification mechanism, triggers the user-mode process to collect and persistently store the error log, so that the user can understand any error that has occurred in the hardware, thus facilitating subsequent maintenance processing.
[0119] Step 206: Load the persistent storage file system. When the system driver of the persistent storage file system determines that there is an error log in the flash memory, according to the preset notification mechanism, it triggers the user-mode process to collect and persistently store the error log.
[0120] Among them, the persistent storage file system is the PStore file system. PStore is a non-volatile storage mechanism used to store key information in the case of system crashes or exceptions, so as to facilitate troubleshooting and analysis after the system restarts.
[0121] Specifically, in the case of the operating system restarting, the PStore file system can be loaded. The system driver corresponding to the PStore file system can determine whether there is an error log in the flash memory through the ERST interface; when it is determined that there is an error log in the flash memory, the system driver corresponding to the PStore file system outputs and clears the error record, and triggers the RAS tracepoint corresponding to the APEI error log to form a trace event. The user-mode process records and persists the error log to the database when it monitors that the RAS tracepoint is triggered.
[0122] In one or more embodiments of the present specification, in the case of an operating system restart, the PStore file system can be loaded by mounting the file directory of the PStore file system, so as to facilitate subsequent acquisition of hardware error logs from the flash memory. The specific implementation is as follows:
[0123] The loading of the persistent storage file system includes:
[0124] The persistent storage file system is loaded by mounting the file directory of the persistent storage file system.
[0125] Among them, in the case where the persistent storage file system is the PStore file system, the file directory of the persistent storage file system can be understood as / sys / fs / postore, where sys is system, fs is file system, and / sys / fs / postore represents the pstore file system directory under the system directory and under the file system directory.
[0126] Specifically, in the case of an operating system restart, the PStore file system can be loaded by mounting the / sys / fs / postore file directory. The PStore file system provides a set of standard file operation interfaces, including opening files, reading and writing files, deleting files, etc.
[0127] The error log collection method provided by the embodiments of the present specification can load the persistent storage file system by mounting the file directory of the persistent storage file system in the case of an operating system restart, so that the system driver corresponding to the persistent storage file system can obtain hardware error logs from the flash memory through the ERST interface.
[0128] In one or more embodiments of the present specification, in the case of an operating system restart, the PStore file system can be loaded by mounting the file directory of the PStore file system, so as to facilitate subsequent acquisition of hardware error logs from the flash memory through the system driver corresponding to the PStore file system. The specific implementation is as follows:
[0129] The operating system includes a system driver of the persistent storage file system running in the kernel;
[0130] Before triggering the user-mode process to collect and persistently store the error logs according to the preset notification mechanism, it further includes:
[0131] The operating system, through the system driver of the persistent storage file system, uses the persistent storage error interface to obtain the error logs from the flash memory and clears the error logs in the flash memory.
[0132] Among them, "driver" generally refers to a device driver, which is a special program that enables a computer to communicate with a device. It is equivalent to the interface of the hardware. Only through this interface can the operating system control the operation of the hardware device.
[0133] Specifically, the operating system can obtain the APEI format error log previously written through the ERST interface from the flash memory in the hardware through the system driver of the persistent storage file system; and in the case where the system driver obtains the error log from the flash memory through the ERST interface, the error log in the flash memory can be cleared for future use.
[0134] In practical applications, the flash memory used for persistent error logging is only 256 kilobytes, while in the user-state persistent database, it is in the terabyte level. Therefore, the flash memory can only be used temporarily. It is necessary to take out the error log from the flash memory, trigger the user-state process to persist it to the database, and delete the error log in the flash memory.
[0135] The error log collection method provided by the embodiments of this specification can, in the case of the operating system restarting, obtain the hardware error log from the flash memory through the system driver of the persistent storage file system by using the ERST interface, so that the user-state process can collect and persistently store the error log, and clearing the error log in the flash memory can ensure that the storage space in the flash memory is not occupied.
[0136] In one or more embodiments of this specification, to enable the user-state process to collect and persistently store the error log, in the case where the user-state process can listen for the error log trace event, the system driver of the persistent storage file system can be used to trigger the error log trace point to generate the error log trace event. The specific implementation method is as follows:
[0137] Triggering the user-state process to collect and persistently store the error log according to the preset notification mechanism includes:
[0138] The operating system, through the system driver of the persistent storage file system, triggers the error log trace point to generate the error log trace event, so that the user-state process can collect and persistently store the error log when it listens for the error log trace event.
[0139] Among them, the error log trace point can be understood as a static probe point predefined by the kernel, which is triggered when the error log is taken out from the flash memory; the trace point is a mechanism used to track code execution in the operating system kernel, which allows developers to monitor and debug the runtime code without affecting the system performance; through the trace point, events can be inserted into the kernel and relevant data can be collected for analyzing and optimizing the system performance and stability, etc.
[0140] Specifically, when retrieving the hardware error log from the flash memory, the system driver of the persistent storage file system can trigger an error log tracepoint to generate an error log trace event, implementing the interaction between the user space and the kernel space, so that when the user space process detects the error log trace event, it can collect and persistently store the hardware error log.
[0141] In practical applications, when the user space process detects the error log trace event, it can collect and persistently store the error log in a database, so that the user can obtain the error log from the database, analyze the data log, locate the root cause, and avoid batch downtime.
[0142] For the error log collection method provided in the embodiments of this specification, the system driver of the persistent storage file system triggers an error log tracepoint to generate an error log trace event, enabling the user space process to collect and persistently store the error log, which is applicable to the reporting of error logs on multiple platforms. Subsequently, the root cause can be located by analyzing the error logs collected during downtime, avoiding batch downtime.
[0143] See Figure 3 , Figure 3 which shows the flowchart of the processing procedure of an error log collection method provided in an embodiment of this specification.
[0144] Step 302: Detect a hardware error and record the error information in the error log.
[0145] Specifically, when the hardware detects a hardware error, it records the error information corresponding to the hardware error in the error record register to form an error log.
[0146] Step 304: Trigger the error interrupt of the hardware.
[0147] Specifically, when the hardware detects a hardware error, it triggers the corresponding interrupt.
[0148] Step 306: In response to the error interrupt of the hardware, write the error log into the memory buffer specified by the hardware error source table.
[0149] Specifically, the firmware, in response to the error interrupt of the hardware, writes the error log of the hardware recorded in the error record register into the memory buffer specified by HEST according to the hardware platform error interface specification.
[0150] Step 308: Send a hardware exception notification to the operating system.
[0151] Specifically, the firmware can also send a hardware exception notification to the operating system so that the operating system can obtain the hardware error log based on the hardware exception notification.
[0152] Step 310: In response to the hardware exception notification, obtain the hardware error log from the memory buffer specified by the hardware error source table.
[0153] Specifically, in response to the hardware exception notification sent by the firmware, the operating system obtains the hardware error log from the memory buffer specified by the HEST.
[0154] Step 312: Write the error log into the flash memory through the persistent storage error interface.
[0155] Specifically, when the operating system determines that the error log meets the persistent storage condition, it writes the error log into the flash memory through the ERST interface.
[0156] Step 314: Restart and load the persistent storage file system.
[0157] Specifically, after the operating system writes the error log into the flash memory, it restarts and loads the Pstore file system.
[0158] Step 316: Through the system driver of the persistent storage file system, obtain the error log from the flash memory using the persistent storage error interface and clear the error log in the flash memory.
[0159] Specifically, the system driver corresponding to the Pstore file system in the operating system determines whether the error log exists in the flash memory through the ERST interface. If it exists, it obtains the error log from the flash memory using the ERST interface and clears the error log in the flash memory.
[0160] Step 318: Trigger the error log tracepoint to generate an error log trace event so that the user-mode process can collect and persistently store the error log when it monitors the error log trace event.
[0161] Specifically, the operating system can trigger the error log tracepoint to generate an error log trace event so that the user-mode process can collect and persistently store the error log when it monitors the error log trace event.
[0162] For the specific implementation, reference can be made to the above embodiments and will not be elaborated here.
[0163] The above is a schematic solution of an error log collection method according to this embodiment. It should be noted that the technical solution of this error log collection method and the technical solution of the error log collection method applied to the data processing device described above belong to the same concept. For the details not described in detail in the technical solution of the error log collection method, reference can be made to the description of the technical solution of the error log collection method applied to the data processing device above.
[0164] See Figure 4 , Figure 4 which shows a flowchart of an error log collection method applied to an operating system of a data processing device provided by an embodiment of this specification.
[0165] Step 402: In response to a hardware exception notification sent by the firmware, obtain the error log of the hardware from a specified memory buffer, where the firmware is installed on the hardware, and the hardware exception notification is that after the firmware responds to the error interrupt of the hardware and writes the error log of the hardware into the specified memory buffer according to the hardware platform error interface specification, it is sent.
[0166] The operating system includes a general hardware error source driver running in the kernel;
[0167] The obtaining the error log of the hardware from a specified memory buffer in response to a hardware exception notification sent by the firmware includes:
[0168] In response to the hardware exception notification sent by the firmware, trigger the general hardware error source driver to determine the corresponding target event processing function according to the notification type carried in the hardware exception notification;
[0169] Obtain the error log of the hardware from the memory buffer specified by the hardware error source table through the target event processing function.
[0170] Step 404: When the error log meets the persistent storage condition, write the error log into the flash memory and restart the operating system.
[0171] The writing the error log into the flash memory when the error log meets the persistent storage condition includes:
[0172] Parse the error log through the target event processing function and determine the error type of the error log according to the parsing result;
[0173] When it is determined that the error type of the error log is the first error type, write the error log into the flash memory through the persistent storage error interface.
[0174] Step 406: Load the persistent storage file system. When it is determined by the system driver of the persistent storage file system that there is an error log in the flash memory, according to a preset notification mechanism, trigger a user-mode process to collect and persistently store the error log.
[0175] The operating system includes a system driver of the persistent storage file system running in the kernel;
[0176] Before triggering the user-mode process to collect and persistently store the error log according to the preset notification mechanism, it further includes:
[0177] Through the system driver of the persistent storage file system, use the persistent storage error interface to obtain the error log from the flash memory and clear the error log in the flash memory.
[0178] Triggering the user-mode process to collect and persistently store the error log according to the preset notification mechanism includes:
[0179] Through the system driver of the persistent storage file system, trigger an error log tracepoint to generate an error log trace event, so that when the user-mode process monitors the error log trace event, it collects and persistently stores the error log.
[0180] For the specific implementation, reference can be made to the above embodiments, which will not be elaborated here.
[0181] The error log collection method provided in the embodiments of this specification is applied to the operating system of a data processing device. Before the operating system crashes, by writing the error log that conforms to the hardware platform error interface specification into the flash memory through the persistent storage error interface, it has universality and is applicable to the persistence of error logs on multiple platforms. After the operating system restarts, obtain the hardware error log from the flash memory, and trigger a user-mode process to collect and persistently store the error log through a preset notification mechanism, so as to realize the collection of error logs when the hardware crashes, which is applicable to the reporting of error logs on multiple platforms. Subsequently, by analyzing the collected error logs at the time of crashing, the root cause can be located to avoid batch crashes.
[0182] Corresponding to the above method embodiments, this specification also provides embodiments of a data processing device. Figure 5 It shows a schematic structural diagram of a data processing device provided in an embodiment of this specification. As Figure 5 shown, the data processing device includes hardware 502, firmware 504 installed on the hardware 502, and an operating system 506, where
[0183] The hardware 502 is configured to record error information corresponding to the hardware error in an error record register when a hardware error is detected, form an error log of the hardware 502, and trigger an error interrupt of the firmware 504;
[0184] The firmware 504 is configured to, in response to the error interrupt of the hardware 502, write the error log of the hardware 502 into a specified memory buffer according to a hardware platform error interface specification, and send a hardware exception notification to the operating system;
[0185] The operating system 506 is configured to, in response to the hardware exception notification sent by the firmware 504, obtain the error log of the hardware 502 from the specified memory buffer, write the error log into a flash memory when the error log meets the persistent storage condition, restart the operating system 506, load a persistent storage file system, and when a system driver of the persistent storage file system determines that there is the error log in the flash memory, trigger a user-mode process to collect and persistently store the error log according to a preset notification mechanism.
[0186] The data processing device provided in the embodiments of this specification has generality by writing an error log that conforms to the hardware platform error interface specification into a flash memory through a persistent storage error interface before the operating system crashes, is applicable to the persistence of error logs of multiple platforms, obtains the hardware error log from the flash memory after the operating system restarts, and triggers a user-mode process to collect and persistently store the error log through a preset notification mechanism, thereby implementing the collection of error logs when a hardware error causes a crash, being applicable to the reporting of error logs of multiple platforms, and subsequently being able to locate the root cause by analyzing the collected error logs at the time of the crash to avoid batch crashes.
[0187] The above is a schematic solution of a data processing device in this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the above error log collection method belong to the same concept. For the details not described in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above error log collection method.
[0188] Figure 6 FIG. shows a structural block diagram of a computing device 600 according to an embodiment of this specification. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.
[0189] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).
[0190] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 6 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.
[0191] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.
[0192] Wherein, the processor 620 is used to execute the following computer-executable instructions, which when executed by the processor, implement obtaining the error log of the hardware from a specified memory buffer in response to the hardware exception notification sent by the firmware,
[0193] Among them, the firmware is installed in the hardware, and the hardware exception notification is sent after the firmware writes the error log of the hardware into the specified memory buffer according to the hardware platform error interface specification in response to an error interrupt of the hardware;
[0194] When the error log meets the persistent storage condition, write the error log into the flash memory and restart the operating system;
[0195] Load the persistent storage file system. When the system driver of the persistent storage file system determines that there is an error log in the flash memory, trigger the steps of collecting and persistently storing the error log by a user-mode process according to a preset notification mechanism.
[0196] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above error log collection method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the error log collection method applied to a data processing device above.
[0197] This specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above error log collection method applied to a data processing device.
[0198] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the above error log collection method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the error log collection method applied to a data processing device above.
[0199] This specification also provides a computer program, which, when executed on a computer, causes the computer to execute the steps of the above error log collection method applied to a data processing device.
[0200] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above error log collection method applied to a data processing device belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the error log collection method applied to a data processing device above.
[0201] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0202] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0203] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0204] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0205] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. An error log collection method is applied to a data processing device, which includes an operating system, hardware, and firmware installed on the hardware. Among them, the operating system, in response to a hardware exception notification sent by the firmware, obtains the error log of the hardware from a specified memory buffer, wherein, the hardware exception notification is sent after the firmware writes the error log of the hardware into the specified memory buffer in accordance with the hardware platform error interface specification in response to an error interrupt of the hardware; when the error log meets the persistent storage condition, write the error log into the flash memory and restart the operating system; load the persistent storage file system, and when the system driver of the persistent storage file system determines that there is an error log in the flash memory, trigger a user-mode process to collect and persistently store the error log according to a preset notification mechanism.
2. According to the error log collection method described in claim 1, before the step of responding to the hardware exception notification sent by the firmware, it further includes: when the hardware detects a hardware error, record the error information corresponding to the hardware error in an error record register to form the error log of the hardware, and trigger an error interrupt of the firmware.
3. According to the error log collection method described in claim 2, after the step of triggering the error interrupt of the firmware, it further includes: the firmware, in response to the error interrupt of the hardware, writes the error log of the hardware recorded in the error record register into a memory buffer specified by a hardware error source table in accordance with the hardware platform error interface specification, and sends the hardware exception notification to the operating system.
4. According to the error log collection method described in claim 1, the operating system includes a general hardware error source driver running in the kernel; the step of obtaining the error log of the hardware from a specified memory buffer in response to the hardware exception notification sent by the firmware includes: the operating system, in response to the hardware exception notification sent by the firmware, triggers the general hardware error source driver to determine a corresponding target event handling function according to the notification type carried in the hardware exception notification; obtain the error log of the hardware from the memory buffer specified by the hardware error source table through the target event handling function.
5. According to the error log collection method described in claim 4, the step of writing the error log into the flash memory when the error log meets the persistent storage condition includes: the operating system parses the error log through the target event handling function and determines the error type of the error log according to the parsing result; when it is determined that the error type of the error log is the first error type, write the error log into the flash memory through a persistent storage error interface.
6. According to the error log collection method described in claim 5, after determining the error type of the error log according to the parsing result, it further includes: When the operating system determines that the error type of the error log is the second error type according to the parsing result, it triggers the user-mode process to collect and persistently store the error log according to the preset notification mechanism.
7. The error log collection method according to claim 1, wherein the loading of the persistent storage file system includes: Loading the persistent storage file system by mounting the file directory of the persistent storage file system.
8. The error log collection method according to claim 1, wherein the operating system includes a system driver of the persistent storage file system running in the kernel; Before triggering the user-mode process to collect and persistently store the error log according to the preset notification mechanism, it further includes: The operating system obtains the error log from the flash memory through the system driver of the persistent storage file system and clears the error log in the flash memory by using the persistent storage error interface.
9. The error log collection method according to claim 1, wherein triggering the user-mode process to collect and persistently store the error log according to the preset notification mechanism includes: The operating system triggers an error log trace point to generate an error log trace event through the system driver of the persistent storage file system, so that the user-mode process collects and persistently stores the error log when it monitors the error log trace event.
10. An error log collection method applied to the operating system of a data processing device, including: Responding to a hardware exception notification sent by the firmware, obtaining the error log of the hardware from a specified memory buffer, wherein the firmware is installed on the hardware, and the hardware exception notification is sent after the firmware writes the error log of the hardware into the specified memory buffer according to the hardware platform error interface specification in response to an error interrupt of the hardware; Writing the error log into the flash memory and restarting the operating system when the error log meets the persistent storage condition; Loading the persistent storage file system, and when the system driver of the persistent storage file system determines that there is an error log in the flash memory, triggering the user-mode process to collect and persistently store the error log according to the preset notification mechanism.
11. The error log collection method according to claim 10, wherein the operating system includes a general hardware error source driver running in the kernel; The obtaining the error log of the hardware from the specified memory buffer in response to the hardware exception notification sent by the firmware includes: Responding to the hardware exception notification sent by the firmware, triggering the general hardware error source driver to determine a corresponding target event processing function according to the notification type carried in the hardware exception notification; Obtaining the error log of the hardware from the memory buffer specified in the hardware error source table through the target event processing function.
12. The error log collection method according to claim 11, wherein writing the error log into the flash memory when the error log meets the persistent storage condition includes: Parse the error log through the target event handling function, and determine the error type of the error log according to the parsing result; In the case where the error type of the error log is determined to be the first error type, write the error log into the flash memory through the persistent storage error interface.
13. The error log collection method according to claim 10, wherein the operating system includes a system driver of the persistent storage file system running in the kernel; Before triggering the user-mode process to collect and persistently store the error log according to the preset notification mechanism, it further includes: Through the system driver of the persistent storage file system, obtain the error log from the flash memory by using the persistent storage error interface, and clear the error log in the flash memory.
14. The error log collection method according to claim 10, wherein triggering the user-mode process to collect and persistently store the error log according to the preset notification mechanism includes: Through the system driver of the persistent storage file system, trigger the error log trace point to generate an error log trace event, so that the user-mode process collects and persistently stores the error log when it monitors the error log trace event.
15. A data processing device, the data processing device includes an operating system, hardware, and firmware installed on the hardware, wherein, The hardware is used to record the error information corresponding to the hardware error in the error record register when a hardware error is detected, form the error log of the hardware, and trigger an error interrupt of the firmware; The firmware is used to respond to the error interrupt of the hardware, write the error log of the hardware into a specified memory buffer according to the hardware platform error interface specification, and send a hardware exception notification to the operating system; The operating system is used to respond to the hardware exception notification sent by the firmware, obtain the error log of the hardware from the specified memory buffer, write the error log into the flash memory when the error log meets the persistent storage condition, restart the operating system, load the persistent storage file system, and when the system driver of the persistent storage file system determines that there is an error log in the flash memory, trigger the user-mode process to collect and persistently store the error log according to the preset notification mechanism.