Hot swap error reporting method, processor architecture, device and storage medium

By introducing data interface services between BIOS, BMC and OS on the ARM platform, the problem of ARM platform lacking separate management when handling hot-swap error reports is solved, effectively blocking hot-swap error reports information is achieved, and user experience and system stability are improved.

WO2025123553A1PCT designated stage expired Publication Date: 2025-06-19INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/089641
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-04-24
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

When ARM platform handles PCIE CE errors, especially hot plug errors, and lacks a separate management design, which makes it difficult to block hot plug errors, and then falsely report unnecessary information, affecting system stability and user experience.

Method used

By introducing data interface services between BIOS, BMC and OS on the ARM platform, breaking the data transmission barrier. The OS determines whether the PCIE CE error message is hot-swap error message based on the device data collected by the BMC and the preset hot-swap error message, and sends blocking information to the BMC to delete the hot-swap error message.

Benefits of technology

It realizes separate management and blocking of hot-swap error reporting information, reduces unnecessary error reporting, significantly improves user experience, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024089641_19062025_PF_FP_ABST
    Figure CN2024089641_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a hot swap error reporting method, a processor architecture, a device and a storage medium. The method is applied to an ARM platform, the ARM platform comprising a BMC, an OS and a BIOS. On the basis of a data interface service configured by the BIOS, data transmission is performed between the BMC and the OS. During operation of the ARM platform, after an SCP in the BIOS detects PCIE CE type error reporting information corresponding to any interface, the SCP sends the PCIE CE type error reporting information to the BMC and the OS; the OS triggers the BMC to acquire device data; on the basis of the device data acquired by the BMC, the error reporting condition of the PCIE CE type error reporting information corresponding to any interface and a preset hot swap error reporting parameter, the OS determines whether the PCIE CE type error reporting information corresponding to any interface is hot swap error reporting information; and when the PCIE CE type error reporting information is hot swap error reporting information, the OS sends masking information corresponding to the hot swap error reporting information to the BMC. The embodiments of the present application aim to improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Hot plug error reporting method, processor architecture, device and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 12, 2023, with application number 202311704052.2, and entitled “A hot-swap error reporting method, processor architecture, device and storage medium,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] Embodiments of the present application relate to the technical field of data processing, and in particular, to a hot plug error reporting method, processor architecture, device, and storage medium. Background Art

[0004] The current ARM (Advanced RISC Machine, a processor architecture) platform processor firmware is mainly divided into two parts: SCP (system control processor) and UEFI (Unified Extensible Firmware Interface). Major hardware manufacturers will complete the RAS (Reliability, Availability and Serviceability, an indicator for evaluating system performance) function in the SCP part. Therefore, due to the limitations of hardware component manufacturers and the ARM platform, it is difficult to customize and modify the main functions of RAS.

[0005] The ARM platform's RAS function lacks a mechanism for handling PCIE CE (peripheral component interconnect express correctable errors), especially the common hot-plug function, which lacks a separate, differentiated design. The ARM platform provides a general control and management solution for PCIE CE, but can only regulate the broad category of PCIE CE errors. PCIE CE includes not only hot-plug errors but also other types of error information. Therefore, the current ARM platform is subject to hardware component limitations, making it difficult to shield hot-plug errors, leading to false reports of unnecessary information such as hot-plug errors. This not only affects stable use and overall server evaluation, but also significantly reduces the user's stability experience, resulting in unnecessary warranty and disputes.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide a hot plug error reporting method, processor architecture, device and storage medium, aiming to improve the user experience.

[0008] In a first aspect, an embodiment of the present application provides a hot-plug error reporting method, which is applied to an ARM platform. The ARM platform includes a BMC (Baseboard Management Controller), an OS (Operating System), and a BIOS (Basic Input Output System). Data is transmitted between the BMC and the OS based on a data interface service configured by the BIOS. The method includes:

[0009] During the operation of the ARM platform, when the SCP in the BIOS detects a PCIE CE error message corresponding to any interface, the SCP sends the PCIE CE error message to the BMC and OS.

[0010] The OS triggers the BMC to collect device data;

[0011] The OS determines whether the PCIE CE error information corresponding to any interface is hot-plug error information based on the device data collected by the BMC, the error status of the PCIE CE error information corresponding to any interface, and the preset hot-plug error parameters;

[0012] When the PCIE CE error message is hot-plug error message, the OS sends shielding information corresponding to the hot-plug error message to the BMC.

[0013] In some embodiments, when the PCIE CE error information is hot plug error information, after the OS sends shielding information corresponding to the hot plug error information to the BMC, the method further includes:

[0014] The BMC deletes the hot plug error information in response to the shielding information corresponding to the hot plug error information.

[0015] In some embodiments, the method further comprises:

[0016] The BMC collects device data at preset intervals and sends it to the OS.

[0017] In some embodiments, the hot plug error parameter includes a hot plug error interval. The OS determines whether the PCIE CE error information corresponding to any interface is hot plug error information based on device data collected by the BMC, the error status of the PCIE CE error information corresponding to any interface, and preset hot plug error parameters, including:

[0018] For the PCIE CE error information corresponding to any interface, the OS determines the error duration between the first error reporting time and the last error reporting time of the PCIE CE error information;

[0019] When the error duration is less than the hot-plug error interval, and the OS determines, based on the device data collected by the BMC, that there is device change information for the interface within the error duration, the PCIE CE error information corresponding to the interface is determined as hot-plug error information, where the device change information includes adding or removing a device.

[0020] In some embodiments, the hot plug error parameter includes a hot plug error threshold. The OS determines whether the PCIE CE error information corresponding to any interface is hot plug error information based on the device data collected by the BMC, the error status of the PCIE CE error information corresponding to any interface, and the preset hot plug error parameter, including:

[0021] For any PCIE CE error message corresponding to an interface, the OS counts the number of PCIE CE error messages reported within a specified time period.

[0022] If the number of errors of the PCIE CE type error information within the calibration time period is less than the hot plug error threshold, and the OS determines that there is device change information of the interface within the calibration time period based on the device data collected by the BMC, the PCIE CE type error information corresponding to the interface is determined as hot plug error information, where the device change information includes adding or removing a device.

[0023] In some embodiments, the method further comprises:

[0024] During the boot process of the ARM platform, the data interface service is registered on the OS and BMC through the BIOS of the ARM platform. The data interface service is used to provide a data transmission interface between the BMC and the OS through the BIOS.

[0025] In some embodiments, during the boot process of the ARM platform, registering the data interface service on the OS and BMC through the BIOS of the ARM platform includes:

[0026] After the ARM platform server is powered on, the BIOS registers the data interface service on the BMC;

[0027] After entering the OS, a first hot-swap error reporting management program in the OS is run, and the first hot-swap error reporting management program accesses a data interface service registered with the BIOS;

[0028] The OS sends activation information to the BMC through the data transmission interface provided by the data interface service;

[0029] The BMC starts a second hot-swap error management program stored in the BMC in response to the activation information.

[0030] In some embodiments, after the ARM platform server is powered on, the method further includes:

[0031] The BIOS sends the hot-plug error parameters currently stored in the BIOS to the BMC.

[0032] In some embodiments, after entering the OS and running the first hot-plug error management program in the OS, the method further includes:

[0033] The first hot-plug error reporting management program accesses and obtains the hot-plug error reporting parameters currently stored in the BIOS.

[0034] In some embodiments, the method further comprises:

[0035] After the ARM platform server is powered on, the BMC checks whether it has stored the hot-plug error parameters obtained from the BIOS.

[0036] The BIOS sends its currently stored hot-plug error parameters to the BMC and registers the data interface service. The data interface service is used to provide a data transmission interface between the BMC and the OS through the BIOS.

[0037] After entering the OS, a first hot-swap error reporting management program in the OS is run, the first hot-swap error reporting management program accesses a data interface service registered with the BIOS, and obtains hot-swap error reporting parameters currently stored in the BIOS;

[0038] The OS sends activation information to the BMC through the data transmission interface provided by the data interface service;

[0039] The BMC starts a second hot-swap error management program stored in the BMC in response to the activation information.

[0040] In some embodiments, before entering the OS, the method further includes:

[0041] BIOS monitors the modification operation of hot-plug error parameters in real time;

[0042] When there is a modification operation of the hot-plug error reporting parameters, the BIOS sends the modified hot-plug error reporting parameters to the OS and BMC respectively.

[0043] In some embodiments, before entering the OS, the method further includes:

[0044] In response to the closing operation of the hot plug error setting option, the BIOS sends hot plug alarm closing information to the OS and the BMC respectively;

[0045] The OS stops executing the first hot plug error management program in response to the hot plug alarm shutdown information;

[0046] The BMC stops executing the second hot-plug error management program in response to the hot-plug alarm shutdown information.

[0047] In some embodiments, the method further comprises:

[0048] During the operation of the ARM platform, the OS responds to the change operation of the hot-plug error reporting parameter, stores the changed hot-plug error reporting parameter, and sends the changed hot-plug error reporting parameter to the BIOS and the BMC respectively.

[0049] In some embodiments, during the next startup of the ARM platform, the method further includes:

[0050] After the ARM platform server is powered on, the BIOS sends the modified hot-plug error parameters stored in its own storage to the BMC and OS respectively.

[0051] In some embodiments, the method further comprises:

[0052] In response to the upgrade operation of the hypervisor, the OS obtains the first hot-plug error reporting hypervisor to be upgraded, and upgrades the current first hot-plug error reporting hypervisor to the first hot-plug error reporting hypervisor to be upgraded.

[0053] In some embodiments, after the OS determines whether the PCIE CE error information is hot plug error information based on device data collected by the BMC, the number of PCIE CE error information errors, and preset hot plug error parameters, the method further includes:

[0054] When the PCIE CE error message is not a hot-swap error message, the OS sends normal processing information corresponding to the PCIE CE error message to the BMC;

[0055] The BMC records and reports the PCIE CE error information in response to the normal processing information corresponding to the PCIE CE error information.

[0056] In some embodiments, when the PCIE CE error message is hot plug error message, the method further includes:

[0057] The OS stores the PCIE CE error information and the timestamp of the PCIE CE error information in the hot plug summary list, so that all hot plug error information can be summarized and viewed.

[0058] In some embodiments, the method further comprises:

[0059] In response to the query operation of the hot plug error, the OS generates a visual hot plug error chart according to the hot plug summary list for display.

[0060] In a second aspect, an embodiment of the present application provides a processor architecture, which includes a BMC, an OS, and a BIOS. The processor architecture is used to execute the hot-plug error reporting method of the first aspect of the embodiment.

[0061] In a third aspect, an embodiment of the present application provides a computer device comprising: at least one processor, and a memory, wherein the memory stores a computer program that can be run on the processor, wherein when the processor executes the computer program, the hot plug error reporting method of the first aspect of the embodiment is executed.

[0062] In a fourth aspect, an embodiment of the present application provides a non-volatile readable storage medium, which stores a computer program, wherein when the computer program is executed by a processor, the hot plug error reporting method of the first aspect of the embodiment is executed. Beneficial effects:

[0063] During operation of the ARM platform, when the SCP in the BIOS detects PCIE CE error information corresponding to any interface, the SCP sends the PCIE CE error information to the BMC and OS; the OS triggers the BMC to collect device data; the OS determines whether the PCIE CE error information corresponding to any interface is hot-plug error information based on the device data collected by the BMC, the error status of the PCIE CE error information corresponding to any interface, and preset hot-plug error parameters; when the PCIE CE error information is hot-plug error information, the OS sends shielding information corresponding to the hot-plug error information to the BMC.

[0064] The hot-plug error reporting method provided by this method is not subject to the architectural limitations of device manufacturers and the ARM platform itself. It breaks the data transmission barrier between the OS and the BMC by configuring a data interface service for data transmission between the OS and the BMC. The OS then determines whether the PCIE CE error message is a hot-plug error message. When the PCIE CE error message is a hot-plug error message, the OS sends shielding information corresponding to the hot-plug error message to the BMC, thereby shielding the hot-plug error message and eliminating the need to report unnecessary hot-plug errors to the user. This can significantly improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate some embodiments of the present application or technical solutions in the prior art, the following briefly introduces the drawings required for describing some embodiments or prior art.

[0066] FIG1 is a flowchart showing a hot plug error reporting method according to some embodiments of the present application;

[0067] FIG2 shows a structural topology diagram of an ARM platform provided in some embodiments of the present application;

[0068] FIG3 is a schematic diagram showing a processor architecture provided by some embodiments of the present application;

[0069] FIG4 shows a schematic diagram of a computer device provided by some embodiments of the present application;

[0070] FIG5 shows a schematic diagram of a non-volatile readable storage medium provided in some embodiments of the present application. DETAILED DESCRIPTION

[0071] The technical solutions in some embodiments of the present application will be described below in conjunction with the drawings in some embodiments of the present application.

[0072] In order to make the purpose, technical solutions and advantages of some embodiments of the present application clearer, each embodiment of the present application will be described in detail below with reference to the accompanying drawings. However, it will be understood by those skilled in the art that in each embodiment of the present application, many technical details are proposed to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation on the specific implementation of the present application. The various embodiments can be combined with each other and referenced to each other under the premise that there is no contradiction.

[0073] ARM: Advanced RISC Machine, a processor architecture;

[0074] CPU: Central Processing Unit, central processing unit;

[0075] BMC: Baseboard Management Controller;

[0076] OS: Operating System, the most common OS in servers is Linux.

[0077] BIOS: Basic Input Output System, generally refers to UEFI;

[0078] UEFI: Unified Extensible Firmware Interface, unified extensible firmware interface;

[0079] SCP: system control processor, system control processor;

[0080] RAS: Reliability, Availability, and Serviceability, which are indicators for evaluating system performance. They include reliability, availability, and maintainability. They primarily refer to the Machine Check Architecture (MCA) mechanism, which is used to detect hardware errors.

[0081] PCIE CE: peripheral component interconnect express Correctable errors, memory correctable errors.

[0082] Hot swapping means inserting or removing modules or boards into or out of a system without shutting down the system power supply, thereby improving the system's reliability, rapid maintainability, redundancy, and timely disaster recovery capabilities.

[0083] In related technologies, the ARM platform only provides a general control and management solution for PCIE CE hot-plug errors. This means that PCIE CE errors on the current ARM platform include not only hot-plug errors but also other types of errors. Hot-plug errors cannot be managed separately. Hot-plug errors are unnecessary information for high-performance requirements. Including hot-plug errors in PCIE CE error messages requires professional technicians to identify and eliminate false hot-plug errors during troubleshooting. False alarms can also result in corresponding error records in the BMC, reducing machine stability. The need for professional technicians to distinguish between them increases maintenance difficulty, maintenance costs, and production costs. This not only impacts stable use and overall server evaluation, but also significantly reduces user stability, leading to unnecessary warranty and disputes.

[0084] Based on this, some embodiments of the present application provide a hot-plug error reporting method, which can manage hot-plug error reporting separately, reduce unnecessary error reporting, and thus improve the user experience.

[0085] 1 , a flowchart of a hot-plug error reporting method provided by some embodiments of the present application is shown. The method is applied to an ARM platform, which includes a BMC, an OS, and a BIOS. In this embodiment, data is transmitted between the BMC and the OS based on a data interface service configured by the BIOS, breaking the data transmission barrier between the OS and the BMC. The method may specifically include the following steps:

[0086] S101: During operation of the ARM platform, when the SCP in the BIOS detects PCIE CE error information corresponding to any interface, the SCP sends the PCIE CE error information to the BMC and the OS.

[0087] In the ARM platform, the BMC includes software based on independent hardware that runs on the server as soon as it is powered on. As long as the server is plugged in, the BMC software can run quickly. The OS and BIOS are programs executed by the CPU of the ARM platform. The BIOS can be divided into SCP and UEFI. Among them, SCP is a program running on a separate small core within the CPU. SCP is used to monitor PCIE CE error information. Specifically, when a device such as a hard disk device is triggered to be removed, the server's end node (end point) will report the PCIE CE error information to the root node (root point). SCP manages the root point and has the function of filtering PCIE CE error information. When reporting PCIE CE error information, SCP sets the register in the CPU. The BMC can obtain the PCIE CE error information by reading the register status. At the same time, the CPU can send PCIE CE error information to the OS through the ACPI interface (Advanced Configuration and Power Management Interface). Each PCIE CE error message carries the corresponding error interface information, which means that the reporting PCIE CE error information interface.

[0088] In this embodiment, when the SCP detects PCIE CE error information corresponding to any interface during operation on the ARM platform, the PCIE CE error information may be hot-plug error information. Therefore, the PCIE CE error information is further judged. However, since the SCP is the underlying firmware of the ARM platform, the CPU manufacturer provides a fixed SCP management policy that is difficult to modify. Therefore, in order to separate the hot-plug error information from the PCIE CE error information, in this embodiment, the PCIE CE error information detected by the SCP is sent to the BMC and the OS at the same time.

[0089] S102: The OS triggers the BMC to collect device data.

[0090] In the ARM platform based on this embodiment, the OS can perform data transmission with the BMC, breaking the data transmission barrier between the two.

[0091] In some feasible implementations, during the booting process of the ARM platform, a data interface service can be registered on the OS and BMC through the BIOS of the ARM platform. The data interface service is used to provide a data transmission interface between the BMC and the OS through the BIOS.

[0092] Specifically, a first hot-plug error management program can be pre-configured in the OS, a second hot-plug error management program can be configured in the BMC, and a hot-plug error setting option can be pre-configured in the BIOS. Through the hot-plug error setting option, the user can independently choose whether to execute the process of shielding hot-plug errors, and the user is allowed to define hot-plug error parameters.

[0093] After the ARM platform server is powered on, the BIOS registers the data interface service on the BMC and sends its currently stored hot plug error parameters to the BMC.

[0094] After entering the OS, the first hot-swap error reporting management program in the OS is run. The first hot-swap error reporting management program accesses the data interface service registered with the BIOS and obtains the hot-swap error reporting parameters currently stored in the BIOS. The OS sends activation information to the BMC through the data transmission interface provided by the data interface service. In response to the activation information, the BMC starts the second hot-swap error reporting management program stored in the BMC. At this time, data transmission is possible between the BMC and the OS.

[0095] Then, when the OS receives the PCIE CE error message sent by the SCP, it triggers the BMC to collect device data. In addition to starting to collect device data after being triggered, in other implementations, the BMC can also collect device data at preset intervals and send it to the OS.

[0096] Because the BMC has a data collection channel independent of the BIOS and has special hardware support, for example, the BMC can directly access the information of PCIE devices on the server through the I2C bus (Inter-Integrated Circuit), such as reading hard disk information, so as to obtain the actual online status of the device. The collected device data is then packaged and passed to the OS through the data interface service registered with the BIOS.

[0097] S103: The OS determines whether the PCIE CE error information corresponding to any interface is hot plug error information based on the device data collected by the BMC, the error status of the PCIE CE error information corresponding to any interface, and preset hot plug error parameters.

[0098] Hot-plug error messages are reported as PCIE CE error messages when the device is unplugged. Hot-plug error messages are also reported as PCIE CE error messages when the device is plugged in due to contact problems such as installation jitter. However, hot-plug error messages will no longer be reported after the device plug-in and unplugging action is completed. Therefore, you can set hot-plug error parameters and combine the device data collected by the BMC to determine whether the PCIE CE error message corresponding to any interface is a hot-plug error message. The hot-plug error parameter can be a hot-plug error interval or a hot-plug error threshold. The hot-plug error interval refers to the time interval between different error reporting times of PCIE CE error messages of any interface, and the hot-plug error threshold refers to the threshold of the number of PCIE CE error messages of any interface.

[0099] In some feasible implementations, whether the PCIE CE error information is hot plug error information may be determined based on the hot plug error interval.

[0100] Specifically, for the PCIE CE type error information corresponding to any interface, the OS determines the error duration between the first error time and the last error time of the PCIE CE type error information; when the error duration is less than the hot plug error interval, and the OS determines based on the device data collected by the BMC that there is device change information on the interface within the error duration, the PCIE CE type error information corresponding to the interface is determined as hot plug error information.

[0101] When hot-plugging any device, the hot-plugging action is completed in a short time, such as within 1 second. The interface that performs the hot-plugging operation within 1 second will report a PCIE CE error message. However, if the OS determines that the PCIE CE error message of any interface is still being reported after 1 second, it indicates that the PCIE CE error message is not caused by the hot-plugging action, but may be caused by poor contact between the device and the interface, etc., which causes the PCIE CE error message to be continuously reported.

[0102] In addition, hot-plugging will inevitably cause the interface to add or remove devices. Therefore, when determining whether the PCIE CE error information is a hot-plugging error information, it is also necessary to analyze whether there is any device change information on the interface during the error period based on the device data collected by the BMC. Device change information includes adding or removing devices, because the device may still be connected to the interface, but the PCIE circuit between the device and the interface is disconnected. This situation does not belong to the error caused by the hot-plugging operation. Therefore, the BMC needs to obtain device change information through the I2C bus to rule out the situation where the device is still connected but the PCIE circuit is disconnected.

[0103] The hot-plug error reporting interval in this embodiment can be customized according to the needs of actual application, and this embodiment does not impose any limitation.

[0104] In some other feasible implementations, it may also be possible to determine whether the PCIE CE error information is hot plug error information based on a hot plug error threshold.

[0105] Specifically, for the PCIE CE type error information corresponding to any interface, the OS counts the number of errors of the PCIE CE type error information within the calibration time period; if the number of errors of the PCIE CE type error information within the calibration time period is less than the hot plug error threshold, and the OS determines based on the device data collected by the BMC that there is device change information of the interface within the calibration time period, that is, when the device is added or reduced, the PCIE CE type error information corresponding to the interface is determined as hot plug error information.

[0106] Since the error frequency is generally fixed at the factory, the number of PCIE CE error messages within a calibrated time period can be counted to determine whether the number of errors within the calibrated time period is less than the hot plug error threshold. For example, the hot plug error threshold can be set to 5. If the number of PCIE CE error messages corresponding to any interface is less than 5 within 1 second, it indicates that no error is reported within 1 second due to the completion of the hot plug operation. If the number of errors is greater than or equal to 5, it indicates that the error is not caused by the hot plug operation, that is, the PCIE CE error message is not a hot plug error message.

[0107] The hot-plug error threshold and the size of the calibration time period in this embodiment can be customized according to the needs of actual application, and this embodiment does not impose any restrictions.

[0108] S104: When the PCIE CE error message is hot-plug error message, the OS sends shielding information corresponding to the hot-plug error message to the BMC.

[0109] After the BMC responds to the shielding information corresponding to the hot-plug error information, it may delete the hot-plug error information of the corresponding interface.

[0110] If the PCIE CE error message is not a hot-swap error message, the OS may allow normal reporting of the PCIE CE error message.

[0111] In some feasible implementations, when the PCIE CE error information is not hot plug error information, the OS may further send normal processing information corresponding to the PCIE CE error information to the BMC; the BMC records and reports the PCIE CE error information in response to the normal processing information corresponding to the PCIE CE error information.

[0112] Although hot-plug error information is no longer reported, the OS can store PCIE CE error information and the timestamp of PCIE CE error information in the hot-plug summary list to facilitate the summary and viewing of all hot-plug error information. When maintenance personnel want to view hot-plug error information, they can perform a query operation on the OS. In response to the hot-plug error query operation, the OS generates a visual hot-plug error chart based on the hot-plug summary list for display.

[0113] This embodiment combines the BIOS, BMC and OS of the ARM platform based on their respective partial functions, thereby realizing functions that are difficult to complete with a single part, that is, realizing the function of separately managing hot-plug type errors that cannot be achieved separately on the ARM platform.

[0114] Referring to Figure 2, a structural topology diagram of the ARM platform provided by some embodiments of the present application is shown. Although the UEFI part of the BIOS cannot directly control PCIE CE type errors, as the main user interaction part, hot plug error setting options and hot plug error parameters can be pre-configured in the BIOS through UEFI, so that users can manage the hot plug error judgment process through common operation forms, and the BIOS will send the hot plug error setting options and hot plug error parameters to the BMC and OS respectively; the SCP in the BIOS is used to obtain PCIE CE type error information and send it to the BMC and OS, and the OS and BMC can exchange data based on the data interface service of the BIOS.

[0115] As the primary user interaction channel, the OS is the system most easily interacted with by users during normal business operations. Therefore, this feature is leveraged to place the main interactive part of the process on the OS side. By aggregating the hot-plug error setting options and hot-plug error parameters provided by the BIOS and the device data collected by the BMC into the OS, the OS determines whether the PCIE CE error message is a hot-plug error message. The OS then feeds back the judgment result to the BMC through the data interface service, informing the BMC that the PCIE CE error message is a hot-plug error message, rather than a normal PCIE CE error message. The BMC then deletes the hot-plug error message to avoid false error messages on the BMC. At the same time, the OS can aggregate hot-plug error messages, making it easier for users to view the error messages of the current hot-plug and easier for engineers to confirm information during debugging.

[0116] Moreover, since the judgment of the main hot-plug error information is located on the OS side, based on the characteristic of easy updating of the OS, the first hot-plug error management program on the OS can also be updated and upgraded in a timely manner, and thus there is no need to rely on the update of the firmware. Specifically, in response to the upgrade operation of the management program, the OS obtains the first hot-plug error management program to be upgraded, and upgrades the current first hot-plug error management program to the first hot-plug error management program to be upgraded.

[0117] At the same time, as the most important user interaction channel, the OS side can also directly update or adjust the hot-plug error reporting parameters through the OS, and then synchronize them to the BIOS and BMC. For example, during the operation of the ARM platform, the OS responds to the change operation of the hot-plug error reporting parameters, stores the changed hot-plug error reporting parameters, and sends the changed hot-plug error reporting parameters to the BIOS and BMC respectively. After the hot-plug error reporting parameters updated or adjusted by the OS are synchronized to the BIOS, during the next ARM platform startup process, when the ARM platform server is powered on, the BIOS still sends the hot-plug error reporting parameters changed by the user through the OS to the BMC and OS, so that the user can configure the hot-plug error reporting parameters on both the BIOS and the OS.

[0118] In some feasible implementations, after the ARM platform server is powered on, the BMC detects whether it has stored the hot-plug error reporting parameters obtained from the BIOS and waits for the registration of the BIOS data interface service; then the BIOS sends its currently stored hot-plug error reporting parameters to the BMC and registers the data interface service, and the BIOS monitors the modification operations of the hot-plug error reporting parameters in real time; when there is a modification operation of the hot-plug error reporting parameters, the BIOS sends the modified hot-plug error reporting parameters to the BMC respectively.

[0119] After entering the OS, a first hot-plug error reporting management program in the OS is run. The first hot-plug error reporting management program accesses the data interface service registered with the BIOS and obtains the modified hot-plug error reporting parameters in the BIOS.

[0120] The OS sends activation information to the BMC through the data transmission interface provided by the data interface service; the BMC responds to the activation information and starts the second hot-swap error management program stored in the BMC.

[0121] In some feasible implementations, before entering the OS, in response to the shutdown operation of the hot plug error setting option, the BIOS sends the hot plug error shutdown information to the OS and the BMC respectively; the OS stops executing the first hot plug error management program in response to the hot plug error shutdown information; the BMC stops executing the second hot plug error management program in response to the hot plug error shutdown information, thereby stopping the hot plug error information judgment process.

[0122] The hot-plug error reporting method provided in this embodiment breaks the data transmission barrier between the OS and the BMC through the data interface service configured between the OS and the BMC for data transmission. The OS then determines whether the PCIE CE error information is hot-plug error information. When the PCIE CE error information is hot-plug error information, the OS sends shielding information corresponding to the hot-plug error information to the BMC, thereby shielding the hot-plug error information and eliminating the need to report unnecessary hot-plug errors to the user. This significantly improves the user experience and has the advantages of strong scalability, strong portability, and strong operability. It can be applied to different hardware platforms and is not limited to the separate management of hot-plug error information. It can also be applied to the management of other types of error information.

[0123] 3 , a schematic diagram of a processor architecture provided in some embodiments of the present application is shown. The processor architecture includes a BMC, an OS, and a BIOS. The processor architecture is used to execute the hot-plug error reporting method in some embodiments.

[0124] 4 , a schematic diagram of a computer device provided in accordance with some embodiments of the present application is shown. A computer device 400 includes at least one processor 401 and a memory 402 , wherein the memory 402 stores a computer program 403 that can be run on the processor 401 , wherein the processor 401 executes the hot-plug error reporting method of some embodiments when executing the computer program 403 .

[0125] 5 , a schematic diagram of a non-volatile readable storage medium provided in some embodiments of the present application is shown. The non-volatile readable storage medium 500 stores a computer program 501 , wherein the computer program 501 , when executed by a processor, executes the hot plug error reporting method of some embodiments.

[0126] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0127] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0128] The present application embodiment is described with reference to the flow chart and / or block diagram of the method, terminal device (system), and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flow chart and / or block diagram and the combination of the process and / or box in the flow chart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for realizing the function specified in one process or multiple processes and / or one box or multiple boxes of the flow chart.

[0129] These computer program instructions may also be stored in a computer non-volatile readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer non-volatile readable storage medium produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0131] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0132] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0133] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A hot plug error reporting method, characterized in that: Applied to an ARM platform, the ARM platform includes a BMC, an OS and a BIOS, the BMC and the OS perform data transmission based on a data interface service configured by the BIOS, and the method includes: During the operation of the ARM platform, when the SCP in the BIOS detects PCIE CE error information corresponding to any interface, the SCP sends the PCIE CE error information to the BMC and the OS; The OS triggers the BMC to collect device data; The OS determines whether the PCIE CE error information corresponding to any one of the interfaces is hot-plug error information according to the device data collected by the BMC, the error status of the PCIE CE error information corresponding to any one of the interfaces, and the preset hot-plug error parameter; When the PCIE CE error information is hot-plug error information, the OS sends shielding information corresponding to the hot-plug error information to the BMC.

2. The method according to claim 1, characterized in that When the PCIE CE error information is hot plug error information, after the OS sends shielding information corresponding to the hot plug error information to the BMC, the method further includes: The BMC deletes the hot plug error information in response to the shielding information corresponding to the hot plug error information.

3. The method according to claim 1, characterized in that The method further comprises: The BMC collects device data at preset intervals and sends the data to the OS.

4. The method according to claim 1, characterized in that: The hot-plug error parameter includes a hot-plug error interval, and the OS determines whether the PCIE CE error information corresponding to any one of the interfaces is hot-plug error information according to the device data collected by the BMC, the error condition of the PCIE CE error information corresponding to any one of the interfaces, and the preset hot-plug error parameter, including: For PCIE CE error information corresponding to any interface, the OS determines the error duration between the first error time and the last error time of the PCIE CE error information; When the error duration is less than the hot plug error interval, and the OS determines, based on the device data collected by the BMC, that there is device change information for the interface within the error duration, the PCIE CE class error information corresponding to the interface is determined as hot plug error information, wherein the device change information includes adding a device or reducing a device.

5. The method according to claim 1, characterized in that The hot-plug error parameter includes a hot-plug error threshold, and the OS determines whether the PCIE CE error information corresponding to any one of the interfaces is hot-plug error information according to the device data collected by the BMC, the error condition of the PCIE CE error information corresponding to any one of the interfaces, and the preset hot-plug error parameter, including: For PCIE CE error information corresponding to any interface, the OS counts the number of errors of the PCIE CE error information within a specified time period; If the number of errors of the PCIE CE type error information within the calibration time period is less than the hot plug error threshold, and the OS determines that there is device change information of the interface within the calibration time period based on the device data collected by the BMC, the PCIE CE type error information corresponding to the interface is determined as hot plug error information, wherein the device change information includes adding a device or reducing a device.

6. The method according to claim 1, characterized in that The method further comprises: During the startup of the ARM platform, a data interface service is registered on the OS and the BMC through the BIOS of the ARM platform, and the data interface service is used to provide a data transmission interface between the BMC and the OS through the BIOS.

7. The method according to claim 6, characterized in that During the booting process of the ARM platform, registering a data interface service on the OS and the BMC through the BIOS of the ARM platform includes: After the server of the ARM platform is powered on, the BIOS registers a data interface service on the BMC; After entering the OS, running a first hot-plug error management program in the OS, wherein the first hot-plug error management program accesses the data interface service registered by the BIOS; The OS sends activation information to the BMC through a data transmission interface provided by the data interface service; The BMC starts a second hot-plug error management program stored in the BMC in response to the activation information.

8. The method according to claim 7, characterized in that After the server of the ARM platform is powered on, the method further includes: The BIOS sends the hot-plug error reporting parameters currently stored in the BIOS to the BMC.

9. The method according to claim 7, characterized in that: After entering the OS and running the first hot-plug error management program in the OS, the method further includes: The first hot-plug error management program accesses and obtains the hot-plug error parameters currently stored in the BIOS.

10. The method according to claim 1, characterized in that The method further comprises: After the server of the ARM platform is powered on, the BMC detects whether it stores the hot-plug error reporting parameters obtained from the BIOS; The BIOS sends the hot-plug error parameters currently stored in the BIOS to the BMC, and registers a data interface service, where the data interface service is used to provide a data transmission interface between the BMC and the OS through the BIOS; After entering the OS, a first hot-plug error management program in the OS is run, wherein the first hot-plug error management program accesses the data interface service registered by the BIOS and obtains the hot-plug error information currently stored in the BIOS. parameter; The OS sends activation information to the BMC through a data transmission interface provided by the data interface service; The BMC starts a second hot-plug error management program stored in the BMC in response to the activation information.

11. The method according to claim 10, characterized in that Before entering the OS, the method further includes: The BIOS monitors the modification operation of the hot-plug error reporting parameter in real time; When there is a modification operation of the hot-plug error reporting parameter, the BIOS sends the modified hot-plug error reporting parameter to the OS and the BMC respectively.

12. The method according to any one of claims 7 to 11, characterized in that: Before entering the OS, the method further includes: In response to a closing operation on a hot-plug error reporting setting option, the BIOS sends hot-plug reporting closing information to the OS and the BMC respectively; The OS stops executing the first hot plug error management program in response to the hot plug error shutdown information; The BMC stops executing the second hot-plug error management program in response to the hot-plug error shutdown information.

13. The method according to claim 1, characterized in that The method further comprises: During the operation of the ARM platform, the OS responds to the change operation of the hot-plug error reporting parameter, stores the changed hot-plug error reporting parameter, and sends the changed hot-plug error reporting parameter to the BIOS and the BMC respectively.

14. The method according to claim 13, characterized in that During the next ARM platform startup process, the method further includes: After the server of the ARM platform is powered on, the BIOS sends the modified hot-plug error reporting parameters stored in the BIOS to the BMC and the OS respectively.

15. The method according to claim 12, characterized in that The method further comprises: In response to the upgrade operation of the management program, the OS obtains the first hot-plug error reporting management program to be upgraded, and upgrades the current first hot-plug error reporting management program to the first hot-plug error reporting management program to be upgraded.

16. The method according to claim 1, characterized in that After the OS determines whether the PCIE CE error information is hot plug error information according to the device data collected by the BMC, the number of error reports of the PCIE CE error information, and a preset hot plug error parameter, the method further includes: When the PCIE CE error message is not hot-plug error message, the OS sends normal processing information corresponding to the PCIE CE error message to the BMC; The BMC records and reports the PCIE CE type error information in response to normal processing information corresponding to the PCIE CE type error information.

17. The method according to claim 1, characterized in that When the PCIE CE error message is hot plug error message, the method further includes: The OS stores the PCIE CE error information and the timestamp of the PCIE CE error information in a hot plug summary list, so as to summarize and view all hot plug error information.

18. The method according to claim 17, characterized in that The method further comprises: In response to the query operation of the hot-plug error, the OS generates a visualized hot-plug error chart according to the hot-plug summary list for display.

19. A processor architecture, characterized in that: The processor architecture includes BMC, OS and BIOS, and the processor architecture is used to execute the hot plug error reporting method described in any one of claims 1-18.

20. A computer device, characterized in that: include: At least one processor, and a memory, wherein the memory stores a computer program that can be run on the processor, wherein the processor executes the hot plug error reporting method as described in any one of claims 1-18 when executing the computer program.

21. A non-volatile readable storage medium, characterized in that: The non-volatile readable storage medium stores a computer program, wherein the computer program, when executed by a processor, executes the hot-plug error reporting method according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Interruption processing method and system, electronic equipment and storage medium

    CN109885521A

  • Memory repairable error reporting method and device, equipment and medium

    CN115033409A

  • Memory CE error reporting method, system and device of ARM architecture server and medium

    CN115391081A

  • Memory fault-tolerant method and device capable of correcting error storm and medium

    CN116954986A

  • Hot plug error reporting method, processor architecture, equipment and storage medium

    CN117389819A