Error reporting method of processor platform and electronic equipment

By obtaining the error identification of the processor platform and building a prediction model, the problem of difficulty in providing early warning of potential failures in existing technologies is solved, and timely detection and prevention of processor platform failures are achieved, thereby improving the system's availability and operation and maintenance efficiency.

CN120803802AActive Publication Date: 2025-10-17LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Patent Information

Application Number
CN202511311814.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

The existing BIOS RAS error reporting mechanism mainly focuses on recording and reporting errors that have occurred, but has limited in-depth analysis of historical error data and the ability to predict potential failures, which can easily miss opportunities to take preventive measures before failures occur.

Method used

By obtaining multiple error identifiers from the processor platform, a prediction model is constructed, and historical data and real-time operating parameters are used to predict future fault types and report them to the processor platform in a timely manner for early warning.

Benefits of technology

It improves the predictive maintenance capabilities of the processor platform, detects impending failures in a timely manner, reduces the risk of unplanned downtime, and improves system availability and operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803802A_ABST
    Figure CN120803802A_ABST
Patent Text Reader

Abstract

The invention discloses an error reporting method of a processor platform and electronic equipment, and relates to the technical field of computer hardware, and the error reporting method comprises the steps that a plurality of error identifiers of the processor platform are acquired, and the error identifiers are used for representing information of fault types of the processor platform; the prediction model is constructed according to the historical data and the real-time operation parameters, the prediction model can improve the predictive maintenance capability of the processor platform, and then the fault type generated by the processor platform in a period of time in the future is determined according to the error identifier and the prediction model. The technical problem that a processor platform is difficult to warn potential faults in advance in the prior art can be solved, and the technical effects that the fault which is about to occur on the processor platform can be detected in time, then the fault is processed in time according to the prediction result, and the chance of taking preventive measures before the fault occurs is prevented from being missed easily are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer hardware, and particularly relates to a processor platform error reporting method and an electronic device. BACKGROUND

[0002] In modern data centers and high-performance computing environments, the reliability, availability and serviceability (RAS) of a system is crucial to ensure business continuity and reduce operation and maintenance costs. At present, in a server platform, a BIOS (Basic Input Output System) is a core component for hardware initialization and management, and a RAS error reporting mechanism thereof mainly relies on standard interfaces such as APEI (ACPI Platform Error Interface) to record and report error information. However, the existing RAS mechanism of the BIOS mainly focuses on the recording and reporting of errors that have occurred, and has limited ability to analyze historical error data in depth and predict potential faults, which is prone to miss the opportunity to take preventive measures before a fault occurs. SUMMARY

[0003] The present application provides a processor platform error reporting method and an electronic device to at least solve the technical problem that a processor platform is difficult to predict potential faults in advance in the related art.

[0004] The present application provides a processor platform error reporting method, comprising: acquiring a plurality of error identifiers of a processor platform, the error identifier being used to represent information of a fault type of the processor platform; constructing a prediction model according to historical data and real-time running parameters, the historical data and the real-time running parameters at least comprising: error logs of a basic input output system of the processor platform, hardware running parameters and environmental monitoring data; determining a fault type generated by the processor platform in a future period of time according to the plurality of error identifiers and the prediction model; and reporting the fault type to the processor platform and giving a warning.

[0005] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any one of the processor platform error reporting methods.

[0006] The error reporting method provided by the application obtains a plurality of error identifiers of a processor platform, which cover the types and characteristics of errors of the processor platform, and provides basic data for subsequent error analysis and prediction. A prediction model is constructed according to historical data and real-time running parameters, which can improve the predictive maintenance capability of the processor platform. The error identifiers and the prediction model are used to determine the fault types generated by the processor platform in a future period of time, so that the technical problem that the processor platform is difficult to early warn potential faults in the related art can be solved, the technical effect that the processor platform can be detected in time to determine that a fault will occur soon can be achieved, and timely processing can be performed according to the prediction result, so that the opportunity to take preventive measures before the fault occurs can be avoided. BRIEF DESCRIPTION OF DRAWINGS

[0007] In order to more clearly illustrate the embodiments of the application, the drawings required in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0008] Figure 1 A hardware structure block diagram of a mobile terminal is shown, which provides an error reporting method of a processor platform according to an embodiment of the application.

[0009] Figure 2 A flowchart of an error reporting method of a processor platform provided by an embodiment of the application is shown.

[0010] Figure 3 A flowchart of another error reporting method of a processor platform provided by an embodiment of the application is shown.

[0011] The above drawings include the following reference signs:

[0012] 102, processor; 104, memory; 106, transmission device; 108, input and output device. DETAILED DESCRIPTION

[0013] The technical solutions in the embodiments of the application will be described clearly and completely in combination with the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0014] It should be noted that in the description of the present application, the term "comprising", "containing" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or inherent to such a process, method, article or apparatus. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0015] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments.

[0016] In combination with the specific application environment architecture or specific hardware architecture on which the error reporting method of the processor platform is dependent, the specific application environment architecture or specific hardware architecture is described here.

[0017] The method embodiments provided in the embodiments of the present application can be executed in a server device or similar computing device. Taking the case of running on a server device, Figure 1 is a hardware structure block diagram of a server device of the error reporting method of the processor platform according to the embodiments of the present application. As shown in Figure 1 , the server device can include one or more (only one is shown in Figure 1 ) processors 102 (the processor 102 can include but is not limited to a central processing unit CPU, a microprocessor MCU or a programmable logic device FPGA processing device) and a memory 104 for storing data, wherein the above-mentioned server device can further include a transmission device 106 for communication function and an input and output device 108. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned server device. For example, the server device can further include more or less components than those shown in Figure 1 , or have a different configuration from Figure 1 .

[0018] The error reporting method of the present application is mainly applied in environments with strict requirements on RAS (Reliability Availability Serviceability) such as data center servers, high-performance computing systems and embedded computing platforms. Specifically, the application environment architecture on which it depends is as follows:

[0019] Distributed computing architecture: In a large-scale server cluster, the present method can cooperate with BIOS and BMC on multiple nodes to realize cross-node fault detection and reporting, and improve the stability and operation and maintenance efficiency of the entire cluster.

[0020] Virtualized environment: This method is applicable to virtualized data centers and can penetrate the virtual layer to provide real-time monitoring and early warning of physical hardware errors, ensuring the continuous operation of virtual machines and services.

[0021] The technical solution of this application is particularly targeted at the following hardware architectures:

[0022] BMC with integrated VGA controller: As a hardware management controller, the integration of a VGA controller is key to efficient data exchange. The BMC's VGA controller provides independent video memory, allowing the BIOS to directly access and write out error data, significantly reducing data transmission latency.

[0023] TLS and AES-256 encryption hardware: To ensure data security and compliance, this approach relies on hardware that supports TLS (Transport Layer Security) and AES-256 (Advanced Encryption Standard). TLS ensures data encryption during network transmission, while AES-256 encrypts data at rest, providing a dual line of defense for data protection.

[0024] AGESA (Generic Encapsulated Software Architecture) microcode-compatible hardware: The hardware should support the platform server's AGESA microcode, which is a core component of the processor's low-level initialization and configuration. AGESA provides a software foundation that is tightly integrated with the processor's error detection and reporting mechanisms.

[0025] Hardware interface configuration: In this application, the BMC in the platform server integrates a VGA controller and has an independent video memory address space. Through the BIOS hardware abstraction layer (HAL) configuration, the BMC video memory address is mapped to the system memory address space, allowing the BIOS to directly access it.

[0026] Some of the technical terms involved are as follows:

[0027] Reliability, Availability, and Serviceability (RAS) is used to describe the ability of a system to maintain normal operation during its runtime. RAS, which consists of three concepts, is the basis for evaluating and designing high-performance, mission-critical systems. RAS emphasizes the ability of a system to run without failure for a predetermined period of time (reliability), to continue running or quickly recover when a failure occurs (availability), and to maintain and upgrade without the need to shut down the system (maintainability). For data centers, cloud computing environments, and enterprise applications, RAS is crucial to ensure business continuity and data integrity. Through the comprehensive optimization of hardware design, firmware functions, management software, and service strategies, RAS aims to maximize system uptime, minimize the impact of failures, support rapid repair and upgrade, and reduce planned and unplanned system downtime.

[0028] Basic Input / Output System (BIOS) is the underlying firmware that runs when a computer starts up, responsible for tasks such as hardware self-check, initialization of peripheral devices, loading of operating system boot programs, etc. BIOS serves as a bridge between computer hardware and operating systems, ensuring that hardware devices can be correctly identified and used by the operating system. During the startup process, BIOS checks hardware configuration, sets system parameters, and is crucial to ensuring the stable operation and efficient startup of the computer.

[0029] Baseboard Management Controller (BMC) is a management chip independent of the main processor, responsible for remote monitoring and management of the system. BMC is a microprocessor embedded on the server motherboard, independent of the main CPU operation, mainly used for remote monitoring and management of system operation status. BMC collects information through various sensors and hardware interfaces, such as temperature, voltage, fan speed, etc., and also supports remote control of servers through the network. BMC is a core component of server RAS (reliability, availability, serviceability) capabilities. It can monitor system health in real time, perform functions such as remote power on / off, fault warning, log recording, etc., thereby improving system availability and maintainability, and reducing the monitoring burden of administrators.

[0030] The Machine Check Architecture (MCA) mechanism is a key processor mechanism used to detect and report hardware errors. It is primarily used to identify and record hardware failures that occur during runtime. MCA is designed to collect as detailed error information as possible when a hardware failure occurs, enabling the operating system or BIOS to take appropriate corrective actions or perform fault analysis, thereby improving system reliability, availability, and maintainability (RAS).

[0031] The Boot Error Record Table (BERT) is a table in the BIOS that records error information that occurs during the boot process. The BERT is a data structure within the BIOS specifically used to record error information that occurs during system startup. It acts as a system log, particularly during system initialization, and can capture many hardware or firmware-level errors. The BERT helps system administrators and engineers quickly diagnose boot problems and provide early indications of hardware failures, facilitating maintenance and debugging. It typically contains information such as the error code, error type, time of occurrence, and location, making it crucial for troubleshooting.

[0032] Transport Layer Security (TLS) is used to provide data encryption and authentication in network communications.

[0033] An embodiment of the present application provides an error reporting method for a processor platform, and the method is described in detail in conjunction with the execution flow of the error reporting method for the processor platform.

[0034] According to a processor platform error reporting method provided by this application, such as Figure 2 As shown, including:

[0035] Step S1, obtaining multiple error identifiers of a processor platform, where the error identifiers are used to characterize information about a fault type of the processor platform;

[0036] Specifically, during the running of the processor platform, error identifications from the MCA mechanism are continuously monitored and collected, which not only contain the types of errors, but also contain detailed information such as occurrence time, CPU number, memory channel number and error address through MCA record registers, providing basic data for subsequent error classification and fault prediction. The error identification is obtained from the hardware error information and the type of error information. The hardware error information includes but is not limited to memory error, CPU error, bus error, etc. Through the BERT table record, the error information at the start of the processor platform can be obtained. Through real-time and comprehensive error identification collection, the health status of the processor platform can be grasped in time, and detailed data support is provided for subsequent intelligent analysis and early warning. In this way, the response speed of the processor platform and the accuracy of fault detection can be improved.

[0037] Step S2, according to the historical data and real-time running parameters, a prediction model is constructed, and the historical data and real-time running parameters at least include: error logs of the basic input and output system of the processor platform, hardware running parameters and environmental monitoring data;

[0038] Specifically, when constructing the prediction model, first, the error logs recorded by the BIOS are analyzed in depth, and the error logs include but are not limited to the BERT table record and the MCA error log. According to the error logs, the types, frequencies, and influence ranges of errors are extracted. At the same time, the hardware running parameters are continuously monitored and provided, including but not limited to: CPU temperature, memory voltage, fan speed and other real-time running parameters, and temperature, humidity, power state and other environmental monitoring data of the machine room. These multi-source data are integrated and input into the machine learning model, and the model can identify potential fault patterns and warning signals after training. Through the prediction model combining historical data and real-time parameters, the future possible hardware failure can be accurately predicted based on the current hardware state and running environment, such as issuing a warning 36 / 72 hours in advance. In this way, the risk of unplanned downtime can be reduced, and the availability and operation and maintenance efficiency of the processor platform are improved.

[0039] Step S3, according to a plurality of error identifications and a prediction model, determining a fault type generated by the processor platform in a future period of time;

[0040] Specifically, when a new hardware error occurs, a new error identifier is detected and fed into a trained predictive model. The model analyzes this data, identifies potential failure modes, and outputs a prediction of the type of failure expected over the next period of time (e.g., the next week), including the specific type of failure and its estimated probability of occurrence. The predictive model can also simultaneously input hardware operating parameters of the current processor platform. This allows for a comprehensive evaluation based on the processor platform's current operating status and predicted failures, resulting in more accurate maintenance measures. By combining real-time error data with the predictive model, maintenance strategies can be dynamically adjusted, prioritizing high-risk failures, reducing response time, and improving the targetedness and efficiency of maintenance.

[0041] Step S4: Report the fault type to the processor platform and issue an early warning.

[0042] Specifically, once the prediction model identifies a possible future fault type, the BIOS immediately reports this fault type information to the BMC via an optimized data exchange mechanism (based on efficient transmission of VGA controller memory). The BMC then sends the warning information to the server management system or notifies maintenance personnel to take preventive measures, such as automatic restart and repair and memory chip isolation.

[0043] Through the aforementioned processor platform error reporting method, multiple error identifiers are obtained for the processor platform. These identifiers cover the types and characteristics of the processor platform's errors, providing basic data for subsequent error analysis and prediction. Based on historical data and real-time operating parameters, a prediction model is constructed. This prediction model can enhance the processor platform's predictive maintenance capabilities. The error identifiers and prediction model are then used to determine the types of failures that the processor platform may experience in the future. This resolves the technical issue of processor platforms being unable to provide early warnings of potential failures in related technologies, achieving the technical effect of timely detecting impending processor platform failures and then promptly addressing them based on the prediction results, avoiding the opportunity to take preventive measures before a failure occurs.

[0044] In some optional implementations, the method further includes: initializing the basic input / output system; obtaining read access to the video memory space, and obtaining the address and capacity of the video memory space based on the read access; and allocating a buffer with a preset capacity within the video memory space. By completing video memory space configuration during the BIOS initialization phase, the efficiency of error data reporting is improved, and the data exchange time between the BIOS and the BMC is reduced from the traditional 300ms to under 50ms, a reduction of 83%. Furthermore, the system startup time is shortened by an average of 15%, reducing unnecessary delays.

[0045] First, get the address mapping, in the BIOS initialization phase, read the VGA controller information of the BMC through the PCI configuration space, get its video memory base address and size;

[0046] Set permissions, set the processor platform memory mapping register (MMR), and give the BIOS read and write permissions for this video memory area;

[0047] Perform buffer allocation, divide a fixed-size ring buffer (such as 16KB) in the BMC video memory for storing error data. The buffer structure is as follows:

[0048] typedef struct {

[0049] uint32_t head; / / write pointer

[0050] uint32_t tail; / / read pointer

[0051] uint8_t data

[16384] ; / / data storage area

[0052] uint8_t flags; / / status flag bit

[0053] } BiosBmcSharedBuffer。

[0054] PCI (Peripheral Component Interconnect) is a high-speed bus standard for connecting computer hardware. Originally developed by Intel, it was designed to provide faster data transfer speeds and better performance than the previous ISA bus. The advent of the PCI bus significantly promoted standardization and compatibility in computer hardware, allowing peripherals (such as graphics cards, sound cards, and network cards) from different manufacturers to communicate with the computer motherboard through a unified interface. Each PCI device has a configuration space, a memory area used to store and retrieve device configuration information. This configuration space contains information such as the device's identification, functions, and status registers. This information is crucial for the operating system or BIOS to initialize and configure the hardware. During computer startup, the BIOS or operating system accesses the PCI configuration space to identify and configure PCI devices. For example, reading information about the VGA controller from the BMC and obtaining its video memory base address and size are both accomplished by accessing specific registers in the PCI configuration space. Access to the PCI configuration space is typically accomplished through the PCI configuration registers, which allow the host to read and write configuration information for PCI devices. The access process involves writing the device's bus number, device number, function number, and register offset address to the PCI Configuration Address Register, and then reading the PCI Configuration Data Register to obtain configuration data or writing new configuration values ​​to it to update the device status.

[0055] In some optional implementations, obtaining multiple error identifiers of the processor platform includes: obtaining multiple error information of the processor platform; and constructing an error structure according to the multiple error information and a preset order.

[0056] Specifically, error detection is first performed. The MCA mechanism of the control processor detects a hardware error and notifies the BIOS through an interrupt. Data packaging is then performed. The control BIOS reads error information from the MCA register and constructs the error information into an error structure according to a certain preset order. The error structure structure is as follows:

[0057] typedef struct {

[0058] uint8_t error_type; / / MCA error type

[0059] uint8_t severity; / / severity level

[0060] uint32_t cpu_id; / / CPU number

[0061] uint32_t bank_id; / / memory channel number

[0062] uint64_t error_address; / / error address

[0063] uint32_t mca_record

[16] ; / / MCA record register value

[0064] uint64_t timestamp; / / timestamp

[0065] uint32_t crc32; / / CRC check code

[0066] } McaIrqData.

[0067] In the case of detecting that there is new error information in the basic input and output system input, the error information is converted into an error identifier according to an error structure body and a preset format. By constructing a structured error identifier, the data interaction between the BIOS and the BMC is simplified, the complexity of data transmission is reduced, the reliability and integrity of transmission are ensured, the real-time performance of error reporting is reduced from an average of 10 seconds to within 1 second, and the performance is improved by 90%.

[0068] In some optional embodiments, the error information is converted into an error identifier according to an error structure body, including:

[0069] The write pointer of the buffer of the processor platform is obtained; after the write pointer is obtained, the information can be allowed to be stored in the data storage area.

[0070] According to the write pointer, the error structure is written into the data storage area of the buffer area, and the error data check value CRC in the data storage area is updated according to the error structure; the VGA memory space of the BMC is used as an information temporary storage medium between the BIOS and the BMC. Specifically, after detecting a hardware error (such as an MCA error, a memory error, etc.), the BIOS packs the error data according to a preset structure format, calculates a CRC check code, and then writes the error data into a specified region of the VGA memory of the BMC. When subsequent reading is required, the error data can be accessed in parallel through a dedicated hardware interface and the VGA memory, replacing the traditional IPMI serial communication mode, improving the data interaction efficiency, and reducing the single interaction time of the BIOS and the BMC from the traditional IPMI of 300 ms to below 50 ms, about 83%. Due to the reduction of the interaction delay between the BIOS and the BMC, the platform startup time is shortened by an average of 15%, from 45 seconds to 38 seconds. To ensure compatibility, the platform also retains IPMI communication as a fallback solution, which automatically switches to the IPMI mode when the VGA memory interaction is abnormal.

[0071] According to a preset period, a plurality of ready flag bits of the buffer area are obtained, it is judged whether the ready flag bit is a target ready flag bit (READY_FLAG=1), and in the case that the ready flag bit is the target ready flag bit, the updated error data check value in the data storage area is read; according to the updated error data check value, it is determined whether the data in the error structure is normal; the above-mentioned CRC is the unique check value of the error data packet, and the BIOS sends the data packet to the BMC, and the CRC check value is also sent, which is used by the BMC as a check credential to ensure that the data is not damaged.

[0072] In the case that it is determined that the data in the error structure is normal, the normal error structure is parsed according to a preset format to obtain an error identifier. The preset format is the structure of the error structure. Through the setting of the ready flag bit and the verification of the CRC, the reliability of error reporting is significantly enhanced, and even in a high-load environment, the fault detection accuracy of the system can be maintained at more than 90%, thereby enhancing the overall stability of the system.

[0073] The above-mentioned MCA mechanism monitors various possible hardware fault points at the processor core level, including but not limited to a memory system (such as a memory controller, a data path and a cache), an I / O subsystem, a system bus, a power management unit and a clock unit, etc. When an error is detected, the MCA triggers an interrupt (Machine Check Interrupt, referred to as MCI), and saves the error information in a group of special registers called Machine Check Registers (MCR).

[0074] The MCA mechanism can report multiple types of errors, which are classified into two categories: Uncorrectable Error (UCE) and Correctable Error (CE).

[0075] Uncorrectable Error (UCE): This type of error usually means that the hardware component is permanently damaged, and the system needs to respond immediately, such as shutting down the affected CPU core or restarting the entire system.

[0076] Correctable Error (CE): In contrast, CE errors are usually transient and can be repaired by the hardware's own error correction capabilities, such as error correction in ECC memory. However, frequent CE errors may also indicate that the hardware is about to experience a UCE error. The error information in this application can be for both types of errors. According to the detected error type, targeted processing is performed.

[0077] In some optional embodiments, the method further comprises:

[0078] After obtaining the error identifier, output an acknowledgement flag (ACK_FLAG = 1) to the basic input / output system;

[0079] In the case where the basic input / output system receives the acknowledgement flag, clear the ready flag and the determined flag of the buffer area of the processor platform.

[0080] Specifically, according to the error structure, the BMC parses the error data and remaps it to the internal buffer memory of the BMC to obtain the error identifier, outputs the acknowledgement flag to the basic input / output system, and obtains the fault type according to the error identifier and the prediction model and then reports the data. By introducing the acknowledgement flag, bidirectional communication confirmation between the BIOS and the BMC is achieved, the timely clearing mechanism of the flag avoids redundant storage of data, improves the utilization rate of the buffer area, and reduces the possibility of data conflict.

[0081] In some optional embodiments, the method further comprises: before outputting the acknowledgement flag to the basic input / output system, in the case where the ready flag is a target ready flag, obtaining the acknowledgement flag according to a preset period, judging whether the acknowledgement flag is a target preset flag, and in the case where the acknowledgement flag is the target preset flag, clearing the ready flag and the determined flag of the buffer area of the processor platform.

[0082] Specifically, after the BIOS writes the ready flag bit, the confirmation flag bit is polled, and after the confirmation flag bit is obtained, the ready flag bit and the confirmation flag bit are cleared to release the buffer space and prepare for new error data, indicating that the interaction is completed. In this way, the dynamic use of the buffer and the fast response of the platform can be guaranteed, the invalid waiting time is reduced, and the real-time performance of the platform is improved.

[0083] In some optional embodiments, the method further comprises:

[0084] Obtaining a historical error log, and performing cleaning and feature extraction on data in the historical error log to obtain a target historical error log, the feature-extracted data at least including: error type and occurrence frequency; the historical error log including but not limited to BERT table records and MCA error logs. The feature-extracted data can also include the impact range, such as single-DIMM failure or cross-CPU node failure.

[0085] Training a machine learning model according to the target historical error log and the error category to obtain an error classification model; the machine learning model including but not limited to random forest or LSTM algorithm;

[0086] Classifying the error identifier according to the error classification model.

[0087] Specifically, when a new error occurs (the occurrence of the error is detected, and the MCA mechanism generates an interrupt to notify the BIOS), the error data is input into the trained model, and the error category is output. Through deep analysis of historical data and model training driven by machine learning, the newly occurring error can be automatically and accurately classified, improving the accuracy and efficiency of error identification. Intelligent error classification reduces the time for administrators to locate key problems by 60%, from an average of 30 minutes to 12 minutes.

[0088] In some optional embodiments, the method further comprises:

[0089] Determining the priority of the plurality of error identifiers according to target features of the plurality of error identifiers, the target features including: the impact range of the error identifier and the load state of the processor platform.

[0090] According to the influence range of the error (such as single DIMM failure or cross-CPU node failure), the current load state of the platform (such as CPU occupancy, memory utilization), and other factors, the error priority is dynamically adjusted. For example, if the current platform load is high, even a slight hardware error can quickly deteriorate into a serious failure, at which time the model will automatically elevate the priority of such errors to ensure timely response. By dynamically adjusting the error priority, those errors that can have a significant impact on overall performance can be prioritized, improving the timeliness of fault response and the level of maintenance of system stability. Especially in high-load scenarios, this mechanism can effectively prevent system crashes caused by neglecting initial errors, further enhancing the availability and reliability of the platform. For example, when the system CPU occupancy reaches 85%, a slight error (such as a single DIMM correctable error) will be elevated to high priority for processing, thereby avoiding potential fault escalation and ensuring stable operation of the system under high pressure.

[0091] In some optional embodiments, a prediction model is constructed based on historical data and real-time running parameters, including:

[0092] The historical data and real-time running parameters include BIOS error logs, hardware running parameters, microcode data of the processor platform, and environmental monitoring data. The BIOS error logs include but are not limited to BERT table records, MCA error logs, and memory self-check results. The hardware running parameters include but are not limited to real-time data such as CPU temperature, memory voltage, and fan speed. The AGESA microcode data includes but is not limited to processor underlying error information such as cache errors and bus errors. The environmental monitoring data includes but is not limited to room temperature, humidity, and power supply status.

[0093] The machine learning model is trained based on the historical data and real-time running parameters to obtain a prediction model. The machine learning model includes at least one of the following: random forest or long short-term memory network. The prediction model obtained by the above method has an early warning accuracy rate of memory failure of over 90% and can issue a warning 72 hours in advance. Preventive maintenance measures can be taken in a timely manner based on the prediction results of the prediction model, unplanned downtime is reduced by 50%, and MTTR (mean time to repair) is reduced from 4 hours to 2.4 hours. The intelligent error classification of the prediction model reduces the time for administrators to locate critical issues by 60%, from an average of 30 minutes to 12 minutes.

[0094] The processor platform of the present application supports AGESA 1.2.0 and above version microcode in hardware; supports the display memory interaction interface of mainstream BMC firmware (such as AST2500, iKVM) in software, while retaining the IPMI 2.0 compatible mode. In terms of data security, TLS encryption transmission and AES-256 storage encryption are used to ensure the confidentiality of error data, and blockchain evidence is used to ensure data tamper resistance; in terms of compliance support, it meets the compliance requirements of ISO 27001, PCI-DSS, etc., and the auditability of error reporting reaches enterprise-level standards; in terms of audit tracking, through blockchain evidence, the generation, transmission and storage of error reports can be traced back to meet the compliance audit requirements.

[0095] In order to optimize error reporting, a multi-level error reporting strategy is introduced. When the BIOS detects a hardware error, a preliminary analysis is first performed to determine whether it needs to be reported to the BMC immediately or whether it can be handled at the BIOS level first. For example, for minor, correctable errors, the BIOS can attempt to repair them locally using mechanisms such as PPR (Power-on Program Package Repair); for serious, uncorrectable errors, the BIOS will immediately trigger the reporting process to ensure that the BMC and upper management platform can respond in a timely manner.

[0096] Users are allowed to customize the trigger conditions and strategies of error warning according to their business needs and platform running environment. Users can set specific error types and parameter thresholds (such as CPU temperature exceeding 80°C, memory error frequency exceeding 10 times / hour, etc.) through the management interface of BIOS and BMC, and when the platform detects errors that meet these conditions, the warning process will be automatically started and appropriate preventive measures will be taken. The user interface is provided for custom settings, enhancing the configurability and user interaction experience of the system, so that non-professionals can easily manage the RAS features of the system.

[0097] In order to enable those skilled in the art to more clearly understand the technical solutions of the present application, the implementation process of the error reporting method of the processor platform of the present application will be described in detail below in conjunction with specific embodiments.

[0098] Embodiment

[0099] As shown in Figure 3 , the control BIOS detects the hardware errors of the platform, generates error data and packages the detected error information, and calculates a new CRC check code: obtain a plurality of error information of the processor platform; construct an error structure according to the plurality of error information and a preset order.

[0100] The control BIOS writes the error structure into the VGA display memory, sets the ready flag of the buffer, detects that the set buffer ready flag is 1, controls the MBC to monitor the ready flag, reads the display memory data, verifies the CRC check code, and after verification, processes the error data (analyzes the error identifier from the error structure): obtains the write pointer of the buffer of the processor platform; writes the error structure into the data storage area of the buffer according to the write pointer, and updates the error data check value in the data storage area according to the error structure, and the data storage area is used for temporarily storing the error structure; according to the preset period, obtain a plurality of ready flags of the buffer, and judge whether the ready flag is the target ready flag, in the case where the ready flag is the target ready flag, read the updated error data check value in the data storage area; according to the updated error data check value, determine whether the data in the error structure is normal; in the case where it is determined that the data in the error structure is normal, according to the preset format, the normal error structure is analyzed to obtain the error identifier, and after obtaining the error identifier, the confirmation flag is set to 1.

[0101] In the process of obtaining the error identifier, the confirmation flag is polled to determine whether it is 1, and in the case where the confirmation flag is 1 and the ready flag is 1, the ready flag (0) and the confirmation flag (0) are cleared: in the case where the ready flag is 1, the confirmation flag is obtained according to the preset period, and it is judged whether the confirmation flag is the target preset flag, and in the case where the confirmation flag is the target preset flag, the ready flag and the determination flag of the buffer of the processor platform are cleared.

[0102] Meanwhile, according to the obtained error identifier and the prediction model, the fault type that the processor platform is likely to produce in the future period of time is determined; the fault type is reported to the processor platform and a warning is given.

[0103] Through the description of the above implementation, those skilled in the art can clearly understand that the method according to the above embodiment can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better implementation.

[0104] The embodiment of the present application further provides an error reporting device of a processor platform. The device comprises: an acquisition module, configured to acquire a plurality of error identifiers of the processor platform, the error identifier being used to represent information of a fault type of the processor platform; a construction module, configured to construct a prediction model according to historical data and real-time running parameters, the historical data and the real-time running parameters at least comprising: an error log of a basic input and output system of the processor platform, a hardware running parameter and environmental monitoring data; a determination module, configured to determine a fault type of the processor platform generated in a future period of time according to the plurality of error identifiers and the prediction model; and a reporting module, configured to report the fault type to the processor platform and perform early warning.

[0105] The features of the embodiment of the error reporting device of the processor platform can be referred to the related description of the embodiment of the error reporting method of the processor platform, which will not be repeated here.

[0106] The embodiment of the present application further provides an electronic device, comprising a memory and a processor, the memory storing a computer program, and the processor being configured to run the computer program to perform the steps in any of the above-mentioned error reporting method embodiments of the processor platform.

[0107] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above-mentioned error reporting method embodiments of the processor platform when running.

[0108] In an example embodiment, the above-mentioned computer readable storage medium can include, but is not limited to: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0109] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned error reporting method embodiments of the processor platform.

[0110] The embodiment of the present application further provides another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above-mentioned error reporting method embodiments of the processor platform.

[0111] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide an enabling description of the applications and are not intended to limit the scope of the applications. Therefore, embodiments or examples described herein are not meant to be limiting, but merely to aid in the understanding of the overall more complete disclosure of the applications. Accordingly, those skilled in the art will recognize that modifications and variations of the more complete description herein can be resorted to without departing from the spirit and scope of the applications. Therefore, it is intended that the applications encompass all such modifications and variations as fall within the scope of the applications. All articles, patents, and other publications that have been cited herein are incorporated herein by reference for the teachings relevant to the sentence and / or paragraph in which the article, patent, and / or publication is mentioned.

[0112] The error reporting method of a processor platform and the electronic device provided by the application are described in detail above. The principles and implementation manners of the application are described by applying specific examples in this paper. The above description of the embodiments is only applicable to help understand the method of the application and its core idea. It should be pointed out that for ordinary skilled in the art, without departing from the principles of the application, the application can be improved and modified in several ways. These improvements and modifications also fall within the protection scope of the claims of the application.

Claims

1. A method for reporting errors on a processor platform, characterized in that: include: Acquire multiple error identifiers of a processor platform, where the error identifiers are used to characterize information about a fault type of the processor platform; Constructing a prediction model based on historical data and real-time operating parameters, wherein the historical data and real-time operating parameters include at least: error logs of a basic input / output system of the processor platform, hardware operating parameters, and environmental monitoring data; determining, based on the plurality of error identifiers and the prediction model, a type of fault that will occur on the processor platform within a future period of time; The fault type is reported to the processor platform and an early warning is issued.

2. The error reporting method according to claim 1, characterized in that: The obtaining of multiple error identifiers of the processor platform includes: Obtaining multiple error information of the processor platform; Constructing an error structure according to the plurality of error messages and a preset order; When it is detected that the basic input and output system has new error information input, the error information is converted into an error identifier according to an error structure and a preset format.

3. The error reporting method according to claim 2, characterized in that: The method further comprises: Initializing the basic input and output system; Obtaining a video memory space read permission, and obtaining the address and capacity of the video memory space according to the read permission; A buffer zone with a preset capacity is divided in the video memory space.

4. The error reporting method according to claim 3, characterized in that: The step of converting the error information into an error identifier according to the error structure includes: Obtaining a write pointer of a buffer of the processor platform; Writing the error structure into the data storage area of ​​the buffer according to the write pointer, and updating the error data check value in the data storage area according to the error structure, wherein the data storage area is used to temporarily store the error structure; According to a preset cycle, a plurality of ready flags of the buffer are obtained, and it is determined whether the ready flag is a target ready flag. If the ready flag is the target ready flag, an updated error data check value in the data storage area is read; Determining whether the data in the error structure is normal according to the updated error data check value; When it is determined that the data in the error structure is normal, the normal error structure is parsed according to the preset format to obtain the error identifier.

5. The error reporting method according to claim 1, wherein: The method further comprises: After obtaining the error identifier, outputting a confirmation flag to the basic input and output system; When the basic input / output system receives the confirmation flag, the ready flag and the confirmation flag of the buffer of the processor platform are cleared.

6. The error reporting method according to claim 5, characterized in that: The method further comprises: Before outputting the confirmation flag to the basic input / output system, if the ready flag is the target ready flag, the confirmation flag is obtained according to a preset period, and it is determined whether the confirmation flag is the target preset flag. If the confirmation flag is the target preset flag, the ready flag and the confirmation flag of the buffer of the processor platform are cleared.

7. The error reporting method according to claim 1, characterized in that: The method further comprises: Obtaining a historical error log, and cleaning and feature extracting data in the historical error log to obtain a target historical error log, wherein the feature-extracted data includes at least: error type and occurrence frequency; Training a machine learning model based on the target historical error logs and error categories to obtain an error classification model; The error identification is classified according to the error classification model.

8. The error reporting method according to claim 7, characterized in that: The method further comprises: Priorities of the plurality of error identifiers are determined according to target features of the plurality of error identifiers, wherein the target features include: an impact range of the error identifier and a load status of the processor platform.

9. The error reporting method according to claim 1, characterized in that: The prediction model is constructed based on historical data and real-time operating parameters, including: Acquiring the historical data and the real-time operating parameters, wherein the historical data and the real-time operating parameters include an error log of the basic input / output system, the hardware operating parameters, microcode data of the processor platform, and the environmental monitoring data; The machine learning model is trained according to the historical data and the real-time operating parameters to obtain the prediction model, wherein the machine learning model includes at least one of the following: a random forest or a long short-term memory network.

10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the error reporting method for a processor platform according to any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Error analysis method and system, electronic equipment and medium

    CN119201525A

  • Fault prediction method and device and baseboard management controller

    CN119883843A

  • Server and processor error processing method and device, medium and program product

    CN120086053A

  • Fault analysis method and device for switch chip

    CN120474904A

  • Cache system

    US20110107143A1

Cited By

  • Multi-BIOS (Basic Input / Output System) starting switching method and equipment, storage medium and computer program product

    CN121433987A

  • Multi-bios boot switching method, device, storage medium and computer program product

    CN121433987B