Alarm method and device capable of correcting errors in server test and electronic equipment

By detecting the system event log and system log in the server test log, determining the correctable error set of the device and alarming, the problem of insufficient CE filtering and alarming accuracy is solved, and CE accurate detection and efficient testing is realized.

CN120492283APending Publication Date: 2025-08-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510668364.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In server testing, the accuracy of error (CE) filtering and alarms can be corrected in the prior art, resulting in extended test cycles and CE omissions.

Method used

By detecting the system event log generated by the substrate management controller and the system log generated by the operating system, the target sensor and/or physical location identifier corresponding to the device that can correct errors are obtained, and an alarm is issued when the preset alarm conditions are met.

Benefits of technology

Accurate detection of correctable errors is achieved, CE missed detection is avoided, filtering and alarm accuracy is improved, and testing efficiency is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492283A_ABST
    Figure CN120492283A_ABST
Patent Text Reader

Abstract

The invention provides an alarm method and device capable of correcting errors in server testing and electronic equipment, and relates to the technical field of computers. The method comprises the steps that in the server testing process, correctable errors in testing logs are detected, and the testing logs comprise a system event log generated by a substrate management controller and a system log generated by an operating system; determining a target sensor and / or a physical location identifier corresponding to the equipment with the correctable error under the condition of detecting that the correctable error exists in the test log; obtaining a correctable error set of the equipment from the test log based on a target sensor or a physical location identifier corresponding to the equipment; and when it is determined that the correctable error set satisfies the preset alarm condition of the device, giving an alarm to correctable errors of the device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to a method, device, and electronic device for alarming correctable errors in server testing. Background Art

[0002] With the widespread adoption of PCIe (Peripheral Component Interconnect Express) 4.0 / 5.0 technology, servers are increasingly integrating multiple high-speed devices. This has increased the load on PCIe links, significantly increasing the incidence of correctable errors (CEs). During factory testing of servers, the CE threshold is set to 1 to maximize exposure to hardware vulnerabilities. However, a large number of CEs are intercepted, extending the testing cycle.

[0003] In related technologies, correctable errors can be filtered and alerted through a single log. However, this approach carries the quality risk of missing correctable errors, and the accuracy of filtering and alerting CE issues needs to be improved. Summary of the Invention

[0004] In view of this, the present application provides an alarm method for correcting errors in server testing.

[0005] One aspect of the present application provides an alarm method for correctable errors in server testing, the method comprising: during the server testing process, detecting correctable errors in a test log, the test log comprising a system event log generated by a baseboard management controller and a system log generated by an operating system; upon detecting the presence of a correctable error in the test log, determining a target sensor and / or physical location identifier corresponding to a device where the correctable error occurred; obtaining a correctable error set for the device from the test log based on the target sensor or physical location identifier corresponding to the device; and upon determining that the correctable error set meets a preset alarm condition for the device, alarming the correctable error of the device.

[0006] One aspect of the present application provides an alarm device for correctable errors in server testing, the device comprising: a detection module for detecting correctable errors in a test log during a server test, the test log comprising a system event log generated by a baseboard management controller and a system log generated by an operating system; a determination module for determining, upon detecting the presence of a correctable error in the test log, a target sensor and / or physical location identifier corresponding to a device where the correctable error occurred; an acquisition module for acquiring a correctable error set of the device from the test log based on the target sensor or physical location identifier corresponding to the device; and an alarm module for alarming the correctable error of the device upon determining that the correctable error set meets a preset alarm condition of the device.

[0007] Another aspect of the present application provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0008] Another aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0009] Another aspect of the present application further provides a computer program product, including a computer program or instructions, which implement the steps of the above method when the computer program or instructions are executed by a processor.

[0010] According to the alarm method for correctable errors in server testing provided by the present application, by detecting test logs including SEL logs and system logs, when correctable errors are detected in the test logs, the target sensor and physical location identifier of the device are obtained. For different devices, the target sensor or physical location identifier can be used to obtain the correctable error sets of different devices respectively, thereby realizing joint detection of multiple logs, avoiding CE missed detection, and thus improving the accuracy of CE filtering and alarming. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:

[0012] Figure 1 The following schematically illustrates an exemplary system architecture to which the server test correctable error alarm method of the present application can be applied;

[0013] Figure 2 The flowchart of the server test correctable error alarm method according to an embodiment of the present application is schematically shown;

[0014] Figure 3 Schematically shows a flow chart of a method for alerting correctable errors in server testing according to another embodiment of the present application;

[0015] Figure 4 A schematic diagram of a CE alarm of an SEL log according to an embodiment of the present application is schematically shown;

[0016] Figure 5 A schematic diagram schematically shows a target correctable error block of a RAS log according to an embodiment of the present application;

[0017] Figure 6 A flowchart of a method for determining a physical location identifier according to an embodiment of the present application is schematically shown;

[0018] Figure 7 A schematic diagram of a CE alarm of a system log according to an embodiment of the present application is schematically shown;

[0019] Figure 8 Schematically shows a flow chart of a method for alerting a correctable error in a server test according to another embodiment of the present application;

[0020] Figure 9 A schematic diagram of a module of a correctable error alarm device in server testing according to an embodiment of the present application is shown; and

[0021] Figure 10 A block diagram of an electronic device suitable for a correctable error alarm method in server testing according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0022] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0023] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0025] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0026] In the embodiments of this application, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of all data involved (including, but not limited to, user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard the security of user personal information, network security, and national security.

[0027] Figure 1 An exemplary system architecture to which the server test correctable error alarm method of the present application can be applied is schematically shown.

[0028] like Figure 1 As shown, the system architecture includes a preset alarm condition management module 110 , a test execution module 120 and a central server processing module 130 .

[0029] The preset alarm condition management module 110 is responsible for managing, invoking, and updating preset alarm conditions, and provides alarm rules for the test execution module 120. The test execution module 120 is used to perform tests on services and transmit test data to the central server processing module 130. The central server processing module 130 is responsible for storing and managing the data transmitted by the preset alarm condition management module 110 and the test execution module 120, and providing corresponding data support for the preset alarm condition management module 110 and the test execution module 120.

[0030] Figure 2 The flowchart of the server test correctable error alarm method according to an embodiment of the present application is schematically shown.

[0031] like Figure 2 As shown, the method includes operations S201 to S204.

[0032] In operation S201 , during a server test, correctable errors in a test log are detected. The test log includes a system event log generated by a baseboard management controller and a system log generated by an operating system.

[0033] In operation S202 , when it is detected that a correctable error exists in the test log, a target sensor and / or a physical location identifier corresponding to the device where the correctable error occurs is determined.

[0034] In operation S203 , a correctable error set of the device is obtained from the test log based on the target sensor or physical location identifier corresponding to the device.

[0035] In operation S204, if it is determined that the correctable error set meets a preset alarm condition of the device, an alarm is issued for the correctable error of the device.

[0036] The System Event Log (SEL) is a log file generated by the Baseboard Management Controller (BMC). It records various events and errors that occur in the server hardware system. Based on the SEL log, you can obtain information such as the specific time of the event, event type, event severity, and sensor number.

[0037] System logs include various logs generated by the operating system kernel, drivers, and applications. For example, dmesg logs and message logs can be used to obtain information such as the specific time of an event, event type, and the physical location of the device.

[0038] For example, logs such as SEL, dmesg, and message may be obtained in real time as initial processing units. CE in the SEL log may be detected first, and then CE in the dmesg and message logs may be detected in sequence.

[0039] In hardware systems like servers, sensors are deployed on key devices and components to monitor their operating status in real time. When an abnormal event occurs on a device, the corresponding target sensor detects the parameter changes associated with the abnormal event, triggering the sensor's corresponding error logging mechanism and recording correctable error information in the system event log. Therefore, if a CE is found in the system event log, information about the device's target sensor can be obtained from the system event log.

[0040] The physical location identifier is used to uniquely identify the physical installation location of the device in the server. The physical location identifier may be, for example, a BDF (Bus-Device-Function).

[0041] In some embodiments, the correctable error set of the device can be extracted from the SEL log using the target sensor of the device. In other embodiments, the correctable error set of the device can also be extracted from the system log using the physical location identifier of the device.

[0042] Preset error reporting conditions refer to a series of rules pre-set for a device to determine whether the device's correctable errors have reached the level that requires an error report. Based on the different device models, package orders, and test items, different preset error reporting conditions can be set for different devices.

[0043] When it is determined that the first correctable error set does not meet the preset error reporting condition of the device, the correctable errors may be filtered, that is, no interception and error reporting are performed, thereby saving testing time.

[0044] According to an embodiment of the present application, by detecting test logs including SEL logs and system logs, when correctable errors are detected in the test logs, the target sensor and physical location identifier of the device are obtained. For different devices, the target sensor or physical location identifier can be used to obtain the correctable error sets of different devices respectively, thereby realizing joint detection of multiple logs, avoiding CE missed detection, and thus improving the accuracy of CE filtering and alarming.

[0045] Figure 3 The flowchart of the method for alarming correctable errors in server testing according to another embodiment of the present application is schematically shown.

[0046] like Figure 3 As shown, the method includes operations S301 to S308.

[0047] In operation S301, a CE in a system event log is detected. If a CE exists in the system event log, operation S302 is performed. If a CE does not exist in the system event log, operation S304 is performed.

[0048] In operation S302 , a target sensor and a physical location identifier of a device are obtained.

[0049] In operation S303, a correctable error set is acquired based on the target sensor. If the correctable error set meets the preset alarm condition, operation S307 is performed; if the correctable error set does not meet the preset alarm condition, operation S304 is performed.

[0050] In operation S304, the CE in the system log is detected. If the CE exists in the system log, operation S305 is executed; if the CE does not exist in the system log, operation S308 is executed.

[0051] In operation S305 , a physical location identifier of the device is obtained.

[0052] In operation S306, a correctable error set is obtained based on the physical location identifier of the device. If the correctable error set meets the preset alarm condition, operation S307 is performed. If the correctable error set does not meet the preset alarm condition, operation S308 is performed.

[0053] In operation S307 , an alarm is issued to the CE.

[0054] In operation S308 , CEs are filtered and other CE detections are continued.

[0055] This enables joint detection of multiple logs, avoids missed CE detection, and improves the accuracy of CE filtering and alarming.

[0056] According to an embodiment of the present application, when CE is detected in the test log, determining the target sensor and / or physical location identifier corresponding to the device where CE occurs may include: when CE is detected in the system event log, determining the target sensor corresponding to the device where CE occurs; based on a mapping relationship library between sensors and physical location identifiers, using the target sensor to query the physical location identifier of the device in the mapping relationship library.

[0057] For example, when a CE alarm is detected in the SEL log, the SEL log may be parsed based on a regular expression to obtain a target sensor corresponding to the device where the CE occurs.

[0058] Figure 4 The diagram schematically shows a CE alarm of the SEL log according to an embodiment of the present application.

[0059] like Figure 4 As shown, when a "Bus Correctable error" alarm is detected in the SEL log, the SEL log can be parsed to obtain the target sensor information "GAI_GPU5_Status".

[0060] The mapping database is a structured data set used to store the correspondence between sensors and the physical location identifiers of devices. Sensors can be used to query the mapping database for their corresponding physical location identifiers. Similarly, physical location identifiers can be used to reversely query the corresponding sensors.

[0061] According to the embodiment of the present application, the physical location identifier of the device where the correctable error occurs can be quickly queried through the mapping relationship library, and the physical location identifier can be used to facilitate the positioning of the CE in other logs.

[0062] According to an embodiment of the present application, when it is determined that the physical location identifier of the device does not exist in the mapping relationship library, the reliability, availability, and serviceability log (RAS log) generated by the baseboard management controller is obtained; based on the reliability, availability, and serviceability log and CE, the physical location identifier of the device is determined.

[0063] Because devices (such as GPUs and FPGAs) may be dynamically assigned to different virtual machines or tasks, their physical location identifiers (BDFs) may change with load changes. Alternatively, if the mapping relationship library is not regenerated after hardware replacement (such as replacing a faulty memory stick or expanding a PCIe card), the old mapping entries will become invalid, resulting in the device's physical location identifier not existing in the mapping relationship library.

[0064] The RAS log contains various key information about system operation, such as hardware failures and software errors. The RAS log contains more detailed information than the SEL log, such as physical location identification information.

[0065] Illustratively, the RAS log can be obtained from the baseboard management controller through a log collector, a script, or an interface.

[0066] Illustratively, determining the physical location of a device based on the RAS log and CE can include: accurately matching the SEL log with the RAS based on information such as the target sensor, timestamp, and log index corresponding to the CE in the SEL log. In the RAS log, the correctable error blocks for the same CE are identified, and device-related information, such as bus, device, and function information, is extracted from the correctable error blocks to obtain the device's physical location.

[0067] Figure 5 A schematic diagram schematically shows a target correctable error block of a RAS log according to an embodiment of the present application.

[0068] like Figure 5 As shown in the figure, based on the target sensor "GAI_GPU5_Status" and the timestamp "04 / 06 / 2025|00:04:16", the correctable error block of the same CE can be matched. By extracting Bus: 0xba, Device: 0x00, and Function: 0x00 from the correctable error block, the physical location identifier of the device is obtained.

[0069] After obtaining the physical location of the device, the upstream and downstream ports of the device can be obtained based on the PCIe topology of the device, and a mapping relationship can be established with the corresponding target sensor. The established mapping relationship can be updated to the mapping relationship library to facilitate rapid matching the next time CE appears.

[0070] According to an embodiment of the present application, by parsing the RAS log and obtaining the physical location identifier that has a mapping relationship with the target sensor, dynamic binding of the sensor and the physical location identifier can be achieved, thereby facilitating the use of the sensor and the physical location identifier to accurately match and filter the CE of different logs.

[0071] Figure 6 A flowchart of a method for determining a physical location identifier according to an embodiment of the present application is schematically shown.

[0072] like Figure 6 As shown, when a CE is detected in the SEL log, the system event log is parsed to determine the target sensor 601 of the device. Based on the target sensor 601 of the device, the physical location identifier 603 of the device is queried in the mapping relationship library 602. If the physical location identifier 603 of the device can be queried, the physical location identifier 603 of the device is obtained. If the physical location identifier 603 of the device cannot be queried, the RAS log 604 is obtained. Based on the RAS log 603, the physical location identifier 605 of the device is obtained.

[0073] According to an embodiment of the present application, obtaining a correctable error set of a device from a test log based on a target sensor or physical location identifier corresponding to the device may include: obtaining a correctable error set from an SEL log based on a target sensor corresponding to the device; and deleting CEs in the system log that are duplicated in the correctable error set based on the physical location identifier of the device.

[0074] For devices with sensors, all CEs that occur on the device can be recorded in the SEL log, and some CEs of the device may also be recorded in the system log. For devices without sensors, CEs are recorded in the system log.

[0075] Schematically, you can use the BMC tool to export the SEL log. Based on the target sensor and error type (such as the correctable error type), filter the SEL log to obtain all CE records of the device. From all CE records, extract key fields such as timestamp, sensor ID, and error description to obtain the correctable error set.

[0076] According to the embodiments of the present application, since the system log and the SEL log may record the same hardware error event repeatedly, resulting in duplicate alarms, the CEs in the system log that are duplicated with the SEL log are deleted based on the physical location identifier. In this way, after the SEL log is checked, the CEs in the system log are checked again to avoid duplicate checks of the same CE.

[0077] According to an embodiment of the present application, deleting the duplicate CEs in the system log and the correctable error set based on the physical location identifier of the device may include: matching the duplicate correctable error blocks in the system log based on the physical location identifier of the device and the timestamp information of each CE in the correctable error set.

[0078] Because CE alarms for the same device may be captured multiple times by different log sources, such as the SEL log and the system log, duplicate records may be recorded. For example, some hardware errors in the system log may also be recorded in the SEL log. Duplicate CE alarms consume storage space and interfere with the alarm system (for example, by triggering the same alarm multiple times).

[0079] Illustratively, the system log can be parsed first. The error type identified in the system log is a correctable error type, and a CE entry containing the physical location identifier of the device is included. Each CE entry in the system log is traversed, and its physical location identifier and timestamp are extracted. Within the correctable error set in the SEL, CE entries with the same physical location and the same or similar timestamps are searched. If a match is found, the CE entry is marked as a duplicate correctable error block. Once the duplicate correctable error block is obtained, it can be deleted from the system log.

[0080] According to the embodiments of the present application, by matching the physical location identifier and the timestamp, duplicate correctable error blocks in the system log can be efficiently identified and deleted, thereby optimizing the management of the system log, avoiding the same error triggering multiple alarms, and improving the processing efficiency of the log alarms.

[0081] According to an embodiment of the present application, deleting CEs in the system log that are repeated in the correctable error set based on the physical location identifier of the device can also include: obtaining the physical location identifier of the upstream device and / or downstream device of the link where the device is located based on the physical location identifier of the device; matching the repeated correctable error blocks in the system log based on the physical location identifier of the upstream device and / or downstream device and the timestamp information of each CE in the correctable error set.

[0082] In a PCIe link, CE events can propagate along the link, causing the same CE event to be recorded by multiple devices. For example, an upstream switch port error may cause multiple downstream devices to report anomalies; or a signal problem on a PCIe device may trigger an error on the link of an upstream device.

[0083] Schematically, all CE entries for the device, the upstream device, and / or the downstream device on the link where the device is located can be filtered from the system log based on the physical location identifier. The CE time series of the device can be constructed in chronological order. If CEs of upstream and downstream devices exist in the same time window, they are merged into the same CE entry. Each CE entry in the system log is traversed to extract its physical location identifier and timestamp. In the correctable error set of the SEL log, CE entries with the same physical location and the same or similar timestamps are searched. If a match is successful, the CE entry in the system log is marked as a duplicate correctable error block and deleted.

[0084] According to an embodiment of the present application, by integrating the CEs of the uplink device and / or downlink device of the link where the device is located, the accuracy of CE deduplication can be effectively improved.

[0085] According to an embodiment of the present application, obtaining a correctable error set of a device from a test log based on a target sensor or physical location identifier corresponding to the device may also include: when it is determined that the CE does not have a corresponding target sensor, obtaining a correctable error set corresponding to the physical location identifier from a system log based on the physical location identifier.

[0086] When a CE is detected in a system log, the system log may be parsed to obtain a physical location identifier of a device where the CE occurs.

[0087] Figure 7 A schematic diagram of a CE alarm of a system log according to an embodiment of the present application is schematically shown.

[0088] like Figure 7 As shown in the figure, according to the "device_id: 0000:1c:00.0" information in the system log, the physical location identifier 0000:1c:00.0 of the device can be obtained.

[0089] Based on the device's physical location, the corresponding target sensor is searched in the mapping relationship library. If the corresponding target sensor can be found, the correctable error set is obtained from the system log based on the target sensor and duplicate CEs in the system log are deleted. If the corresponding target sensor cannot be found, the correctable error set is obtained from the system log based on the device's physical location.

[0090] Illustratively, the SEL log can be checked first, and then the system log can be checked after the SEL log check is complete. Because duplicate CEs in the system log have already been deleted during the SEL check, CE alarms that do not appear in the SEL log can be detected from the system log based on the physical location identifier, thus achieving comprehensive detection of CE alarms.

[0091] According to an embodiment of the present application, obtaining a correctable error set corresponding to the physical location identifier from a system log based on the physical location identifier may include: obtaining a first correctable error set from the current system log based on the physical location identifier; obtaining a second correctable error set, the second correctable error set being obtained from the historical system log based on the physical location identifier; and obtaining a correctable error set based on the first correctable error set and the second correctable error set.

[0092] Historical system logs refer to system logs at historical time periods during the server testing process.

[0093] The historical system log may be a system log generated in a time period before the current time period. Based on the physical location identifier of the device, the second correctable error set is obtained from the historical system log of the previous time period and saved, for example, in a central server.

[0094] The present application is not limited to this. The historical system log may also be the system log before the restart in the event that a restart occurs during the test process. Since the service test is usually a diskless system, if a system restart occurs during the test process, the system log will not be retained after the restart. By obtaining the second correctable error set from the historical system log, the integrity of the CE can be ensured.

[0095] After obtaining the first correctable error set from the current system log, the saved second correctable error set can be directly obtained. The union of the first correctable error set and the second correctable error set is used as the correctable error set. The correctable error set is then used to overwrite the saved second correctable error set, which is then used as the updated second correctable error set. This allows the second correctable error set to be directly obtained during subsequent testing.

[0096] According to an embodiment of the present application, by integrating the first correctable error set of the current system log and the second correctable error set of the historical system log, the integrity of CE detection in the system log can be ensured, and CE omission can be avoided.

[0097] Figure 8 The flowchart of the method for alarming correctable errors in server testing according to another embodiment of the present application is schematically shown.

[0098] like Figure 8 As shown, the method includes operations S801 to S808.

[0099] In operation S801, a CE in a test log is detected. If it is determined that a corresponding target sensor exists for the CE, operation S802 is performed. If it is determined that a corresponding target sensor does not exist for the CE, operation S805 is performed.

[0100] In operation S802, a correctable error set is obtained based on the target sensor corresponding to the device. If the correctable error set meets the preset alarm condition, operation S808 is performed. If the correctable error set does not meet the preset alarm condition, operation S809 is performed.

[0101] In operation S803 , duplicate correctable error blocks are matched in the system log based on the physical location identifier of the device.

[0102] In operation S804, duplicate correctable error blocks in the system log are deleted.

[0103] In operation S805 , a first correctable error set is obtained from a current system log based on the physical location identifier.

[0104] In operation S806 , a second correctable error set is obtained from a historical system log based on the physical location identifier.

[0105] In operation S807, a correctable error set is obtained based on the first correctable error set and the second correctable error set. If the correctable error set meets the preset alarm condition, operation S808 is performed. If the correctable error set does not meet the preset alarm condition, operation S809 is performed.

[0106] In operation S808 , an alarm is issued to the CE.

[0107] In operation S809 , the CEs are filtered and other CEs are detected.

[0108] In this way, the joint detection of multiple logs is achieved, avoiding missed CE detection, and the filtering of repeated CE is achieved, avoiding the problem of repeated alarms during multiple log detection.

[0109] According to an embodiment of the present application, the preset alarm conditions include the statistical range of CE and the threshold conditions for triggering the alarm; when it is determined that the correctable error set meets the preset alarm conditions of the device, alarming the CE of the device may include: obtaining the number of CEs within the statistical range from the correctable error set; when the number of CEs meets the threshold conditions, alarming the CE of the device.

[0110] CE statistics can be calculated by various scopes, including full test, test item, minute, and hour. A full test scope counts all CEs in the correctable error set. A test item scope counts all CEs after the start time of the current test item. Minute and hourly scopes count CEs every minute or hour.

[0111] In related technologies, when the statistical range of CE is per minute and per hour, the seconds or minute-second part of the CE alarm timestamp is conventionally removed to count the number of CEs per minute or per hour. However, it is impossible to count the number of CEs that span minutes or hours, which may easily cause quality risks.

[0112] According to an embodiment of the present application, when the statistical scope is characterized as statistics collected at predetermined time intervals, the number of CEs in the predetermined time interval is obtained from the correctable error set in a sliding window manner. This allows detection of all CEs and avoids CE omissions.

[0113] For example, using a sliding window approach to count the number of CEs every 60 seconds or 60 minutes, the calculation formula is as follows:

[0114]

[0115] N(t) is the number of CEs in each time interval, t is the timestamp of the current traversed CE, k is the total number of CEs, t i is the timestamp of the i-th CE, and 60 is the time interval. I is an indicator function that determines whether the timestamp t_i of the i-th log is within the window [t, t+60). If so, it returns 1, otherwise it returns 0.

[0116] The number of CEs counted in each time interval is compared with the threshold condition. If the counted number is greater than or equal to the threshold, an alarm is issued; otherwise, the CE is filtered out and other CEs are processed until all CEs are processed and other server tests are continued.

[0117] In some embodiments, the preset alarm conditions also include conditional information such as the device model name, package order, test items, automatic adjustment, filter status, and update time. Among them, the model name and package order are used to limit the server type, and only matching servers will execute the above alarm method. The test item identifier is used to execute the test item of the filtering rule. When it is empty and All, all test items are applicable. The automatic adjustment identifier indicates whether to automatically adjust according to the predicted threshold. Some servers have strict control over the tested CE and do not allow automatic adjustment of the CE threshold. The filter status indicates the status of the current CE filtering rule, which is unavailable when invalid. The update time indicates the update time of the CE threshold filtering rule.

[0118] According to an embodiment of the present application, the alarm method may further include: inputting device attribute information and test condition information into a threshold prediction model to output a predicted threshold; and updating the threshold condition using the predicted threshold. The threshold prediction model is trained by: obtaining historical test data; obtaining device attribute information, test condition information, and the historical number of correctable errors from the historical test data where correctable errors occurred; and training an initial threshold prediction model using the device attribute information and test condition information as input data and the historical number of correctable errors as a label to obtain a threshold prediction model.

[0119] Device attribute information may include the device model, package order, sensor, physical location identifier, etc. Test condition information may include test items, test start time, test end time, etc.

[0120] Exemplarily, after obtaining historical test data, the test data is parsed to obtain device attribute information such as device model, package order, sensor, physical location identification, and other information, as well as test condition information such as test items, test start time, and test end time. Repeated CE filtering of the same log and repeated CE filtering of multivariate logs are performed on the CE data to obtain the number of historical correctable errors. The device attribute information and test condition information are uniquely encoded and normalized to obtain input data. The initial threshold prediction model can use a random forest regression model, which is trained with the number of historical correctable errors as a label to obtain a threshold prediction model. The present application is not limited to this, and the initial threshold prediction model can also be a gradient boosting tree model, a Bayesian regression model, and the like.

[0121] It is also possible to periodically obtain new test data, use the new data to retrain, and dynamically update the prediction threshold.

[0122] Threshold generation formula: Predicted Threshold = μ + k * σ

[0123] Among them, μ is the threshold mean predicted by the model; σ is the threshold standard deviation predicted by the model; k is the confidence interval coefficient, and the initial setting of k can be set to 1.96 to cover the 95% confidence level.

[0124] After obtaining the predicted threshold, you can match the predicted threshold with the device based on information such as model, package order, CE location, threshold conditions, etc. Automatic adjustment type rules directly modify the threshold, while non-automatic adjustment rules prompt recommended thresholds, which are manually reviewed for modification.

[0125] According to the embodiments of this application, by dynamically updating threshold conditions through a threshold prediction model, reliable CE threshold adjustment recommendations can be provided, providing effective data support for CE analysis. This reduces the time cost of analyzing large amounts of data and improves problem-solving efficiency. The automatic adjustment mode can further save labor and development costs, improving production efficiency.

[0126] Figure 9 The module diagram of the correctable error alarm device in server testing according to an embodiment of the present application is schematically shown.

[0127] like Figure 9 The device 900 shown includes a detection module 910 , a determination module 920 , an acquisition module 930 and an alarm module 940 .

[0128] A detection module 910 is configured to detect correctable errors in a test log during a server test, wherein the test log includes a system event log generated by a baseboard management controller and a system log generated by an operating system;

[0129] A determination module 920 is configured to, when detecting that a correctable error exists in the test log, determine a target sensor and / or a physical location identifier corresponding to a device where the correctable error occurs;

[0130] An acquisition module 930 is configured to acquire a correctable error set of the device from the test log based on a target sensor or a physical location identifier corresponding to the device;

[0131] The alarm module 940 is configured to generate an alarm for the correctable error of the device when it is determined that the correctable error set meets a preset alarm condition of the device.

[0132] According to an embodiment of the present application, the determination module 920 includes a target sensor determination submodule and a query submodule.

[0133] The target sensor determination submodule is used to determine the target sensor corresponding to the device where the correctable error occurs when a correctable error is detected in the system event log.

[0134] The query submodule is used to query the physical location identifier of the device in the mapping relationship library based on the mapping relationship library of sensors and physical location identifiers using the target sensor.

[0135] According to an embodiment of the present application, the determination module 920 further includes an acquisition submodule and a physical location identification determination submodule.

[0136] The acquisition submodule is used to acquire the reliability, availability and serviceability logs generated by the baseboard management controller when it is determined that the physical location identifier of the device does not exist in the mapping relationship library.

[0137] The physical location identification determination submodule is used to determine the physical location identification of the device based on reliability, availability and serviceability logs and correctable errors.

[0138] According to an embodiment of the present application, the acquisition module 930 includes a first acquisition submodule and a deletion submodule.

[0139] The first acquisition submodule is configured to acquire a correctable error set from a system event log based on a target sensor corresponding to the device.

[0140] The deletion submodule is used to delete the correctable errors in the system log that are repeated in the correctable error set based on the physical location identifier of the device.

[0141] According to an embodiment of the present application, the deletion submodule includes at least one of a first deletion unit and a second deletion unit:

[0142] The first deleting unit is configured to match repeated correctable error blocks in the system log based on a physical location identifier of the device and timestamp information of each correctable error in the correctable error set.

[0143] The second deleting unit includes a first acquiring subunit and a first matching subunit.

[0144] The first acquiring subunit is configured to acquire the physical location identifier of an uplink device and / or a downlink device of a link where the device is located based on the physical location identifier of the device.

[0145] The first matching subunit is configured to match repeated correctable error blocks in the system log based on the physical location identifier of the upstream device and / or the downstream device and the timestamp information of each correctable error in the correctable error set.

[0146] According to an embodiment of the present application, the acquisition module 930 further includes a second acquisition submodule, a third acquisition submodule, and a merging submodule.

[0147] The second acquisition submodule is configured to acquire a first correctable error set from a current system log based on the physical location identifier when it is determined that the correctable error has no corresponding target sensor.

[0148] The third acquisition submodule is configured to acquire a second correctable error set, where the second correctable error set is acquired from a historical system log based on the physical location identifier.

[0149] The merging submodule is configured to obtain a correctable error set based on the first correctable error set and the second correctable error set.

[0150] According to an embodiment of the present application, the preset alarm conditions include a statistical range of correctable errors and a threshold condition for triggering an alarm; the alarm module 940 includes a correctable error number acquisition submodule and an alarm submodule.

[0151] The correctable error number acquisition submodule is used to obtain the correctable error number within a statistical range from the correctable error set.

[0152] The alarm submodule is used to issue an alarm for the correctable errors of the device when the number of correctable errors meets a threshold condition.

[0153] According to an embodiment of the present application, the device further includes:

[0154] The prediction module is used to input device attribute information and test condition information into the threshold prediction model and output the predicted threshold.

[0155] The updating module is used to update the threshold condition using the predicted threshold.

[0156] Among them, the threshold prediction model is obtained through training of the following modules:

[0157] The test data acquisition module is used to obtain historical test data.

[0158] An information acquisition module is used to obtain attribute information of devices where correctable errors have occurred, test condition information, and the number of historical correctable errors from historical test data;

[0159] The training module is used to take device attribute information and test condition information as input data and the number of historical correctable errors as labels to train the initial threshold prediction model to obtain a threshold prediction model.

[0160] According to the embodiments of the present application, any number of modules, submodules, units, and subunits, or at least part of the functions of any number of them, can be implemented in one module. According to the embodiments of the present application, any one or more of the modules, submodules, units, and subunits can be split into multiple modules for implementation. According to the embodiments of the present application, any one or more of the modules, submodules, units, and subunits can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present application, one or more of the modules, submodules, units, and subunits can be at least partially implemented as a computer program module, which can perform the corresponding functions when the computer program module is executed.

[0161] For example, any number of the detection module 910, determination module 920, acquisition module 930, and alarm module 940 can be combined into a single module / unit / sub-unit, or any one of these modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functionality of one or more of these modules / units / sub-units can be combined with at least part of the functionality of other modules / units / sub-units and implemented in a single module / unit / sub-unit. According to an embodiment of the present application, at least one of the detection module 910, determination module 920, acquisition module 930, and alarm module 940 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware by any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the detection module 910 , the determination module 920 , the acquisition module 930 , and the alarm module 940 may be at least partially implemented as a computer program module, and when the computer program module is executed, the corresponding function may be executed.

[0162] Figure 10 A block diagram of an electronic device suitable for a correctable error alarm method in server testing according to an embodiment of the present application is schematically shown.

[0163] Figure 10 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present application is schematically shown. Figure 10 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0164] Electronic device is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device may also refer to various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the present application described and / or claimed herein.

[0165] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. RAM 1003 may also store various programs and data required for the operation of device 1000. Computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0166] Multiple components in electronic device 1000 are connected to I / O interface 1005, including: an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0167] Computing unit 1001 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1001 performs the various methods and processes described above, such as the alarm method. For example, in some embodiments, the alarm method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the alarm method described above may be performed. Alternatively, in other embodiments, computing unit 1001 may be configured to perform the alarm method via any other suitable means (e.g., via firmware).

[0168] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0169] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable test device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0170] In the context of this application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0171] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0172] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0173] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0174] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.

[0175] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.

Claims

1. A method for warning of correctable errors in server testing, characterized in that: include: During the server test process, correctable errors in the test log are detected, wherein the test log includes a system event log generated by the baseboard management controller and a system log generated by the operating system; In the case where a correctable error is detected in the test log, determining a target sensor and / or a physical location identifier corresponding to a device where the correctable error occurs; Acquire a correctable error set of the device from the test log based on a target sensor or a physical location identifier corresponding to the device; If it is determined that the correctable error set meets a preset alarm condition of the device, an alarm is issued for the correctable error of the device.

2. The alarm method according to claim 1, characterized in that: When a correctable error is detected in the test log, determining a target sensor and / or a physical location identifier corresponding to a device where the correctable error occurs includes: In a case where a correctable error is detected in the system event log, determining a target sensor corresponding to a device where the correctable error occurs; Based on a mapping relationship library between sensors and physical location identifiers, the physical location identifier of the device is queried in the mapping relationship library using the target sensor.

3. The alarm method according to claim 2, further comprising: When it is determined that the physical location identifier of the device does not exist in the mapping relationship library, obtaining a reliability, availability, and serviceability log generated by a baseboard management controller; A physical location identification of the device is determined based on the reliability, availability, and serviceability logs and the correctable errors.

4. The alarm method according to any one of claims 1 to 3, characterized in that: The acquiring, from the test log, a correctable error set of the device based on a target sensor or a physical location identifier corresponding to the device, includes: Obtaining the correctable error set from the system event log based on a target sensor corresponding to the device; and Based on the physical location identifier of the device, correctable errors in the system log that are repeated in the correctable error set are deleted.

5. The method according to claim 4, wherein the deleting, based on the physical location identifier of the device, correctable errors in the system log that are duplicated in the correctable error set comprises at least one of the following situations: Scenario 1: matching repeated correctable error blocks in the system log based on the physical location identifier of the device and the timestamp information of each correctable error in the correctable error set; Scenario 2: Based on the physical location identifier of the device, the physical location identifier of the uplink device and / or downlink device of the link where the device is located is obtained; Based on the physical location identifier of the upstream device and / or the downstream device and the timestamp information of each correctable error in the correctable error set, duplicate correctable error blocks are matched in the system log.

6. The alarm method according to claim 1, characterized in that: Acquiring a correctable error set of the device from the test log based on a target sensor or a physical location identifier corresponding to the device, further comprising: If it is determined that the correctable error has no corresponding target sensor, obtaining a first correctable error set from a current system log based on the physical location identifier; Obtaining a second correctable error set, where the second correctable error set is obtained from a historical system log based on the physical location identifier; The correctable error set is obtained based on the first correctable error set and the second correctable error set.

7. The alarm method according to claim 1, characterized in that: The preset alarm conditions include the statistical range of correctable errors and the threshold conditions for triggering the alarm; The step of, when determining that the correctable error set satisfies a preset alarm condition of the device, issuing an alarm for the correctable error of the device, includes: Obtaining the number of correctable errors within the statistical range from the correctable error set; When the number of correctable errors meets the threshold condition, an alarm is issued for the correctable errors of the device.

8. The alarm method according to claim 7, characterized in that: The method further comprises: Input device attribute information and test condition information into a threshold prediction model and output a predicted threshold; updating the threshold condition using the predicted threshold, The threshold prediction model is trained in the following way: Get historical test data; Acquire attribute information of devices where correctable errors occur, test condition information, and the number of historical correctable errors from the historical test data; The device attribute information and test condition information are used as input data, and the number of historical correctable errors is used as a label to train an initial threshold prediction model to obtain the threshold prediction model.

9. An alarm device for correctable errors in server testing, characterized in that: include: A detection module, configured to detect correctable errors in a test log during a server test, wherein the test log includes a system event log generated by a baseboard management controller and a system log generated by an operating system; a determination module, configured to, when detecting that a correctable error exists in the test log, determine a target sensor and / or a physical location identifier corresponding to a device where the correctable error occurs; an acquisition module, configured to acquire a correctable error set of the device from the test log based on a target sensor or a physical location identifier corresponding to the device; The alarm module is configured to generate an alarm for the correctable errors of the device when it is determined that the correctable error set meets a preset alarm condition of the device.

10. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 8.