Method and apparatus for controlling server fault information, storage medium and electronic device

By detecting and filtering PCIe CE error descriptions on the server, the problem of low server stability was solved, and timely processing and cleanup of fault information were achieved, thereby improving the server's operational stability.

WO2026113672A1PCT designated stage Publication Date: 2026-06-04INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2025-10-13
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

In existing technologies, the accumulation of PCIe CE errors in servers leads to low operational stability and the inability to report errors in a timely manner, making it impossible to distinguish which error information needs to be reported.

Method used

By querying the number and time of candidate description information in the server register, detecting whether the number matches the threshold, filtering out description information to be deleted, and determining whether to transmit updated description information to the controller, timely processing and cleanup of fault information can be achieved.

Benefits of technology

This improves the stability of server operation, ensures that maintenance personnel can handle important fault information in a timely manner, and avoids excessive fault information in the registers from affecting server operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025127363_04062026_PF_FP_ABST
    Figure CN2025127363_04062026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method and apparatus for controlling server fault information, a storage medium and an electronic device. The method comprises: querying the number of candidates of candidate description information accumulated in a register of a server at a current time, and extracting an initial time at which a first piece of description information among the candidate description information was stored in the register; detecting whether a first matching condition is satisfied between the number of candidates and a target number threshold to obtain a first detection result, and detecting whether a second matching condition is satisfied between a target duration between the current time and the initial time and a duration threshold to obtain a second detection result; on the basis of the first detection result and the second detection result, filtering the candidate description information for target description information to be deleted, and determining whether to transmit, to a controller in the server, updated description information stored in the register; and deleting the target description information, and when it is determined that the updated description information is to be transmitted to the controller, transmitting the updated description information to the controller.
Need to check novelty before this filing date? Find Prior Art

Description

Methods and devices for controlling server fault information, storage media and electronic devices

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411706757.2, filed on November 26, 2024, entitled “Control Method and Apparatus, Storage Medium and Electronic Device for Server Fault Information”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of computers, and more specifically, to a method and apparatus for controlling server fault information, a storage medium, and an electronic device. Background Technology

[0004] The server has a bus port deployed on it. The bus port can be used, but is not limited to, to connect bus devices. For example, during the operation of the server, the bus port and bus devices may fail, such as CE (Corrected Error).

[0005] Taking a server with a PCIe (PCI Express, high-speed serial computer expansion bus standard) port as an example, in related technologies, for the processing of PCIe CEs, Intel CPUs (Central Processing Units) have set an AER (Advanced Error Capabilities and Control Register) to identify a threshold for the number of CEs. When the number of PCIe CEs accumulates to the threshold, the error information is written into the AER register. In AMD CPUs, only a threshold is set for CEs to be reported to the BMC (Baseboard Management Controller) and OS (Operating System). This threshold is the common threshold for all devices on the server and is often large, making it impossible to report in a timely manner.

[0006] Understandably, the way PCIe CE is handled in related technologies can lead to the continuous accumulation of CEs, which may eventually evolve into UCEs (Uncorrectable Errors). It is also impossible to distinguish which error messages need to be reported, and they cannot be reported in a timely manner, resulting in low stability of server operation. Summary of the Invention

[0007] This application provides a method and apparatus for controlling server fault information, a storage medium, and an electronic device, in order to at least solve the problem of low stability in the operation of servers in related technologies.

[0008] According to one embodiment of this application, a method for controlling server fault information is provided. The method is applied to target firmware in a server. The method includes: querying the number of candidate description information accumulated in a register in the server at the current time, and extracting the initial time for storing the first description information in the candidate description information into the register, wherein the initial time is earlier than the current time, and the register is configured to store description information of faults occurring on a bus port on the server and description information of faults occurring on a bus device connected to the bus port; detecting whether the number of candidates and a target number threshold satisfy a first matching condition, obtaining a first detection result, and detecting the target number between the current time and the initial time. The second detection result is obtained by determining whether the duration and duration threshold meet the second matching condition. The target quantity threshold is determined by the target firmware based on the impact of bus port and bus device failures on the server's operating performance. Based on the first and second detection results, target description information to be deleted is filtered from the candidate description information, and it is determined whether to transmit the updated description information stored in the register to the controller in the server. The updated description information is the description information stored in the register after the target description information, and the candidate description information includes the updated description information. The target description information is deleted, and if it is determined that the updated description information should be transmitted to the controller, the updated description information is transmitted to the controller.

[0009] In one exemplary embodiment, the server includes an adjustment device connected to target firmware. Before detecting whether a first matching condition is met between the candidate number and a target number threshold and obtaining a first detection result, the method further includes: receiving a target adjustment request, wherein the target adjustment request is used to request that the number threshold corresponding to the number of description information stored in the register be adjusted to the target number threshold; and sending the target adjustment request to the adjustment device, wherein the adjustment device is configured to execute the target adjustment request.

[0010] In one exemplary embodiment, the server includes a target interface and receives a target adjustment request, including: detecting a first editing operation performed on a first field on the target interface and detecting a second editing operation performed on a second field on the target interface, wherein the first field is used to set the degree of impact of bus port and bus device failures on the server's operating performance, the second field is used to set a threshold value for the number of descriptive information stored in a register corresponding to the degree of impact of bus port and bus device failures on the server's operating performance, the first editing operation is used to set the degree of impact of bus port and bus device failures on the server's operating performance to a target degree of impact, and the second editing operation is used to adjust the threshold value for the number of descriptive information stored in the register corresponding to the target degree of impact to a target threshold value; generating a target adjustment request corresponding to the first editing operation and the second editing operation, and transmitting the target adjustment request to the target firmware.

[0011] In one exemplary embodiment, after sending a target adjustment request to the adjustment device, the method further includes: defining a threshold identifier in a configuration file, wherein the value of the threshold identifier is used to identify the quantity threshold corresponding to the quantity of descriptive information stored in the register; and calling an entry function to adjust the value of the threshold identifier to the target quantity threshold.

[0012] In one exemplary embodiment, the server includes a clock chip connected to target firmware. The clock chip is configured to record the current time of the server. Before querying the number of candidate descriptions accumulated in the registers of the server at the current time, the method includes: detecting a target address of the clock chip, wherein the target address is configured to extract the time recorded in the clock chip; and extracting the time recorded in the clock chip as the current time by accessing the target address.

[0013] In an exemplary embodiment, querying the candidate number of candidate description information accumulated in the registers of the query server at the current time includes: when the bus ports include N bus ports, querying the candidate number of candidate description information recorded in the registers at the current time by performing the following steps, wherein the registers are configured to record description information of faults occurring in the N bus ports and the bus devices connected to each of the N bus ports, where N is a positive integer; detecting N sets of port description information of the N bus ports, wherein the i-th set of port description information in the N sets of port description information is used to indicate the i-th bus port in the N bus ports, where i is a positive integer less than or equal to N; calling a target interrupt to poll the number of description information corresponding to the N bus ports and the bus devices connected to each of the N bus devices accumulated in the registers at the current time based on the N sets of port description information, to obtain N numbers; and performing a sum-value operation on the N numbers to obtain the candidate number.

[0014] In one exemplary embodiment, detecting whether a first matching condition is satisfied between the number of candidates and a target number threshold to obtain a first detection result includes: detecting whether the number of candidates is greater than or equal to the target number threshold; if the number of candidates is detected to be greater than or equal to the target number threshold, determining that the first detection result indicates that the number of candidates and the target number threshold satisfy the first matching condition; and if the number of candidates is detected to be less than the target number threshold, determining that the first detection result indicates that the number of candidates and the target number threshold do not satisfy the first matching condition.

[0015] In one exemplary embodiment, detecting whether a second matching condition is met between a target duration and a duration threshold between the current time and the initial time to obtain a second detection result includes: detecting whether the target duration is less than or equal to the duration threshold; if the target duration is detected to be less than or equal to the duration threshold, determining that the second detection result indicates that the target duration and the duration threshold meet the second matching condition; and if the target duration is detected to be greater than the duration threshold, determining that the second detection result indicates that the target duration and the duration threshold do not meet the second matching condition.

[0016] In one exemplary embodiment, filtering target description information to be deleted from candidate description information based on a first detection result and a second detection result includes: filtering description information of the target quantity threshold as target description information when the first detection result indicates that a first matching condition is met between the candidate quantity and the target quantity threshold, and the second detection result indicates that a second matching condition is met between the target duration and the duration threshold; and determining candidate description information as target description information when the first detection result indicates that the first matching condition is not met between the candidate quantity and the target quantity threshold, and / or the second detection result indicates that the second matching condition is not met between the target duration and the duration threshold.

[0017] In one exemplary embodiment, determining whether to transmit updated description information stored in a register to a controller in a server includes: determining to transmit updated description information to the controller if a first detection result indicates that a first matching condition is met between the candidate number and the target number threshold, and a second detection result indicates that a second matching condition is met between the target duration and the duration threshold; and determining not to transmit updated description information to the controller if the first detection result indicates that the first matching condition is not met between the candidate number and the target number threshold, and / or the second detection result indicates that the second matching condition is not met between the target duration and the duration threshold.

[0018] In one exemplary embodiment, the controller includes a data table, and the server includes a recording function that transmits update description information to the controller, including: calling the recording function to record update description information into the data table one by one, and transmitting update description information to the controller one by one.

[0019] According to another embodiment of this application, a control device for server fault information is provided. The device is applied to target firmware in a server. The device includes: a first processing module configured to query the number of candidate description information accumulated in a register in the server at the current time, and extract the initial time for storing the first description information in the candidate description information into the register, wherein the initial time is earlier than the current time, and the register is configured to store description information of faults occurring in a bus port on the server and description information of faults occurring in a bus device connected to the bus port; and a first detection module configured to detect whether a first matching condition is met between the number of candidates and a target number threshold, obtain a first detection result, and detect the target number between the current time and the initial time. The second detection result is obtained by determining whether the target duration and the duration threshold satisfy a second matching condition. The target quantity threshold is determined by the target firmware based on the impact of faults in the bus port and bus device on the server's operating performance. The second processing module is configured to filter target description information to be deleted from candidate description information based on the first and second detection results, and determine whether to transmit updated description information stored in the register to the controller in the server. The updated description information is the description information stored in the register after the target description information, and the candidate description information includes the updated description information. The third processing module is configured to delete the target description information and, if it is determined that the updated description information should be transmitted to the controller, transmit the updated description information to the controller.

[0020] According to yet another embodiment of this application, a non-volatile computer-readable storage medium is also provided, wherein a computer program is stored in the non-volatile computer-readable storage medium, and the computer program is configured to perform the steps in any of the above method embodiments when it is run.

[0021] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0022] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0023] This application addresses the issue that bus ports and connected bus devices on a server may experience malfunctions during operation. A target quantity threshold is determined by the target firmware based on the impact of these malfunctions on server performance. This threshold is specifically set for each bus port and connected bus device. The system automatically filters out descriptions to be deleted from stored descriptions in registers based on whether the candidate quantity matches the target quantity threshold and whether the target duration between the current time and the initial time matches the duration threshold. It then determines whether to transmit updated descriptions (e.g., descriptions not selected for deletion) to the controller. If it is determined that updated descriptions should be transmitted to the controller, then the updated descriptions are transmitted. This approach allows maintenance personnel to promptly receive and process descriptions of critical faults when a large number of malfunctions occur, and to clean up stored descriptions in registers in a timely manner, preventing excessive accumulation and impact on server operation. Therefore, it addresses the issue of low server stability and improves overall server stability. Attached Figure Description

[0024] Figure 1 is a hardware structure block diagram of a server device according to an embodiment of the present application of a server fault information control method;

[0025] Figure 2 is a flowchart of a server fault information control method according to an embodiment of this application;

[0026] Figure 3 is a schematic diagram of an optional target interface according to an embodiment of this application;

[0027] Figure 4 is a schematic diagram of threshold clearing and reporting of optional server fault description information according to this embodiment;

[0028] Figure 5 is a flowchart of an optional 24-hour quantitative clearing of PCIe CE errors according to an embodiment of this application;

[0029] Figure 6 is a structural block diagram of a server fault information control device according to an embodiment of this application. Detailed Implementation

[0030] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0032] The terms used in the embodiments of this application are explained as follows:

[0033] IPMI: Intelligent Platform Management Interface. IPMI can intelligently monitor, control, and automatically report the operational status of a large number of servers across different operating systems, firmware, and hardware platforms, thereby reducing server system costs.

[0034] BMC: Baseboard Management Controller.

[0035] BIOS: Basic Input and Output System, is a set of programs embedded in a ROM (Read-Only Memory) chip on the computer's motherboard. It stores the computer's most important basic input and output programs, power-on self-test programs, and system startup programs. It also contains specific information for reading and writing system settings.

[0036] AMD: Advanced Micro Devices.

[0037] Intel: Intel Corporation.

[0038] PCIe: PCI Express, a high-speed serial computer expansion bus standard, is a technically fast peripheral component interconnect express, but is usually abbreviated as PCIe or PCI-E. It is a standard type of connection for internal computer devices.

[0039] CE: Corrected Error. An error can be corrected.

[0040] RAS stands for Reliability, Availability, and Serviceability. As a whole, RAS ensures the entire system operates reliably for as long as possible without going offline, and possesses sufficiently robust fault tolerance mechanisms. This is an indispensable component for application environments such as large data centers, network centers like stock exchanges, telecommunications data centers, and bank database centers.

[0041] APEI: Advanced Platform Error Interfaces, a unified and efficient interface used to transmit error information between hardware and upper-layer software.

[0042] RTC: Real-time Clock / Calendar Chip. A high-performance, low-power real-time clock circuit with RAM (Random Access Memory) that can keep track of year, month, day, day of the week, hour, minute, and second, and has leap year compensation function.

[0043] FW First: Firmware First, meaning that errors are handled by the BIOS first.

[0044] OS First: Operating System First, meaning that errors are handled by the OS first.

[0045] Runtime protocols: In BIOS, runtime protocols refer to protocols that can be invoked when the machine is in the OS phase.

[0046] CE and UCE errors: In the application scope of PCIe RAS, the most common types include: Error Detection, Error Correction, Error Reporting, and Hot-plug Support. Error Detection typically uses parity checks and cyclic redundancy checks for rapid error detection. Error Correction (ECC) is a set of extra bits added during data storage or transmission to detect and correct certain types of errors. When an error is detected, if the location of the error is known, ECC information can be used to automatically correct the error without retransmitting the data. Error Reporting commonly involves two types of errors: UCE and CE. UCE stands for Uncorrectable Error, which refers to errors that cannot be automatically corrected by the hardware. When an uncorrectable error is detected, the hardware reports it to the operating system or device driver via AER. UCE may be caused by hardware failure, severe signal attenuation, etc. UCE usually requires user intervention to resolve, such as replacing the faulty component or restarting the system. CE stands for Correctable Error, referring to errors that can be automatically corrected at the hardware level. When a correctable error is detected, the hardware will attempt to fix it and report it to the operating system or device driver via AER. CE is usually caused by noise, transient interference, etc.

[0047] The methods and embodiments provided in this application can be executed in a server device or a similar computing device. Taking a server device as an example, FIG1 is a hardware structure block diagram of a server device according to an embodiment of this application for controlling server fault information. As shown in FIG1, the server device may include one or more (only one is shown in FIG1) processors 102 (processors 102 may include, but are not limited to, microprocessors MCUs or programmable logic devices FPGAs, etc.) and a memory 104 configured to store data. The server device may also include a transmission device 106 configured for communication functions and an input / output device 108. It will be understood by those skilled in the art that the structure shown in FIG1 is only illustrative and does not limit the structure of the server device. For example, the server device may also include more or fewer components than shown in FIG1, or have a different configuration than shown in FIG1.

[0048] The memory 104 can be configured to store computer programs, such as application software programs and modules, like the computer program corresponding to the server fault information control method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the server device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0049] The transmission device 106 is configured to receive or transmit data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module configured to communicate wirelessly with the Internet.

[0050] This embodiment provides a method for controlling server fault information. Figure 2 is a flowchart of the method for controlling server fault information according to an embodiment of this application. The method is applied to the target firmware in the server. As shown in Figure 2, the process includes the following steps:

[0051] Step S202: Query the number of candidate description information accumulated in the register of the server at the current time, and extract the initial time of storing the first description information in the candidate description information into the register. The initial time is earlier than the current time. The register is configured to store the description information of the fault that occurred on the bus port of the server and the description information of the fault that occurred on the bus device connected to the bus port.

[0052] Step S204: Detect whether the candidate number and the target number threshold meet the first matching condition to obtain the first detection result, and detect whether the target duration between the current time and the initial time meets the second matching condition to obtain the second detection result. The target number threshold is determined by the target firmware based on the degree of impact of the failure of the bus port and bus device on the server's operating performance.

[0053] Step S206: Based on the first detection result and the second detection result, filter the target description information to be deleted from the candidate description information, and determine whether to transmit the updated description information stored in the register to the controller in the server. The updated description information is the description information stored in the register after the target description information, and the candidate description information includes the updated description information.

[0054] Step S208: Delete the target description information, and if it is determined that updated description information will be transmitted to the controller, transmit updated description information to the controller.

[0055] Through the above steps, the bus ports on the server and the bus devices connected to them may experience failures during operation. The target quantity threshold is determined by the target firmware based on the impact of these failures on the server's performance. This threshold is understood to be a separate threshold set for each bus port and its connected devices. Based on whether the candidate quantity meets the first matching condition and the target quantity threshold, and whether the target duration between the current time and the initial time meets the second matching condition, description information to be deleted is automatically filtered from the description information stored in the register. It is then determined whether to transmit updated description information (e.g., description information not selected for deletion) to the controller. If it is determined that updated description information should be transmitted to the controller, then it is transmitted. In this way, when a large number of failures occur, maintenance personnel can promptly receive and process description information of the failures requiring attention, and can also promptly clean up the description information stored in the register, preventing excessive accumulation of description information that could affect server operation. Therefore, this addresses the issue of low server stability and improves overall server stability.

[0056] In the technical solution provided in step S202 above, the bus port may include, but is not limited to, a port that conforms to the target bus protocol, and the bus device may include, but is not limited to, the same bus protocol as the bus port. For example, the target bus protocol may include, but is not limited to, the PCIe protocol, the I2C (Inter-Integrated Circuit, two-wire serial bus) protocol, etc. Taking the PCIe protocol as an example, the bus port may include, but is not limited to, a PCIe port, and the bus device may include, but is not limited to, a PCIe device, such as a network card.

[0057] Optionally, in this embodiment, the target firmware may be configured, but is not limited to, to manage and control description information of faults occurring on the bus port on the server and description information of faults generated by the bus devices connected to the bus port. The target firmware may include, but is not limited to, a system management module, such as a BIOS.

[0058] Optionally, in this embodiment, the fault description information may be, but is not limited to, indicating the type of fault that occurred in the bus port and bus device, the bus port that failed, the bus device that failed, and the error message that occurred, etc.

[0059] Optionally, in this embodiment, the description information of the faults occurring in the bus port and the faults occurring in the bus devices connected to the bus port may include, but are not limited to, faults such as CE and UCE. One fault may correspond to one description information. Taking, but is not limited to, the target bus protocol including the PCIe protocol, the bus port including the PCIe port, and the bus device including the PCIe device as an example, the description information of the faults occurring in the bus port and the faults occurring in the bus devices connected to the bus port may include, but is not limited to, the description information of the CE faults occurring in the PCIe port and the PCIe device.

[0060] Optionally, in this embodiment, the initial time can be, but is not limited to, representing the time when the first description information in the candidate description information is stored in the register. The current time and the initial time can be, but are not limited to, represented by year, month, day, hour, minute, second, etc. For example, the initial time can be, but is not limited to, represented by a series of specific time parameters, such as being stored and used in the form of T0D (day), T0H (hour), T0M (minute), and T0S (second). The current time can be, but is not limited to, represented by a series of specific time parameters, such as being stored and used in the form of T2D (day), T2H (hour), T2M (minute), and T2S (second).

[0061] In one exemplary embodiment, the server includes a clock chip connected to target firmware. The clock chip is configured to record the current time of the server, which may be done, but is not limited to, before querying the number of candidate descriptive information accumulated in the registers of the server at the current time, by: detecting a target address of the clock chip, wherein the target address is used to extract the time recorded in the clock chip; and extracting the time recorded in the clock chip as the current time by accessing the target address.

[0062] Optionally, in this embodiment, the clock chip may be configured, but is not limited to, to record the current time of the server in real time. The target address may be, but is not limited to, an access address that accesses the current time of the server recorded in the clock chip. The target address may be accessed through a register or storage unit inside the clock chip that is configured to store and update time data, and the current time of the server may be extracted to obtain the current time.

[0063] Optionally, in this embodiment, the clock chip can, but is not limited to, continue to accurately track time when the system is powered off using a backup power supply, thereby achieving real-time time recording.

[0064] Through the embodiments of this application, the current time is read from the target address provided by the clock chip (RTC) to identify when the duration threshold has been reached, thereby realizing the periodic quantitative cleaning of fault description information and improving the operation and maintenance efficiency of the server.

[0065] In one exemplary embodiment, the number of candidate descriptions accumulated in the registers of the server at the current time can be queried in the following manner, but not limited to: when the bus ports include N bus ports, the following steps are performed to query the number of candidate descriptions recorded in the registers at the current time, wherein the registers are configured to record descriptions of faults occurring in the N bus ports and the bus devices connected to each of the N bus ports, where N is a positive integer: detecting N sets of port descriptions for the N bus ports, wherein the i-th set of port descriptions in the N sets of port descriptions is used to indicate the i-th bus port in the N bus ports, where i is a positive integer less than or equal to N; calling the target interrupt to poll the number of descriptions corresponding to the N bus ports and the bus devices connected to each of the N bus ports accumulated in the registers at the current time based on the N sets of port descriptions, to obtain N counts; performing a sum-value operation on the N counts to obtain the candidate count.

[0066] Optionally, in this embodiment, the port description information may include, but is not limited to, the identifier of the bus port. For example, it may include, but is not limited to, the BDF (Bus Number, Device Number, and Function Number, abbreviated as BDF) of the bus port, etc.

[0067] Optionally, in this embodiment, the i-th quantity among the N quantities may be, but is not limited to, 0, or the i-th quantity may be greater than 0. The i-th quantity is the number of description information of the faults of the i-th bus port and the i-th bus device connected to it stored in the register, and the N bus ports include the i-th bus port.

[0068] Optionally, in this embodiment, the target interrupt may be used, but is not limited to, to issue an interrupt signal, such as an SMI interrupt, when a fault is detected (e.g., a PCIe CE fault). For example, when the BIOS detects a PCIe CE error in the AER register via the SMM protocol, it may, but is not limited to, trigger an SMI interrupt.

[0069] In the technical solution provided in step S204 above, the degree of influence can be calculated, but is not limited to, according to the following formula (1):

[0070] Where R represents the degree of impact, S1 represents the server's operating performance when the bus port and the bus device connected to the bus port are not faulty, and S2 represents the server's operating performance after the bus port and the bus device connected to the bus port have failed.

[0071] As an alternative example, server performance can be determined, but is not limited to, by server processing speed, the accuracy of server data transmission, the speed at which server responds to requests, and so on.

[0072] Optionally, in this embodiment, the threshold values ​​corresponding to the degree of impact of bus port and bus device failures on server performance may be different. For example, when the degree of impact of bus port and bus device failures on server performance is 20%, the corresponding threshold value is 200, and when the degree of impact of bus port and bus device failures on server performance is 40%, the corresponding threshold value is 600.

[0073] Optionally, in this embodiment, the first detection result may include, but is not limited to, the first matching condition being met between the number of candidates and the target number threshold, or the first matching condition not being met between the number of candidates and the target number threshold; the second detection result may include, but is not limited to, the second matching condition being met between the target duration and the duration threshold between the current time and the initial time, or the second matching condition not being met between the target duration and the duration threshold between the current time and the initial time.

[0074] In one exemplary embodiment, the server includes an adjustment device connected to target firmware, which may, but is not limited to, obtain a first detection result before detecting whether a first matching condition is met between the candidate quantity and a target quantity threshold by: receiving a target adjustment request, wherein the target adjustment request is used to request that the quantity threshold corresponding to the quantity of description information stored in the register be adjusted to the target quantity threshold; and sending the target adjustment request to the adjustment device, wherein the adjustment device is configured to execute the target adjustment request.

[0075] Optionally, in this embodiment, a fault diagnosis driver, such as a RAS fault diagnosis driver, may be deployed on the adjustment device, and the adjustment device may execute the target adjustment request through the fault diagnosis driver, but is not limited to that.

[0076] Optionally, in this embodiment, the target adjustment request may be, but is not limited to, requesting that the quantity threshold corresponding to the quantity of description information already stored in the register be adjusted from the initial quantity threshold to the target quantity threshold. For example, the initial quantity threshold may be greater than the target quantity threshold, or the initial quantity threshold may be less than the target quantity threshold.

[0077] In one exemplary embodiment, the server includes a target interface, which may, but is not limited to, receive a target adjustment request in the following manner: detecting a first editing operation performed on a first field on the target interface, and detecting a second editing operation performed on a second field on the target interface, wherein the first field is used to set the degree of impact of bus port and bus device failures on the server's operating performance, the second field is used to set a threshold value for the number of descriptive information stored in the register corresponding to the degree of impact of bus port and bus device failures on the server's operating performance, the first editing operation is used to set the degree of impact of bus port and bus device failures on the server's operating performance to a target degree of impact, and the second editing operation is used to adjust the threshold value for the number of descriptive information stored in the register corresponding to the target degree of impact to a target threshold value; generating a target adjustment request corresponding to the first editing operation and the second editing operation, and transmitting the target adjustment request to the target firmware.

[0078] Optionally, in this embodiment, there may be, but is not limited to, a corresponding relationship between the degree of influence and the quantity threshold. Users may, but are not limited to, edit the degree of influence and the quantity threshold corresponding to the degree of influence on the target interface. For example, by default, the quantity threshold corresponding to a degree of influence of 10% is 100, and the quantity threshold corresponding to a degree of influence of 20% is 300. Users may, but are not limited to, adjust the threshold corresponding to a degree of influence of 10% to 150.

[0079] Figure 3 is a schematic diagram of an optional target interface according to an embodiment of this application. As shown in Figure 3, the user can perform a first editing operation on a first field (e.g., degree of influence) on the target interface. For example, a list is displayed on the target interface showing the degree of influence that the user can select, such as 10%, 20%, 30%...100%, etc. The user can perform a second editing operation on a second field (e.g., quantity threshold) on the target interface. For example, a list is displayed on the target interface showing the quantity thresholds corresponding to the degree of influence that the user can select, such as 100, 300, 500, etc.

[0080] For example, a user can, but is not limited to, edit the impact level to 10% and the quantity threshold to 100, which corresponds to 10%. In such a case, after clicking the confirmation operation on the target interface, a target adjustment request corresponding to the first and second editing operations is generated and transmitted to the target firmware. It is understood that the target adjustment request can, but is not limited to, requesting that the quantity threshold be set to the quantity threshold of 100, which corresponds to the impact level of 10%.

[0081] It should be noted that, due to the correlation between the degree of impact and the quantity threshold, when the degree of impact is edited, the quantity threshold will automatically adjust according to this correlation to match the new degree of impact setting. The target interface also allows users to perform independent editing operations on the first field (degree of impact) and the second field (quantity threshold) according to specific needs. For example, a user can edit the degree of impact on the target interface, resulting in a degree of impact of 10% and a quantity threshold of 100. Users can also adjust the quantity threshold accordingly, such as increasing or decreasing it from 100.

[0082] Through the embodiments of this application, selectable PCIe CE over-limit thresholds are available on the Setup (equivalent to the target interface). Thresholds can be set periodically and quantitatively according to customer needs. By allowing users to edit the impact level and quantity thresholds, maintenance personnel can adjust the thresholds based on historical error data and the current system status, which can avoid unnecessary alarms and error handling, save maintenance time and costs, and improve the system's configuration flexibility and intelligence level.

[0083] In one exemplary embodiment, after sending the target adjustment request to the adjustment device, the following steps may be taken, but are not limited to: defining a threshold identifier in a configuration file, wherein the value of the threshold identifier is used to identify the quantity threshold corresponding to the quantity of descriptive information stored in the register; and calling an entry function to adjust the value of the threshold identifier to the target quantity threshold.

[0084] Optionally, in this embodiment, the value of the threshold identifier may, but is not limited to, be the same as the quantity threshold corresponding to the quantity of descriptive information stored in the register. When the quantity threshold is updated, the value of the threshold identifier may, but is not limited to, be updated by calling the entry function. It is understood that the value of the threshold identifier may change dynamically.

[0085] In one exemplary embodiment, a first detection result may be obtained by detecting, but is not limited to, whether the candidate number and the target number threshold satisfy a first matching condition in the following manner: detecting whether the candidate number is greater than or equal to the target number threshold; if the candidate number is detected to be greater than or equal to the target number threshold, determining that the first detection result is used to indicate that the candidate number and the target number threshold satisfy the first matching condition; if the candidate number is detected to be less than the target number threshold, determining that the first detection result is used to indicate that the candidate number and the target number threshold do not satisfy the first matching condition.

[0086] Optionally, in this embodiment, the target quantity threshold may change dynamically in different detection cycles. The duration of the detection cycle is equal to the duration threshold. For example, the target quantity threshold in the first detection cycle is less than the target quantity threshold in the second detection cycle. In such a case, the same candidate quantity may satisfy the first matching condition with the target quantity threshold in the first detection cycle, but the same candidate quantity may not satisfy the first matching condition with the target quantity threshold in the second detection cycle.

[0087] In one exemplary embodiment, a second detection result may be obtained by detecting, but is not limited to, whether the target duration between the current time and the initial time and the duration threshold satisfy a second matching condition in the following manner: detecting whether the target duration is less than or equal to the duration threshold; if the target duration is detected to be less than or equal to the duration threshold, determining that the second detection result indicates that the target duration and the duration threshold satisfy the second matching condition; if the target duration is detected to be greater than the duration threshold, determining that the second detection result indicates that the target duration and the duration threshold do not satisfy the second matching condition.

[0088] Optionally, in this embodiment, the current time and the time of the first recorded fault description information can be standardized using the following formula (2), that is, the duration between the current time and the initial time can be converted into hours, and then the duration between the two times can be calculated by subtraction, with the unit being hours. T=(T2D*24+T2H)-(TOD*24+TOH) Formula (2)

[0089] Where T represents the duration between the current time and the initial time, T2D represents the number of days in the current time, T2H represents the number of hours in the current time, T0D represents the number of days in the time when the fault description information was first recorded, and T0H represents the number of hours in the time when the fault description information was first recorded.

[0090] Optionally, in this embodiment, when calculating the duration between the current time and the initial time, the duration between the current time and the initial time can also be converted into minutes or seconds for calculation, but is not limited to.

[0091] By converting the current time and the time when the fault was first recorded to the same time scale, the time between the current time and the time when the fault description information was first recorded can be accurately calculated, ensuring the timeliness and effectiveness of the management of fault description information.

[0092] In the technical solution provided in step S206 above, description information of faults generated by bus ports and bus devices can be written into registers in real time, but is not limited to. At the current time, the number of updated description information stored in the register can be, but is not limited to, 0, or the number of updated description information stored in the register can be greater than 0.

[0093] In one exemplary embodiment, target description information to be deleted can be filtered from candidate description information based on a first detection result and a second detection result in the following manner: when the first detection result indicates that a first matching condition is met between the number of candidates and the target number threshold, and the second detection result indicates that a second matching condition is met between the target duration and the duration threshold, the description information of the target number threshold is filtered from the candidate description information as target description information; when the first detection result indicates that the first matching condition is not met between the number of candidates and the target number threshold, and / or the second detection result indicates that the second matching condition is not met between the target duration and the duration threshold, the candidate description information is determined as target description information.

[0094] Optionally, in this embodiment, if the first detection result satisfies the first matching condition and the second detection result satisfies the second matching condition, it can be indicated that the number of fault descriptions accumulated within the time threshold has exceeded the target number threshold. In this case, the descriptions of the target number threshold are selected from the candidate descriptions as the target descriptions.

[0095] If the first detection result does not meet the first matching condition, and / or the second detection result does not meet the second matching condition, it can be indicated that the number of fault description information accumulated within the time threshold has not yet reached the target number threshold. In such a case, all candidate description information is determined as target description information.

[0096] In one exemplary embodiment, it may be determined, but not limited to, whether to transmit the updated description information stored in the register to the controller in the server in the following ways: if a first detection result indicates that a first matching condition is met between the candidate number and the target number threshold, and a second detection result indicates that a second matching condition is met between the target duration and the duration threshold, then it is determined whether to transmit the updated description information to the controller; if a first detection result indicates that the first matching condition is not met between the candidate number and the target number threshold, and / or a second detection result indicates that the second matching condition is not met between the target duration and the duration threshold, then it is determined whether to transmit the updated description information to the controller.

[0097] Optionally, in this embodiment, it is possible, but not limited to, determining to transmit updated description information to the controller when the first detection result indicates that the candidate number and the target number threshold meet a first matching condition, and the second detection result indicates that the target duration and the duration threshold meet a second matching condition. Figure 4 is a schematic diagram of threshold cleaning and reporting of description information for an optional server fault according to this embodiment. As shown in Figure 4, it is possible, but not limited to, taking a number threshold including 200, a candidate number of candidate description information recorded in the register at the current time being 400, a duration threshold of 24 hours, and the target firmware including BIOS as an example. In this case, the candidate number is greater than the number threshold, and the duration between the current time and the initial time is less than or equal to 24 hours. In this case, the BIOS will filter out description information 1 to description information 200 as description information to be deleted, and determine to transmit description information 201 to description information 400 stored in the register to the controller. It can be understood that the updated description information includes description information 201 to description information 400.

[0098] In the technical solution provided in step S208 above, if it is determined that updated description information will be transmitted to the controller, the description information of the threshold number stored in the register is deleted, and the description information exceeding the threshold number is immediately transmitted to the controller; if it is determined that updated description information will not be transmitted to the controller, all description information stored in the register is deleted.

[0099] For example, if the duration threshold is 24 hours and the quantity threshold is 200, and 300 error messages have accumulated in the register by the 18th hour, then the 201st to 300th messages in the register will be immediately transmitted to the controller, and the 1st to 200th messages will be deleted.

[0100] In related technologies, current server PCIe CE processing requires a CE filtering mechanism to allow customers to dynamically and selectively detect CE reporting. Simply preventing PCIe CE reporting would prevent customers from perceiving risks in certain hardware links. The technical solution in this application can be used, but is not limited to, server products based on the AMD platform x86 architecture (or other server products; this application makes no limitation on this). By using RTC clock readings and setting thresholds periodically and quantitatively according to customer needs, it avoids the risk of machine downtime caused by a large accumulation of PCIe CEs and CE storms. It achieves dynamic processing of the number of PCIe CEs within a certain period, greatly reducing the risk of CE storms, selectively and promptly notifying maintenance personnel of serious machine anomalies, and reducing ineffective investment by customer maintenance personnel.

[0101] In one exemplary embodiment, the controller includes a data table, and the server includes a recording function that can, but is not limited to, transmit update description information to the controller by calling the recording function to record update description information into the data table one by one, and transmitting update description information to the controller one by one.

[0102] Optionally, in this embodiment, the controller may, but is not limited to, detect faults in the bus port and bus device based on the received updated description information, and promptly repair the corresponding faults. The controller may, but is not limited to, include BMC, OS, etc.

[0103] Optionally, in this embodiment, the above method further includes: when the number of candidate description information stored in the register at the current time is greater than or equal to a number threshold, detecting whether the number of candidate description information stored in the register at the current time is greater than or equal to an upper limit value, wherein the upper limit value is greater than the target number threshold, and the upper limit value is the threshold corresponding to the case where the impact of the failure of the bus port and the bus device connected to the bus port on the operating performance of the server is greater than or equal to the impact threshold.

[0104] If the number of candidate description information is greater than or equal to the upper limit, the description information stored in the register after the target description information is determined as the updated description information, the updated description information is transmitted to the controller, and the description information of the upper limit value is deleted.

[0105] According to the embodiments of this application, if the number of candidate description information stored in the register at the current time is greater than or equal to the upper limit, it can be indicated that a CE storm may have occurred, resulting in a large number of faults in a very short time. In such cases, all description information stored in the register after the target description information can be directly transmitted to the controller, avoiding the waste of time to repeatedly check whether reporting is required, and avoiding the waste of time to filter the error information that needs to be reported from the candidate description information. Instead, it directly reports when a CE storm is detected, improving the efficiency of reporting error information to the controller and making it easier for maintenance personnel to quickly handle the CE storm.

[0106] To better understand the control process of the server fault information control method in the embodiments of this application, the control process of the server fault information control method in the embodiments of this application will be explained and described below in conjunction with optional embodiments, which may be applied to, but is not limited to, the embodiments of this application.

[0107] If the CPU can write PCIe CEs into the AER register normally, the technical solution of this application can be implemented by, but is not limited to, the following steps: 1) The BIOS (equivalent to the target firmware) needs to set a threshold for the number of PCIe CEs to indicate when the accumulated CEs reach the threshold within 24 hours and report the PCIe CEs to the BMC and OS; 2) Read the current time according to the clock chip (RTC) address to indicate when the 24-hour duration has been reached; 3) Execute a mechanism to clean up PCIe CEs on time according to the time obtained from the RTC.

[0108] Figure 5 is a flowchart of an optional 24-hour quantitative clearing of PCIe CE errors according to an embodiment of this application. As shown in Figure 5, the explanation can be made by taking, but is not limited to, the target bus protocol including the PCIe protocol, the bus port including the PCIe port, the bus device including the PCIe device, and the fault including the CE fault.

[0109] When a PCIe CE occurs, the BIOS (equivalent to the target firmware) calls the AMD platform's Ras fault diagnosis driver (equivalent to a device tuner) to handle the PCIe CE fault. At this point, we need to set a threshold in the Setup (equivalent to the target interface) to indicate how many CE errors can be cleared within 24 hours (equivalent to a duration threshold, which could also be 12 hours or 30 minutes, etc., this application does not limit this). We need to define the threshold identifier displayed in the Setup file in the sd file (equivalent to the configuration file), and then initialize and assign a value to this threshold at the PCIe CE processing entry function. During this process, we can specify how many CE errors should be missed within 24 hours according to our needs, and associate this threshold with the Setup to achieve selectivity in the number of errors cleared.

[0110] Next, the BIOS calls the SMI interrupt (equivalent to a target interrupt) to handle CE errors in the AER register. When the CE error handling method in the AER is FW First, meaning the CE error is not triggered by the device driver layer, the BIOS will use the SMM protocol to call the SMI interrupt to poll the BDF of the PCIe Root port to confirm whether there are errors in the corresponding PCIe port. When the number of PCIe CE errors exceeds the threshold, the BIOS will pass the errors to the OS and synchronize them to the BMC via SMI. When a PCIe CE error occurs, the BIOS will store it in a list and accumulate the count using the Runtime protocol, thus confirming how many PCIe CE errors are recorded in the CPU's AER register under the current configuration. When the BIOS starts polling for PCIe CE errors, it records the current time as T0D, T0H, T0M, and T0S. Then, when the PCIe CE error exceeds the threshold, the BIOS reads the current time based on the clock chip address, accurate to the day, hour, minute, and second of the month, and records it as T2D, T2H, T2M, and T2S. The BIOS calculates whether the time (T2D*24+T2H)-(T0D*24+T0H) exceeds 24 hours to determine whether the CE error needs to be reported.

[0111] After polling all PCIe CEs, the BIOS triggers the CE logging function. If the value of (T2D*24+T2H)-(T0D*24+T0H) does not exceed 24 hours, the BIOS will record each CE exceeding the threshold into the OS's BERT table (equivalent to a data table) and report each record to the BMC. The BMC records this in the log system and then subtracts the threshold number of PCIe CE errors from the AER register. The remaining PCIe CEs proceed to the next round of accumulation. If the value of (T2D*24+T2H)-(T0D*24+T0H) exceeds 24 hours, the BIOS will not report CE errors to the OS and BMC. It will then clear the PCIe CE errors in the AER register and begin the next round of CE accumulation. After this 24-hour cycle is completed, the BIOS assigns T2D, T2H, T2M, and T2S to T0D, T0H, T0M, and T0S, starting a new cycle count.

[0112] This application primarily optimizes the existing PCIe CE error handling mechanism by utilizing the RTC clock (equivalent to a clock chip) read time confirmation interval to periodically and quantitatively clear PCIe CE errors, bringing the following three benefits:

[0113] (1) When a large number of PCIe CE errors occur on the machine, they can be reported in a timely manner within a certain period of time, so that maintenance personnel can be aware of them and investigate the cause of the errors in a timely manner, which greatly reduces the risk of CE storm;

[0114] (2) A small number of PCIe CEs is a normal phenomenon. Not all PCIe CEs need to be perceived by the customer. This technical solution allows customers to selectively perceive more serious errors and promptly check whether the machine link is abnormal. This can save a lot of investment from maintenance personnel and reduce the risk of serious machine abnormalities.

[0115] (3) The introduction of this technical solution allows customers to selectively perceive PCIe CE error reporting within a day and dynamically and intelligently control PCIe CE reporting, which greatly helps to improve the competitiveness of the machine.

[0116] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0117] This embodiment also provides a control device for server fault information, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0118] Figure 6 is a structural block diagram of a server fault information control device according to an embodiment of the present application. As shown in Figure 6, the device is applied to the target firmware in the server, and the device includes:

[0119] The first processing module 602 is configured to query the number of candidate description information accumulated in the register of the server at the current time, and extract the initial time of storing the first description information in the candidate description information into the register, wherein the initial time is earlier than the current time, and the register is configured to store the description information of the fault that occurred on the bus port of the server and the description information of the fault that occurred on the bus device connected to the bus port.

[0120] The first detection module 604 is configured to detect whether the number of candidates and the target number threshold meet a first matching condition, and obtain a first detection result; and to detect whether the target duration between the current time and the initial time and the duration threshold meet a second matching condition, and obtain a second detection result. The target number threshold is determined by the target firmware based on the degree of impact of the failure of the bus port and bus device on the operating performance of the server.

[0121] The second processing module 606 is configured to filter target description information to be deleted from candidate description information based on the first detection result and the second detection result, and determine whether to transmit the updated description information stored in the register to the controller in the server, wherein the updated description information is the description information stored in the register after the target description information, and the candidate description information includes the updated description information.

[0122] The third processing module 608 is configured to delete the target description information and, if it is determined that updated description information will be transmitted to the controller, transmit updated description information to the controller.

[0123] Through the aforementioned device, the bus ports on the server and the bus devices connected to them may experience failures during operation. The target quantity threshold is determined by the target firmware based on the impact of these failures on the server's performance. This threshold is understood to be a separate threshold set for each bus port and its connected devices. Based on whether the candidate quantity meets the first matching condition and the target quantity threshold, and whether the target duration between the current time and the initial time meets the second matching condition, description information to be deleted is automatically filtered from the description information stored in the register. It is then determined whether to transmit updated description information (e.g., description information not selected for deletion) to the controller. If it is determined that updated description information should be transmitted to the controller, then the updated description information is transmitted. In this way, when a large number of failures occur, maintenance personnel can promptly receive and process description information of the failures requiring attention, and can also promptly clean up the description information stored in the register, preventing excessive accumulation of description information that could affect server operation. Therefore, this addresses the problem of low server stability and improves overall server stability.

[0124] In one exemplary embodiment, the server includes an adjustment device connected to the target firmware, and the apparatus further includes:

[0125] The receiving module is configured to receive a target adjustment request before obtaining a first detection result, based on whether a first matching condition is met between the number of candidate detectors and the target number threshold. The target adjustment request is used to request that the number threshold corresponding to the number of description information stored in the register be adjusted to the target number threshold.

[0126] The sending module is configured to send a target adjustment request to the adjustment device, wherein the adjustment device is configured to execute the target adjustment request.

[0127] In one exemplary embodiment, the server includes a target interface and a receiving module, comprising:

[0128] The first detection unit is configured to detect a first editing operation performed on a first field on the target interface and a second editing operation performed on a second field on the target interface. The first field is used to set the degree of impact of bus port and bus device failures on the server's operating performance. The second field is used to set a threshold number of descriptive information stored in the register corresponding to the degree of impact of bus port and bus device failures on the server's operating performance. The first editing operation is used to set the degree of impact of bus port and bus device failures on the server's operating performance to a target degree of impact. The second editing operation is used to adjust the threshold number of descriptive information stored in the register corresponding to the target degree of impact to the target threshold number.

[0129] The first processing unit is configured to generate target adjustment requests corresponding to the first and second editing operations, and transmit the target adjustment requests to the target firmware.

[0130] In one exemplary embodiment, the apparatus further includes:

[0131] The definition module is configured to define a threshold identifier in the configuration file after sending a target adjustment request to the adjustment device. The value of the threshold identifier is used to identify the quantity threshold corresponding to the quantity of descriptive information stored in the register.

[0132] The adjustment module is configured to call the entry function to adjust the value of the threshold identifier to the target quantity threshold.

[0133] In one exemplary embodiment, the server includes a clock chip connected to target firmware, the clock chip being configured to record the current time of the server, and the device includes:

[0134] The second detection module is configured to detect the target address of the clock chip before querying the number of candidate description information accumulated in the register of the query server at the current time, wherein the target address is used to extract the time recorded in the clock chip;

[0135] The extraction module is configured to access the target address and extract the time recorded in the clock chip as the current time.

[0136] In one exemplary embodiment, the first processing module includes:

[0137] When there are N bus ports, the following steps are performed to query the number of candidate descriptions recorded in the register at the current time, where the register is configured to record descriptions of faults occurring in the N bus ports and the bus devices connected to each of the N bus ports, where N is a positive integer:

[0138] The second detection unit is configured to detect N sets of port description information for N bus ports, wherein the i-th set of port description information in the N sets of port description information is used to indicate the i-th bus port among the N bus ports, and i is a positive integer less than or equal to N;

[0139] The calling unit is configured to call the target interrupt based on the number of N bus ports and the number of description information corresponding to the bus devices connected to each of the N bus ports that have been accumulated in the current time, according to the N sets of port description information polling registers;

[0140] The execution unit is configured to perform a sum-value operation on N quantities to obtain candidate quantities.

[0141] In one exemplary embodiment, the first detection module includes:

[0142] The third detection unit is configured to detect whether the number of candidates is greater than or equal to the target number threshold.

[0143] The first determining unit is configured to, when the number of candidates detected is greater than or equal to a target number threshold, determine a first detection result to indicate that the number of candidates and the target number threshold meet a first matching condition; and when the number of candidates detected is less than the target number threshold, determine a first detection result to indicate that the number of candidates and the target number threshold do not meet the first matching condition.

[0144] In one exemplary embodiment, the first detection module further includes:

[0145] The fourth detection unit is configured to detect whether the target duration is less than or equal to a duration threshold;

[0146] The second determining unit is configured to, when the target duration is detected to be less than or equal to a duration threshold, determine a second detection result to indicate that the target duration and the duration threshold meet a second matching condition; and when the target duration is detected to be greater than the duration threshold, determine a second detection result to indicate that the target duration and the duration threshold do not meet the second matching condition.

[0147] In one exemplary embodiment, the second processing module includes:

[0148] The filtering unit is configured to filter the description information of the target number threshold from the candidate description information as the target description information when the first detection result indicates that the candidate number and the target number threshold meet a first matching condition and the second detection result indicates that the target duration and the duration threshold meet a second matching condition.

[0149] The third determining unit is configured to determine the candidate description information as the target description information when the first detection result indicates that the candidate number does not meet the first matching condition with the target number threshold, and / or the second detection result indicates that the target duration does not meet the second matching condition with the duration threshold.

[0150] In one exemplary embodiment, the second processing module further includes:

[0151] The fourth determining unit is configured to determine to transmit updated description information to the controller when the first detection result indicates that the candidate number and the target number threshold meet a first matching condition and the second detection result indicates that the target duration and the duration threshold meet a second matching condition.

[0152] The fifth determining unit is configured to determine not to transmit update description information to the controller when the first detection result indicates that the candidate number does not meet the first matching condition and / or the second detection result indicates that the target duration does not meet the second matching condition.

[0153] In one exemplary embodiment, the controller includes a data table, the server includes a recording function, and a third processing module includes:

[0154] The second processing unit is configured to call the recording function to record the updated description information into the data table one by one, and transmit the updated description information to the controller one by one.

[0155] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0156] Embodiments of this application also provide a non-volatile computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when it is run.

[0157] In one exemplary embodiment, the aforementioned non-volatile computer-readable storage medium may include, but is not limited to, various non-volatile readable storage media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0158] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0159] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0160] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0161] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0162] Embodiments of this application also provide a computer program that includes computer instructions stored in a non-volatile computer-readable storage medium; a processor of a computer device reads the computer instructions from the non-volatile computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.

[0163] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0164] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0165] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A method for controlling server fault information, characterized in that, The method, applied to target firmware in a server, includes: The system queries the number of candidate descriptions accumulated in the registers of the query server at the current time, and extracts the initial time for storing the first description in the candidate descriptions into the register, wherein the initial time is earlier than the current time. The register is configured to store descriptions of faults occurring on the bus port of the server and descriptions of faults occurring in the bus devices connected to the bus port. The system detects whether the candidate number and the target number threshold satisfy a first matching condition to obtain a first detection result, and detects whether the target duration between the current time and the initial time and the duration threshold satisfy a second matching condition to obtain a second detection result. The target number threshold is determined by the target firmware based on the degree of impact of the faults of the bus port and the bus device on the operating performance of the server. Based on the first detection result and the second detection result, target description information to be deleted is filtered from the candidate description information, and it is determined whether to transmit the updated description information stored in the register to the controller in the server, wherein the updated description information is the description information stored in the register after the target description information, and the candidate description information includes the updated description information; Delete the target description information, and if it is determined that the updated description information should be transmitted to the controller, transmit the updated description information to the controller.

2. The method according to claim 1, characterized in that, The server includes an adjustment device connected to the target firmware. Before detecting whether a first matching condition is met between the candidate number and the target number threshold, and obtaining a first detection result, the method further includes: Receive a target adjustment request, wherein the target adjustment request is used to request that the quantity threshold corresponding to the quantity of description information stored in the register be adjusted to the target quantity threshold; The target adjustment request is sent to the adjustment device, wherein the adjustment device is configured to execute the target adjustment request.

3. The method according to claim 2, characterized in that, The server includes a target interface, and receiving the target adjustment request includes: The system detects a first editing operation performed on a first field on the target interface and a second editing operation performed on a second field on the target interface. The first field is used to set the impact of faults in the bus port and bus device on the server's operating performance. The second field is used to set a threshold value for the number of descriptive information stored in the register corresponding to the impact of the faults in the bus port and bus device on the server's operating performance. The first editing operation sets the impact of the faults in the bus port and bus device on the server's operating performance to a target impact level. The second editing operation adjusts the threshold value for the number of descriptive information stored in the register corresponding to the target impact level to the target threshold value. Generate the target adjustment request corresponding to the first editing operation and the second editing operation, and transmit the target adjustment request to the target firmware.

4. The method according to claim 2, characterized in that, After sending the target adjustment request to the adjustment device, the method further includes: Define a threshold identifier in the configuration file, wherein the value of the threshold identifier is used to identify the quantity threshold corresponding to the quantity of descriptive information stored in the register; The entry function is called to adjust the value of the threshold identifier to the target quantity threshold.

5. The method according to claim 1, characterized in that, The server includes a clock chip connected to the target firmware. The clock chip is configured to record the current time of the server. Before querying the number of candidate descriptions accumulated in the registers of the query server at the current time, the method includes: The target address of the clock chip is detected, wherein the target address is used to extract the time recorded in the clock chip; By accessing the target address, the time recorded in the clock chip is extracted and used as the current time.

6. The method according to claim 1, characterized in that, The number of candidate descriptions accumulated in the registers of the query server at the current time includes: When the bus ports include N bus ports, the following steps are performed to query the number of candidate description information entries recorded in the register at the current time, wherein the register is configured to record description information of faults occurring in the N bus ports and the bus devices connected to each of the N bus ports, where N is a positive integer: Detect the N sets of port description information of the N bus ports, wherein the i-th set of port description information in the N sets of port description information is used to indicate the i-th bus port in the N bus ports, and i is a positive integer less than or equal to N; The target interrupt is invoked to poll the register based on the N sets of port description information at the current time to obtain the number of description information corresponding to the N bus ports and the bus devices connected to each of the N bus ports. Perform a summation operation on the N quantities to obtain the candidate quantities.

7. The method according to claim 1, characterized in that, The step of detecting whether the candidate number and the target number threshold satisfy a first matching condition to obtain a first detection result includes: Detect whether the number of candidates is greater than or equal to the target number threshold; If the number of candidates is detected to be greater than or equal to the target number threshold, the first detection result is determined to indicate that the number of candidates and the target number threshold satisfy the first matching condition; if the number of candidates is detected to be less than the target number threshold, the first detection result is determined to indicate that the number of candidates and the target number threshold do not satisfy the first matching condition.

8. The method according to claim 1, characterized in that, The step of detecting whether the target duration between the current time and the initial time satisfies a second matching condition with a duration threshold, and obtaining a second detection result, includes: Detect whether the target duration is less than or equal to the duration threshold; If the target duration is detected to be less than or equal to the duration threshold, the second detection result is determined to indicate that the target duration and the duration threshold satisfy the second matching condition; if the target duration is detected to be greater than the duration threshold, the second detection result is determined to indicate that the target duration and the duration threshold do not satisfy the second matching condition.

9. The method according to claim 1, characterized in that, The step of filtering target description information to be deleted from the candidate description information based on the first detection result and the second detection result includes: When the first detection result indicates that the candidate quantity and the target quantity threshold satisfy the first matching condition, and the second detection result indicates that the target duration and the duration threshold satisfy the second matching condition, the description information of the target quantity threshold is selected from the candidate description information as the target description information; If the first detection result indicates that the first matching condition is not met between the candidate number and the target number threshold, and / or the second detection result indicates that the second matching condition is not met between the target duration and the duration threshold, the candidate description information is determined as the target description information.

10. The method according to claim 1, characterized in that, The step of determining whether to transmit the updated description information stored in the register to the controller in the server includes: If the first detection result indicates that the candidate number and the target number threshold satisfy the first matching condition, and the second detection result indicates that the target duration and the duration threshold satisfy the second matching condition, then it is determined to transmit the updated description information to the controller. If the first detection result indicates that the first matching condition is not met between the candidate number and the target number threshold, and / or the second detection result indicates that the second matching condition is not met between the target duration and the duration threshold, it is determined that the updated description information will not be transmitted to the controller.

11. The method according to claim 1, characterized in that, The controller includes a data table, the server includes a recording function, and transmitting the updated description information to the controller includes: The recording function is called to record the update description information into the data table one by one, and the update description information is transmitted to the controller one by one.

12. The method according to claim 1, characterized in that, The degree of influence is calculated using the following formula: Where R represents the degree of impact, S1 represents the server's operating performance when the bus port and the bus device connected to the bus port are not faulty, and S2 represents the server's operating performance after the bus port and the bus device connected to the bus port have failed.

13. The method according to claim 8, characterized in that, The target duration is calculated using the following formula: T=(T2D*24+T2H)-(T0D*24+T0H); Wherein, T represents the target duration between the current time and the initial time, T2D represents the number of days of the current time, T2H represents the number of hours of the current time, T0D represents the number of days of the time when the fault description information was first recorded, and T0H represents the number of hours of the time when the fault description information was first recorded.

14. The method according to claim 1, characterized in that, The method further includes: If the number of candidate description information stored in the register at the current time is greater than or equal to a number threshold, it is further detected whether the number of candidate description information stored in the register at the current time is greater than or equal to an upper limit value, wherein the upper limit value is greater than the target number threshold, and the upper limit value is the threshold corresponding to the condition that the impact of the failure of the bus port and the bus device connected to the bus port on the operating performance of the server is greater than or equal to the impact degree threshold. If the number of candidate description information is detected to be greater than or equal to the upper limit, the description information stored in the register after the target description information is determined as the updated description information, the updated description information is transmitted to the controller, and the description information of the upper limit value is deleted.

15. The method according to claim 2, characterized in that, A fault diagnosis driver is deployed on the adjustment component, and the adjustment component is configured to execute the target adjustment request through the fault diagnosis driver.

16. A control device for server fault information, characterized in that, The device is applied to target firmware in a server, and the device includes: The first processing module is configured to query the number of candidate descriptions accumulated in the register of the server at the current time, and extract the initial time for storing the first description information in the candidate description information into the register, wherein the initial time is earlier than the current time, and the register is configured to store description information of faults occurring on the bus port of the server and description information of faults occurring on the bus device connected to the bus port. The first detection module is configured to detect whether the candidate number and the target number threshold satisfy a first matching condition to obtain a first detection result, and to detect whether the target duration between the current time and the initial time and the duration threshold satisfy a second matching condition to obtain a second detection result, wherein the target number threshold is determined by the target firmware based on the degree of impact of the faults of the bus port and the bus device on the operating performance of the server; The second processing module is configured to filter target description information to be deleted from the candidate description information based on the first detection result and the second detection result, and determine whether to transmit the updated description information stored in the register to the controller in the server, wherein the updated description information is the description information stored in the register after the target description information, and the candidate description information includes the updated description information; The third processing module is configured to delete the target description information and, if it is determined that the updated description information should be transmitted to the controller, transmit the updated description information to the controller.

17. The apparatus according to claim 16, characterized in that, The server includes an adjustment device connected to the target firmware, and the device further includes: The receiving module is configured to receive a target adjustment request before detecting whether a first matching condition is met between the candidate quantity and the target quantity threshold and obtaining a first detection result, wherein the target adjustment request is used to request that the quantity threshold corresponding to the quantity of description information stored in the register be adjusted to the target quantity threshold. The sending module is configured to send the target adjustment request to the adjustment device, wherein the adjustment device is used to execute the target adjustment request.

18. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 15.

19. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 15.

20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1 to 15.