Server configuration method, server management method, and server, program product, device and medium

WO2026179150A1PCT designated stage Publication Date: 2026-09-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/123512
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2025-09-24
Publication Date
2026-09-03

Smart Images

  • Figure CN2025123512_03092026_PF_FP_ABST
    Figure CN2025123512_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and in particular to a server configuration method, a server management method, and a server, a program product, a device and a medium. The server configuration method comprises: in response to booting of a server, a basic input / output system initializing a data space shared by a first mode and a second mode, and invoking initialization data in the data space to initialize the first mode and the second mode, wherein the first mode runs on a baseboard management controller, and the second mode runs on an operating system; the baseboard management controller or the operating system using one of the first mode and the second mode as a target detection mode, and feeding back the target detection mode to a programmable logic device; and the programmable logic device enabling a path between a fault signal terminal and an execution entity of the target detection mode, and performing fault detection on the server by using the target detection mode. The present method can be compatible with a plurality of server fault detection modes, so as to facilitate optimization of server performance.
Need to check novelty before this filing date? Find Prior Art

Description

Server configuration and management methods, servers, software products, equipment and media

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510230755.9, filed on February 28, 2025, entitled "Server Configuration and Management Method, Server, Program Product, Device and Media", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of computer technology, and in particular to a server configuration method, a server management method, a server, a program product, a device, and a medium. Background Technology

[0004] A server is an internet-connected device in a network environment that provides computing power and runs software applications. It typically provides computing or application services to other client devices such as personal computers and smartphones. For example, a server may have the capability to respond to service requests, provide services, and ensure service availability.

[0005] To improve the stability of server operation, fault detection strategies such as static fault detection can be used to detect server faults. However, the inventors realized that with the increasing complexity and diversification of server applications, traditional static fault detection strategies are difficult to meet the current needs for efficient and accurate fault detection. Summary of the Invention

[0006] Therefore, it is necessary to provide a server configuration method, server management method, server, program product, device and media that are compatible with multiple server fault detection modes to address the above-mentioned technical problems and thus help optimize server performance.

[0007] On one hand, a server configuration method is provided, comprising: in response to server startup, a basic input / output system initializes a data space shared by a first mode and a second mode, and calls initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on a baseboard management controller, and the second mode runs on an operating system; the basic input / output system or the baseboard management controller uses one of the first mode and the second mode as a target detection mode, and feeds back the target detection mode to a programmable logic device (PLD); in response to acquiring the target detection mode, the PLD selects a path between a fault signal terminal and the operating entity of the target detection mode, so as to use the target detection mode to perform fault detection on the server; wherein the fault signal terminal is used to transmit correctable errors, and the operating entity is the baseboard management controller or the operating system.

[0008] On the other hand, a server management method is provided, which includes: monitoring the operating parameters of the server; and evaluating whether the server is faulty based on the operating parameters using a target detection mode; wherein the target detection mode is activated using the server configuration method as described in any of the above embodiments.

[0009] On another front, a server is provided, comprising: a server body and a control system; the control system is located on the server body and is used to implement the following steps: in response to server startup, the basic input / output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein, the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input / output system or the baseboard management controller uses one of the first mode and the second mode as the target detection mode, and feeds back the target detection mode to the programmable logic device (PLD); in response to acquiring the target detection mode, the PLD selects the path between the fault signal terminal and the running entity of the target detection mode, so as to use the target detection mode to perform fault detection on the server; wherein, the fault signal terminal is used to transmit correctable errors, and the running entity is the baseboard management controller or the operating system. Alternatively, the following steps are implemented: monitoring the server's operating parameters; evaluating whether the server is faulty based on the operating parameters using the target detection mode; wherein, the target detection mode is activated using the server configuration method as described in any of the above embodiments.

[0010] On another front, a computer-readable instruction product is provided, comprising computer-readable instructions that, when executed by a processor, perform the following steps: in response to server startup, a basic input / output system initializes a data space shared by a first mode and a second mode, and calls initialization data within the data space to initialize the first mode and the second mode; wherein the first mode runs on a baseboard management controller, and the second mode runs on an operating system; the basic input / output system or the baseboard management controller uses one of the first mode or the second mode as a target detection mode and feeds back the target detection mode to a programmable logic device (PLD); in response to acquiring the target detection mode, the PLD selects a path between a fault signal terminal and the operating entity of the target detection mode to perform fault detection on the server using the target detection mode; wherein the fault signal terminal is used to transmit correctable errors, and the operating entity is the baseboard management controller or the operating system. Alternatively, the following steps are implemented: monitoring the server's operating parameters; evaluating whether the server is faulty based on the operating parameters using the target detection mode; wherein the target detection mode is activated using the server configuration method as described in any of the above embodiments.

[0011] On another front, a computer device is provided, comprising a memory and one or more processors. The memory stores computer-readable instructions, which, when executed by the one or more processors, cause the one or more processors to perform the following steps: in response to server startup, a basic input / output system initializes a data space shared by a first mode and a second mode, and initializes the first mode and the second mode by calling initialization data in the data space; wherein the first mode runs on a baseboard management controller, and the second mode runs on an operating system; the basic input / output system or the baseboard management controller uses one of the first mode and the second mode as a target detection mode and feeds the target detection mode back to a programmable logic device (PLD); in response to acquiring the target detection mode, the PLD selects a path between a fault signal terminal and the operating entity of the target detection mode to perform fault detection on the server using the target detection mode; wherein the fault signal terminal is used to transmit correctable errors, and the operating entity is the baseboard management controller or the operating system. Alternatively, the following steps are implemented: monitoring the server's operating parameters; evaluating whether the server is faulty based on the operating parameters using the target detection mode; wherein the target detection mode is activated using the server configuration method as described in any of the above embodiments.

[0012] On another front, a non-transitory computer-readable storage medium is provided, on which computer-readable instructions are stored. When executed by a processor, these computer-readable instructions perform the following steps: in response to server startup, initializing multiple fault detection modes respectively; selecting one of the multiple fault detection modes as the target detection mode; configuring the fault detection environment of the target detection mode to activate the target detection mode. Alternatively, the following steps are implemented: monitoring the server's operating parameters; evaluating whether the server is faulty based on the operating parameters using the target detection mode; wherein the target detection mode is activated using the server configuration method as described in any of the above embodiments. Attached Figure Description

[0013] Figure 1 is a flowchart illustrating an embodiment of the server management method of this application;

[0014] Figure 2 is a flowchart illustrating the first embodiment of the server configuration method of this application;

[0015] Figure 3 is a flowchart illustrating the second embodiment of the server configuration method of this application;

[0016] Figure 4 is a flowchart illustrating the third embodiment of the server configuration method of this application;

[0017] Figure 5 is a structural schematic diagram of an embodiment of the server management system of this application;

[0018] Figure 6 is a flowchart illustrating the fourth embodiment of the server configuration method of this application;

[0019] Figure 7 is a schematic diagram of the structure of a server according to an embodiment of this application;

[0020] Figure 8 is a structural schematic diagram of an embodiment of the computer device of this application;

[0021] Figure 9 is a schematic diagram of the structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0023] To address the technical problem that existing fault detection strategies in related technologies are ill-suited to the current demands for efficient and accurate fault detection, this application provides a server configuration method, a server management method, a server, a program product, a device, and a medium. The server configuration method includes: initializing multiple fault detection modes in response to server startup; selecting one of the multiple fault detection modes as the target detection mode; and configuring the fault detection environment for the target detection mode to activate it. Thus, this application can be compatible with multiple server fault detection modes, which helps improve the accuracy and timeliness of server fault detection, thereby improving server performance. The specific working principle of this application is illustrated below with examples.

[0024] Please refer to Figure 1, which is a flowchart illustrating an embodiment of the server management method of this application.

[0025] S101: Monitor the server's operating parameters.

[0026] In some embodiments of this application, the operating parameters include relevant data generated during server operation.

[0027] For example, sensors can be used to sense the server and obtain sensing data, which can be used as one of the operating parameters; resource usage information can also be used as one of the operating parameters; or the server's computing efficiency, response efficiency, etc. can be monitored and computing power-related data can be used as one of the operating parameters, without any limitation.

[0028] S102: Evaluate whether the server is faulty based on the target detection mode using the operating parameters; wherein, the target detection mode is activated using the server configuration method.

[0029] In some embodiments of this application, target detection modes can be used to detect whether a server currently has a fault; and / or to analyze whether a server has a fault trend.

[0030] In some embodiments of this application, the server can be compatible with multiple fault detection modes, and the target detection mode is one of the multiple fault detection modes.

[0031] The server can be compatible with multiple fault detection modes, allowing it to choose one as the target detection mode. This means it can flexibly select the fault detection mode that best suits the current business and application scenarios. Furthermore, it can switch to a different target detection mode if the current mode's performance is unsatisfactory. This improves the adaptability of the target detection mode to the current scenario, enhances the timeliness and accuracy of fault detection, and ultimately ensures the server's operational reliability and stability, while also improving server performance.

[0032] Please refer to Figure 2, which is a flowchart of the first embodiment of the server configuration method of this application.

[0033] S201: In response to server startup, the basic input / output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein, the first mode runs on the baseboard management controller, and the second mode runs on the operating system.

[0034] In some embodiments of this application, the first mode and the second mode are two fault detection modes; in other words, the server's fault detection modes include at least the first mode and the second mode. The first mode operates on the baseboard management controller, and the second mode operates on the operating system.

[0035] As the name suggests, the fault detection mode has a fault detection function. In some embodiments of this application, multiple fault detection modes can be preset, and the fault detection principles of the multiple fault detection modes can be different. That is, the fault detection principles of the first mode and the second mode can be different. The detailed working principles of the first mode and the second mode will be described in detail later, and will not be repeated here.

[0036] When the server starts, it can initialize various fault detection modes, including the first mode and the second mode, so that the server can basically support fault detection modes. In other words, it can provide basic configuration for using fault detection modes as target detection modes in the future, which will help simplify the configuration work when setting them as target detection modes and improve the switching efficiency of fault detection modes as target detection modes.

[0037] Meanwhile, in some embodiments of this application, a data space is also created. As the name suggests, the data space can be used to store data. In some embodiments of this application, the data space created can be in a shared state, meaning that both the first mode and the second mode can access the data within the data space. In some embodiments of this application, the basic input / output system can also initialize the data space shared by both the first mode and the second mode, so that when in the first mode or the second mode, the data space can be accessed to obtain the data therein.

[0038] S202: The basic input / output system or baseboard management controller takes one of the first mode and the second mode as the target detection mode and feeds back the target detection mode to the programmable logic device.

[0039] In some embodiments of this application, the Basic Input / Output System (PIOS) can use either a first mode or a second mode as the target detection mode. Alternatively, the Baseboard Management Controller (BMC) can use either a first mode or a second mode as the target detection mode. That is, in some embodiments of this application, the selection / control of the fault detection mode as the target detection mode can be performed by either the PIOS or the BMC. For example, during the initial stage of server startup, the PIOS can select the fault detection mode as the target detection mode; after the server startup is complete, the BMC can select the fault detection mode as the target detection mode.

[0040] When either the first mode or the second mode is selected as the target detection mode, the target detection mode can be fed back to the programmable logic device (PLD), so that the PLD can adaptively control the activation of the target detection mode.

[0041] S203: In response to acquiring the target detection mode, the programmable logic device selects the path between the fault signal terminal and the operating entity of the target detection mode to perform fault detection on the server using the target detection mode; wherein, the fault signal terminal is used to transmit correctable errors, and the operating entity is the baseboard management controller or the operating system.

[0042] In some embodiments of this application, the programmable logic device (PLD) can be connected to the operating entities of the first mode and the second mode respectively; that is, the PLD can be connected to the baseboard management controller and the operating system respectively. Furthermore, the PLD can also be connected to a fault signal terminal to transmit error information fed back from the first mode and the second mode to the fault signal terminal.

[0043] The programmable logic controller (PLC) responds to the target detection mode feedback from the baseboard management controller (BMC) or basic input / output system (PIS), and can select the path between the fault signal terminal and the operating entity of the target detection mode. That is, when the first mode is the target detection mode, the PLC can select the path between the fault signal terminal and the baseboard management controller; when the second mode is the target detection mode, the PLC can select the path between the fault signal terminal and the operating system.

[0044] As can be seen, in some embodiments of this application, the server can be configured to be compatible with multiple fault detection modes. That is, the server can be at least compatible with a first mode running on the baseboard management controller and a second mode running on the operating system. After the server starts, multiple fault detection modes can be initialized separately, thus providing a basic application environment for the fault detection modes. The first mode and the second mode can share a data space, which is equivalent to normalizing the first mode and the second mode, and helps ensure the synchronization of the data acquired by the two modes. Furthermore, this application uses a programmable logic device to hardware isolate the first mode and the second mode, which helps ensure the independent operation and reliability of the first mode and the second mode. Thus, one of the first mode and the second mode can be selected as the current target detection mode of the server, and the target detection mode can be configured to activate it, using the target detection mode to perform fault detection on the server. Thus, according to the current fault detection needs of the server, the fault detection mode can be flexibly selected from multiple fault detection modes as the target detection mode, thereby improving the accuracy and timeliness of server fault detection by improving the adaptability of the target detection mode to the actual business scenario, and thus improving server performance.

[0045] In other words, after the server starts, multiple fault detection modes can be initialized separately, thus providing a basic application environment for these modes. Subsequently, one of these fault detection modes can be selected as the server's current target detection mode, configured to activate it, and used to perform fault detection on the server. Therefore, this application is compatible with multiple fault detection modes, allowing for flexible selection of the target mode based on the server's current fault detection needs. By improving the adaptability of the target detection mode to the actual business scenario, the accuracy and timeliness of server fault detection can be improved, thereby enhancing server performance.

[0046] Furthermore, when initializing multiple fault detection modes separately, the initialization data and status of each fault detection mode can be configured to enable the server to support the fault detection function of each mode. During initialization, when the first and second modes can share data in the data space, and the baseboard management controller and operating system obtain data from the data space to initialize the first and second modes, they can acquire the data required by the fault detection mode they are running. Compared to passively receiving data, even with a shared data space, this method helps isolate the initialization data transmission between fault detection modes. Furthermore, isolation between the first and second modes can be achieved through methods such as programmable logic devices. This ensures reliable data transmission for each fault detection mode, reduces the impact of data transmission interference on the reliability of each fault detection function configuration, and facilitates the relative independence of the fault detection functions between each mode, reducing the risk of mutual interference.

[0047] As illustrated in the embodiments, the fault detection modes may include at least a first mode and a second mode. For example, the first mode performs correctable error handling in the baseboard management controller; the second mode performs error handling through the platform operation mechanism. Furthermore, the fault detection modes may also include other fault detection modes such as static fault detection modes, which are not limited here.

[0048] For example, the first mode could include features such as RAS (Reliability, Availability, and Serviceability), which is one of the most important functions of a server.

[0049] The second mode could be something like PRM (Platform Runtime Mechanism).

[0050] Please refer to Figure 3, which is a flowchart of the second embodiment of the server configuration method of this application.

[0051] S301: Reads the device options for fault detection and provides fault feedback via an interrupt signal.

[0052] In some embodiments of this application, in response to the initial stage of server startup, the basic input / output system can read the fault detection settings options to obtain the preset fault detection mode identified by the settings options, which can be used as the target detection mode.

[0053] The server startup process can include an initial phase and a transition phase. The initial phase (Post phase) indicates that the server system is controlled by a basic input / output system (BIOS). The transition phase (Runtime phase) indicates that the server system is controlled by an operating system (OS).

[0054] In other words, in some embodiments of this application, setting options can be added within the server to support the ability to pre-set the server's default fault detection mode, so that the server's target detection mode setting can be achieved even without external intervention and without adaptive matching strategies.

[0055] For example, a Fault Diagnosis Select option can be pre-created, allowing selection of a default fault detection mode, i.e., the preset fault detection mode. During this process, either the first or second mode can be set to the preset fault detection mode, i.e., set to RAS offload or PRM. Furthermore, faults in the post-stage can be reported via traditional SMI (System Management Interrupt).

[0056] In other words, during the initial stage of server startup, before fault detection is performed using the target detection mode, the basic input / output system can trigger a system management interrupt to provide fault feedback via the interrupt signal.

[0057] S302: Verify whether the baseboard management controller supports the first mode.

[0058] In some embodiments of this application, when it is determined that the baseboard management controller supports the first mode, step S303 is executed; when it is determined that the baseboard management controller does not support the first mode, step S304 is executed.

[0059] In other words, during the initialization of the fault detection mode, it's also possible to verify whether the drive control module of the fault detection mode supports its functionality. If the drive control module supports the fault detection mode, the initialization data for the fault detection mode is sent to it. If the drive control module does not support the fault detection mode, initialization of the unsupported fault detection mode is unnecessary, improving initialization efficiency and indicating that the fault detection mode cannot be used as a target detection mode. In simpler terms, it can handshake with the BMC (baseboard management controller) to confirm whether the BMC supports the RAS offload function.

[0060] Some embodiments of this application provide a server configuration method with good generalization performance, aiming to enable the server to be compatible with multiple fault detection functions, ideally supporting all fault detection modes. Furthermore, some embodiments of this application consider that in some cases, the server may not support the corresponding fault detection function. Therefore, the server's support for various fault detection modes can be explicitly verified in advance. This reduces the likelihood of invalid fault detection occurring when the corresponding driver control module does not support the implementation of a fault detection mode, thus lowering the risk that the target detection mode cannot detect server faults. This improves the reliability of the server configuration method and ensures that the target detection mode can reliably operate for server fault detection.

[0061] Furthermore, whether each fault detection mode is supported by the corresponding drive control module can be determined at different stages, which will not be elaborated here.

[0062] In some embodiments of this application, it can be verified whether the baseboard management controller supports a first mode during the initial stage of server startup; wherein, the server system is controlled by the basic input / output system during the initial stage. In response to the baseboard management controller supporting the first mode and the first mode being a target detection mode, an enable command is sent to the baseboard management controller to enable the execution of corrective error handling functions.

[0063] S303: Send the first parameter combination of the first mode initialization to the baseboard management controller.

[0064] In some embodiments of this application, a first combination of parameters for initialization of the first mode can be configured.

[0065] The first parameter combination includes functional support capability parameters, correctable error handling strategy parameters, bus topology, and memory topology. This first parameter combination is sent to the baseboard management controller (BMC), allowing the BMC to initialize functions and support configurations related to the first mode. In simpler terms, it prepares RAS offload-related data to be sent to the BMC. In some embodiments of this application, the RAS function can be offloaded to the OS processing, thereby improving the stability and reliability of the server system. Furthermore, real-time dynamic updates of components such as the SMM driver (a firmware driver) can be performed to ensure the server system remains up-to-date and reduce security issues caused by software omissions or outdated drivers.

[0066] S304: Initialize the message handler for the second mode.

[0067] In some embodiments of this application, the message processor of the second mode can be initialized.

[0068] The transmission port associated with the second mode can be configured to execute corresponding code commands when the associated register signal is triggered. A second parameter combination for second mode initialization can also be configured; wherein the second parameter combination includes correctable error handling strategy parameters, bus topology, and memory topology. In other words, a PRM handler can be initialized. Here, the handler is a message processing device; that is, in some embodiments of this application, the message processor may include a PRM handler.

[0069] S305: Sends the target detection mode configuration information to the programmable logic device (PLD) so that it can configure the control module connected to the fault signal terminal based on the configuration information.

[0070] In some embodiments of this application, the second mode of target detection is used as an example for illustration.

[0071] In other words, in some embodiments of this application, configuration information for the target detection mode can be obtained. This configuration information is then sent to a programmable logic device (PLD), which configures the control module connected to the fault signal terminal based on the configuration information. The control module includes a central processing unit (CPU) and a baseboard management controller.

[0072] In other words, the fault detection mode of the BIOS setup can be obtained as the target detection mode, and the corresponding configuration information can be sent to the CPLD via methods such as VGPIO. If so, the CPLD can configure the fault signal terminal to be connected to the CPU or BMC according to the fault detection mode transmitted by the BIOS.

[0073] VGPIO stands for Voltage Control GPIO (General Purpose Input / Output).

[0074] CPLD (Complex Programmable Logic Device) is a type of programmable logic device with high density characteristics.

[0075] The CPU (Central Processing Unit) is the core of the server system for computation and control.

[0076] The fault signal terminal can be defined as an "error pin", meaning that the CPLD can configure error pin #0 to the CPU or BMC.

[0077] In some embodiments of this application, steps S301 to S305 may occur during the post-startup phase of the server, i.e., the initial startup phase of the server. The post-startup phase / initial phase indicates that the server's BIOS controls the startup process during this phase.

[0078] As described above, when the first mode is the target detection mode, an enable command can be sent to the BMC. Specifically, the BIOS can send a command such as "Send Start Cmd w / Offload Features" to the BMC, which indicates that the features related to the first function are enabled. If so, when the first mode is activated, the BMC can poll the error pin to collect error information and report it.

[0079] S306: Building management software for a second mode.

[0080] In some embodiments of this application, the management software is associated with the operating system, and the management software is also used to monitor and handle server failures when the second mode is the target detection mode. The management software may be PRM SW (software), etc., and is not limited thereto.

[0081] A second mode of management software can be built to trigger the management software to obtain functional support information.

[0082] As explained earlier, the verification of driver control module support for the first mode can be performed in the post-processing stage. In this step, the verification of driver control module support for the second mode can be performed in the runtime stage, which represents the stage during server startup where the OS (Operating System) controls the server system. The driver control modules for the first and second modes are different. In other words, among multiple fault detection modes, at least some fault detection modes can share the same driver control module, while different fault detection modes are permissible.

[0083] S307: Evaluate whether the operating system supports the second mode.

[0084] In some embodiments of this application, when the evaluation determines that the operating system supports the second mode, step S308 is executed; when the evaluation determines that the operating system does not support the second mode, step S309 is executed.

[0085] Specifically, in response to the second mode being used as the target detection mode, the operating system's functional support information can be obtained before entering the server system to assess whether the operating system supports the second mode. Upon determining that the operating system supports the second mode, further settings for the second mode are then implemented. These settings include controlling the feedback of correctable errors through the second mode.

[0086] In other words, a PRM switch can be triggered before booting into the system to check the OS's support for PRM. Here, "Boot" can refer to the process where the server system loads the operating system from a storage device when the server is powered on, i.e., the system is in the boot process.

[0087] It should be noted that in some embodiments of this application, the second mode is exemplified by the target detection mode preset in the BIOS. When the preset target detection mode is the first mode, after executing step S305, an enable signal can be sent to the baseboard management controller to activate the first mode, and the baseboard management controller polls the error pin to collect fault information.

[0088] S308: Set the second mode and perform fault detection through the second mode.

[0089] In some embodiments of this application, if the system supports PRM, PRM settings can be configured to control CE (correctable error) to be reported through PRM.

[0090] S309: Feedback to the programmable logic device that the operating system does not support the second mode, so that the target detection mode is updated to the first mode.

[0091] In some embodiments of this application, in response to a determination that the operating system does not support the second mode, a path can be established between the fault signal terminal and the baseboard management controller; wherein, the fault signal terminal is used to transmit correctable errors. The correctable errors are fed back through the first mode, sending an enable signal to the baseboard management controller to update the target detection mode to the first mode.

[0092] In other words, if the OS does not support PRM, it can send a command to the CPLD to control the error pin to switch the connection to the BMC, and the CE can report it through RAS offload.

[0093] In some embodiments of this application, when the operating system does not support the second mode and the second mode is used as the target detection mode, it can be promptly detected that the second mode is not supported, which can be considered as not being able to reliably implement or even implement the fault detection function. It should be noted that this is considered as the second mode may not be able to be implemented or not being able to be reliably implemented, not that the second mode is definitely not able to be implemented. At this time, the target detection mode can be switched in time, so that the first mode supported by its drive control module is used as the target detection mode, which is conducive to ensuring the reliable implementation of the server fault detection function.

[0094] In some embodiments of this application, in response to the occurrence that neither the first mode nor the second mode is supported, the server can be fault detected by means of static fault detection, and / or the relevant situation can be reported.

[0095] S310: Send an enable signal to the board management controller to poll the status of the fault signal terminal to collect fault information.

[0096] In some embodiments of this application, the BIOS may also send an enable signal to the BMC to enable the RAS offload function. When RAS offload is activated, the BMC may poll the error pin status as described above to collect error information and report it.

[0097] Thus, server fault detection is performed using either the first or second mode as the target detection mode until the reporting is completed. This could refer to the server being shut down or stopped from use, etc., without any specific limitation.

[0098] Please refer to Figure 4, which is a flowchart of the third embodiment of the server configuration method of this application.

[0099] S401: Monitoring platform management interface.

[0100] In some embodiments of this application, the platform management interface is used for communication with external servers. In simpler terms, users can input control commands through the platform management interface.

[0101] For example, the platform management interface could be an intelligent platform management interface, etc.

[0102] S402: Obtain switching command.

[0103] In some embodiments of this application, in response to the initial stage of startup via the server, i.e., when the operating system takes over control, the server, host, operating system, etc., can obtain externally input switching instructions so that the baseboard management controller can obtain switching instructions. These switching instructions can originate from outside the server, such as the platform management interface described in step S401.

[0104] Thus, the baseboard management controller can parse the fault detection mode to be switched identified by the switching command and use it as the target detection mode. In some embodiments of this application, it can also support external input switching commands to switch the fault detection mode currently used as the target detection mode on the server, thereby enabling the server's target detection mode to adapt to the user's real-time needs. Furthermore, in some embodiments of this application, configuration is performed in advance using related intelligent platform management interface commands, i.e., in response to passing the initial stage, switching the fault detection mode used as the target detection mode is allowed through the platform management interface, thus achieving simple and free switching between multiple fault detection modes.

[0105] In other words, a solution can be provided that allows users to proactively switch fault detection modes, thereby improving the efficiency of server fault identification. Users can actively switch to the appropriate fault detection mode based on their business usage scenarios and customized strategies, using this as the target detection mode, thus reducing fault handling time and improving business continuity. Furthermore, timely adjustments to hardware registers can improve the effectiveness of server resource utilization, reduce the risk of resource waste, and ultimately improve overall operational efficiency.

[0106] For example, PRM can be enabled through a preset interface.

[0107] S403: Parse the fault detection mode to be switched identified by the switching command and use it as the target detection mode.

[0108] S404: Specifies the transmission port, triggering a system management interrupt through the specified transmission port.

[0109] In some embodiments of this application, a VGPIO may be specified to allow SMI to be triggered via VGPIO.

[0110] S405: Configure the target detection mode indicated by the switching command.

[0111] In some embodiments of this application, taking the indicated target detection mode as PRM as an example, one can enter the PRM enable handler.

[0112] S406: Verify whether the drive control module supports target detection mode.

[0113] In some embodiments of this application, when it is determined that the drive control module supports the target detection mode indicated by the switching command, step S407 is executed; when it is determined that the drive control module does not support the target detection mode indicated by the switching command, step S409 is executed.

[0114] S407: Feeds back the target detection mode to the programmable logic unit to enable the path between the fault signal terminal and the drive control module.

[0115] S408: Fault detection using a new target detection mode.

[0116] In some embodiments of this application, it can be determined whether the OS supports PRM. If it does, a command can be sent to the CPLD to switch the error pin connection to the CPU and execute PRM to report CE.

[0117] S409: Exit system management interrupt and resume the original target detection mode for fault detection.

[0118] In some embodiments of this application, for example, it can be determined whether the OS supports PRM. If the OS does not support PRM, the SMI can be exited and an error returned, and CE reporting can be performed through the original target detection mode.

[0119] In some embodiments of this application, the original target detection mode is taken as the first mode or other fault detection modes. In some embodiments of this application, the functionality of the switching instruction can also be simplified, i.e., the switching instruction only has a switching function, thereby reducing the data volume of the switching instruction. That is, when a switching instruction is received, the fault detection mode used as the target detection mode is switched. For example, taking the fault detection modes as including a first mode and a second mode, and the original target detection mode as the first mode, when a switching instruction is received, the target detection mode is switched to the second mode.

[0120] In layman's terms, when an SMI is triggered by an interface such as VGPIO, regardless of whether the current target detection mode is the first or second mode, it can switch to the other. That is, if the current target detection mode is PRM, it can switch to RAS offload; if the current target detection mode is RAS offload, it can switch to PRM.

[0121] In some embodiments of this application, adaptive identification of server status can also be supported to spontaneously switch to a fault detection mode that better matches the current server status. This proactive switching of fault detection modes can effectively reduce server operation risks, alleviate the workload of maintenance personnel, and through real-time monitoring and data analysis, it can help to discover potential server faults in advance for corresponding preventive maintenance, reduce the risk of actual server faults, thereby reducing server downtime caused by faults and improving business continuity.

[0122] Specifically, the baseboard management controller can acquire the server's status parameters. In response to acquiring the server's status parameters, the baseboard management controller can analyze the status parameters to evaluate whether to switch the fault detection mode as the target detection mode. In response to determining to switch the fault detection mode as the target detection mode, the baseboard management controller can control the programmable logic device to select other fault detection modes as the new target detection mode; wherein, other fault detection modes refer to fault detection modes that are not currently being used as the target detection mode.

[0123] For example, as illustrated in the examples above, fault detection modes may include a first mode and a second mode.

[0124] In response to the target detection mode being in the first mode, the baseboard management controller (BMS) or other modules such as the central processing unit (CPU) can obtain the resource utilization parameters of the BMS and compare them with the resource utilization threshold. If the resource utilization parameters reach the threshold, it can be considered that the BMS is under a certain burden in fault detection, and the fault detection mode is switched to the target detection mode. Therefore, it can switch to the second mode as the target detection mode. In other words, the programmable logic will select the path between the fault signal terminal and the operating system, switching to the second mode as the target detection mode.

[0125] In other words, in some embodiments of this application, considering that the baseboard management controller is one of the important control modules of the server, it can typically use sensors to monitor the server. To ensure the reliable implementation of the basic functions of baseboard management control, such as server monitoring, when the resource utilization of the baseboard management controller reaches a certain level, the fault detection function can be transferred to reduce the burden on the baseboard management controller. This not only ensures the stability of the baseboard management controller but also helps to ensure the reliability of the fault detection function, thereby helping to ensure the reliability of the server.

[0126] Furthermore, the fault detection mode can be pre-set with several control options to pre-configure its availability. For example, a fault detection mode can be associated with one control option, and its availability can be toggled by different values ​​of that control option. Alternatively, it can be set using two or more control options; this is not limited here.

[0127] For example, a first control option can be defined for the first mode, and a second control option can be defined for the second mode.

[0128] When both the first and second control options are enabled, one of the default modes (first mode or second mode) can be used as the target detection mode. For example, if the first mode is the default mode, then the first mode can be used as the target detection mode.

[0129] When the first control option is turned on and the second control option is turned off, the first mode can be used as the target detection mode.

[0130] When the first control option is off and the second control option is on, the second mode can be used as the target detection mode.

[0131] When both the first and second control options are turned off, the system management interrupt can be used as the target detection mode.

[0132] Taking the specific control options as an example, the control option for PRM can be defined as PRMsuppot; the control option for RAS offload can be defined as oobrassupport.

[0133] When both OOBRAS Support and PRM Support are enabled, you can use either PRM or the default mode in RAS offload as the target detection mode. For example, if PRM is the default mode, you can use PRM as the target detection mode.

[0134] When oobrassupport is on and PRMsupport is off, RAS offload can be used as the target detection mode.

[0135] When oobrassupport is off and PRMsupport is on, PRM can be used as the target detection mode.

[0136] When both oobrassupport and PRMsupport are turned off, SMI can be used as the target detection mode.

[0137] This option allows you to isolate whether oobrassupport (the first control option) is enabled from whether initialization data transmission is performed. In other words, disabling oobrassupport does not affect the transmission of RasOffLoad initialization data to the BMC, thus reducing the risk of PRM mode affecting address resolution on the BMC side, even if rasoffload is disabled.

[0138] In some embodiments of this application, the collaborative operation of BIOS settings, hardware interfaces such as VGPIO and CPLD, and firmware such as BMC and PRM SW enables flexible switching between the Post-stage and Runtime stages of fault diagnosis strategies. This facilitates adaptation to different business scenario requirements, allowing for the implementation of fault diagnosis strategies that best suit the current server usage scenario, thereby improving the timeliness and reliability of fault detection. Compatibility design between the PRM and RAS Offload schemes can be implemented in the Runtime stage to ensure appropriate fault detection modes are matched to different business scenarios, guaranteeing stable business operation. Furthermore, traditional SMI fault reporting can be performed in the Post-stage, while fault detection and processing can be performed via PRM and RAS Offload under the OS. Switching between the two can also be achieved through platform management interface commands.

[0139] Please refer to Figures 5 and 6. Figure 5 is a structural schematic diagram of an embodiment of the server management system of this application, and Figure 6 is a flowchart of a fourth embodiment of the server configuration method of this application.

[0140] S501: Basic Input / Output System Initialization First Mode.

[0141] In some embodiments of this application, as described above, the first mode may operate on the baseboard management controller.

[0142] When initializing the first mode, the Basic Input / Output System (BIOS) can query the Baseboard Management Controller (BMC) to verify whether the BMC supports the first mode. This pre-verifies the BMC's ability to reliably operate the first mode, enabling reliable fault detection of the server when the first mode is used as a target detection mode. In response to the BMC's feedback that it supports the first mode, the BIOS can send a first parameter combination of the first mode to the BMC to initialize the first mode using this parameter combination.

[0143] Specifically, the basic input / output system can unload the first mode and install it onto the baseboard management controller; and send the first parameter combination to the baseboard management controller. In this way, the baseboard management controller can load the first mode using the acquired first parameter combination, enabling it to run the first mode for server fault detection.

[0144] For example, the basic input / output system can retain target configuration control; where the target configuration control is used to control whether the first mode is enabled or disabled.

[0145] The basic input / output system unloads the first mode function and sends the associated data of the first mode function to the baseboard management controller.

[0146] The first parameter combination includes functional support capability parameters, correctable error handling strategy parameters, bus topology, and memory topology.

[0147] In response to the first mode being the target detection mode, the basic input / output system sends an enable command to the board management controller, causing the board management controller to enable the execution of the corrective error handling function. The enable command carries a function mask, which identifies the first mode function to be activated.

[0148] In other words, the BIOS can enable or disable RAS Offload through specific BIOS knobs (i.e., target configuration controls). During the handshake process, it sends other RAS configuration knob values ​​to the BMC, influencing the BMC's handling of RAS functions through these feedback values. For example, it can send various error thresholds, backup policies, and memory-related configurations (such as mirroring mode). During the POST (Post-Processing) phase, the BIOS performs only the corresponding hardware programming based on the RAS configuration-related knobs (such as setting preset leaky bucket thresholds). Simultaneously, the BMC can determine the RAS processing procedures or functions it executes based on the configuration transmitted by the BIOS. For example, it can decide whether to install backup firmware based on the error correction technology knob for storage devices.

[0149] Therefore, during the server startup phase, the BIOS and BMC can determine the RAS functions offloaded to the BMC through a handshake process, involving the transfer of one or more data structures such as memory topology (MEM_TOPOLOGY), RAS policies (RAS_POLICIES), and IIO topology (IIO_TOPOLOGY). During this process, data interaction can be achieved through platform management interface commands, such as commands to acquire capabilities, send data, and initiate RAS. Simultaneously, the BIOS and BMC can each have corresponding processing APIs to handle data requests, data transmission, and function activation notifications during the handshake process. At runtime, the BMC and host communicate via MMBI to support various RAS functions. The MMBI protocol, as the transport and protocol layer, can carry RAS-related request and response messages, including commands such as PCODE Mailbox Passthru, Runtime sPPR, WHEA Log Passthru, and EDPC Port Info (all RAS-related commands), and can also set the sending and response data formats for each command.

[0150] Furthermore, the baseboard management controller can initialize the library initialization routine for the first mode. The baseboard management controller identifies the first mode function to be started carried by the function mask, initializes the calling interface of the first mode function, and enables the execution of corrective error handling functions.

[0151] Taking the first mode, RAS offload, as an example, during the initialization of services such as RAS-manager, basic event registration and initialization calls can be executed. For example, platform management interface command handlers and inter-process communication interfaces such as DBus (Desktop Bus) can be registered. A series of RAS library initialization routines can also be called, and rasLibInit (a RAS library initialization routine) can be executed as the main initialization routine.

[0152] In this case, the Basic Input / Output System can handshake with the Baseboard Management Controller via the Platform Management Interface commands to control the offloading of RAS functions, transmit relevant data and policies, and notify the Baseboard Management Controller to start processing RAS events.

[0153] For example, the Get Capabilities command (CMD*, 0x22) queries the BMC for RAS capabilities. The BMC returns a list of capabilities, and the BIOS determines the RAS function to be enabled (i.e., the first mode function) based on the list of capabilities returned by the BMC.

[0154] Send Data Command (Send Data CMD*, 0x23): The BIOS sends data and policies related to the unloading function to the BMC, such as memory topology (MEM_TOPOLOGY), RAS policy (RAS_POLICIES), and IIO topology (IIO_TOPOLOGY). Data types are distinguished by different subtypes, such as 0x3 for memory topology, 0x5 for RAS policy, and 0x6 for IIO topology. Each type of data is preceded by a corresponding "Start" subcommand (e.g., 0xFD) and followed by a "Transfer Complete" subcommand (e.g., 0xFE) after transmission.

[0155] Start RAS command (Start RAS CMD*, 0x24): The BIOS notifies the BMC to start processing RAS events and transmits the RAS function startup mask, indicating the RAS function to be started.

[0156] Furthermore, as exemplified above, the first mode is the RAS offload mode. When the Basic Input / Output System initializes the first mode, it can initialize the RAS library, register the command handlers for the platform management interface, etc.

[0157] S502: Basic Input / Output System Initialization Second Mode.

[0158] In some embodiments of this application, as described above, the second mode may run on an operating system.

[0159] Thus, before initializing the second mode, the basic input / output system can determine whether the operating system supports the second mode, and in response to support the second mode, create the operating conditions for the second mode.

[0160] Specifically, the basic input / output system initializes the message processor for the second mode.

[0161] The Basic Input / Output System can register message handlers with a second-mode configuration protocol, allowing the server to invoke message handlers based on the registration information within the configuration protocol when under operating system control.

[0162] Furthermore, as described above, the Elementary Input / Output System (Elementary Input / Output System) can also be configured with a transmission port associated with the second mode, so that a corresponding code command is executed when the associated register signal is triggered. The second parameter combination for configuring the Elementary Input / Output System for the second mode initialization will not be elaborated upon here.

[0163] Furthermore, as illustrated in the previous example, the second mode can be PRM. When the basic input / output system initializes the second mode, it can register macros such as PRM Handler through PRM_MODULE_EXPORT (a macro used to define the module export interface) and install relevant information of PRM Module (such as PRMConfigProtocol (a parameter and setting of PRM used to configure network protocols) into the corresponding environment.

[0164] It should be noted that in some embodiments of this application, the first mode can be initialized first, and the second mode can be initialized after the first mode is initialized. A reasonably planned initialization sequence for fault detection modes helps ensure compatibility between the configurations of the first and second modes and reduces the risk of configuration conflicts, thus ensuring the reliable operation of the functions of the first and second modes subsequently. Furthermore, in some embodiments of this application, the first mode runs on the baseboard management controller, and the second mode runs on the host-side operating system. This facilitates data interaction between the two modes and also allows them to interact and provide error feedback through MMBI (memory map BMC interface). The initialization process of the shared data space between the first and second modes is described in detail below.

[0165] S503: Basic Input / Output System Initialization Shared Data Space.

[0166] In some embodiments of this application, considering that the first mode and the second mode may involve some of the same data, data structures, etc. during the function implementation and / or initialization phases, a data space that the first mode and the second mode can share is created in some embodiments of this application to store data that may be shared by the two and need to be synchronized.

[0167] In other words, in response to server startup, the Basic Input / Output System (PIS) can initialize both the first mode and the second mode, and can also initialize the data space shared by the first mode and the second mode. In some embodiments of this application, the initialization of the data space and the initialization of the first mode and the second mode can be parallel, interleaved, or the data space can be initialized first, so that the first mode and the second mode can call the initialization data in the data space for initialization during initialization.

[0168] Specifically, the basic input / output system initializes the data to be shared.

[0169] The basic input / output system transmits the data to be shared to the preset data space within the baseboard management controller, and maps the data space to the memory settings information of the basic input / output system, so that the baseboard management controller and the basic input / output system can access the data space based on the memory settings information and call the data to be shared to initialize the first mode and the second mode.

[0170] For example, the data to be shared may include correctable error handling policy parameters and at least one of the memory topology.

[0171] In other words, the memory topology can be initialized and the RAS configuration can be initialized during the server startup phase. The initialized memory topology and RAS configuration are then transferred to the shared data space, so that the data in the data space can meet the requirements of both the first mode and the second mode. This makes it convenient for both the first and second modes to call the required RAS data to complete the initialization of their own modules, and reduces the risk of repeated initialization due to data inconsistency.

[0172] In detail, initializing the correctable error handling strategy parameters can be considered as initializing the RAS library. Specifically, the RAS library provides several initialization APIs (Application Programming Interfaces). As mentioned above, rasLibInit must be called before calling other library APIs. Therefore, rasLibInit can be considered the main library initialization routine, used to initialize common components within the library, such as setting default cache property values.

[0173] In addition, each RAS function has its own initialization API, which must be called to initialize the API before calling the corresponding function's "Handle Events" API.

[0174] The following section also describes other information related to the RAS library. For example, RAS function support. Specifically, in some embodiments of this application, multiple RAS functions can be supported, such as memory pass-through certification reports, post-packaging repair, detection and correction of errors in memory, image failover, multiple protocol reporting, low-rate communication recovery, etc. Each function can be triggered through specific associated APIs, and the triggering conditions for each function can be different. For example, memory pass-through certification reports can trigger ERRO (fault signal) through a preset level of leaky bucket configuration.

[0175] Some embodiments of this application may also involve event reporting and logging. The RAS library reports errors and events using preset formats and supports various types of preset format logs, including standard logs (such as platform memory errors, PCIe (peripheral component interconnect express, a high-speed serial computer expansion bus standard) errors, etc.) and non-standard logs (such as memory spare events, cloud server errors, etc.).

[0176] The RAS library can automatically register two internal monitors: WheaLogger and JournalLogger. WheaLogger and JournalLogger are the names of these internal monitors.

[0177] WheaLogger is responsible for adding standard preset format records to the memory area within the RAS library's internal hardware error architecture for storing hardware error records. It can also be configured to send preset format records to a buffer in the OS's hardware error source notification mechanism. JournalLogger can decode all received preset format records into readable strings and output them to standard output.

[0178] Furthermore, the RAS library also offers playback and recording capabilities. Specifically, the RAS library can record external inputs during error handling, such as register accesses, topology and policy data, and MMBI interactions. In playback mode, the RAS library can also retrieve data from pre-recorded files to simulate execution.

[0179] S504: The basic input / output system feeds back the target detection mode to the programmable logic device.

[0180] In some embodiments of this application, taking the Basic Input / Output System (PIS) control during the initial stage of server startup as an example of a fault detection mode as the target detection mode, the PIS can select the fault detection mode as the target detection mode based on whether the running entity supports the fault detection mode it needs to run, and / or, based on a preset fault detection mode. The specific selection method has been described above and will not be repeated here.

[0181] In response to the selection of a fault detection mode as the target detection mode, the basic input / output system can feed back the target detection mode to the programmable logic device (PLD), so that the PLD can adaptively select the fault detection mode as the target detection mode.

[0182] For example, when the first mode is selected as the target detection mode, the basic input / output system can feed back relevant information about the target detection mode being the first mode to the programmable logic device.

[0183] S505: The path between the programmable logic controller's fault signal terminal and the main body of the target detection mode.

[0184] In some embodiments of this application, in response to acquiring a target detection mode, the programmable logic device can select the path between the fault signal terminal and the operating entity of the target detection mode, so as to use the target detection mode to perform fault detection on the server; wherein, the fault signal terminal is used to transmit correctable errors, and the operating entity is a baseboard management controller or an operating system.

[0185] As illustrated in Figure 5, the operating system may include an ESPI interface, a GPIO0 interface, and an ERRO interface. The baseboard management controller includes an ESPI interface and a GPIO interface. The programmable logic device includes a GPIO0 interface, a GPIO1 interface, and a GPIO2 interface.

[0186] The operating system's ESPI interface can be considered as an Enhanced Serial Peripheral Interface (ESPI), used to connect to the ESPI interface of the board management controller via link L3.

[0187] As explained earlier, the second mode runs on the operating system. The operating system's GPIO0 interface is connected to the programmable logic device's GPIO1 interface. That is, when the programmable logic device activates link L4 between its GPIO0 and GPIO1 interfaces, it establishes a path between the fault signal terminal and the operating entity of the second mode (i.e., the operating system). The operating system's GPIO0 interface connects to the operating system's ERRO interface sequentially through the programmable logic device's GPIO1 interface, link L4, the programmable logic device's GPIO0 interface, and link L1. The ERRO interface can represent the fault signal terminal, and it can be located on the CPU, without conflicting with the "error pin" mentioned earlier.

[0188] The GPIO interface of the baseboard management controller can be connected to the GPIO2 interface of the programmable logic controller (PLC). Similarly, when the PLC activates link L2 between its GPIO0 and GPIO2 interfaces, it establishes a path between the fault signal terminal and the operating entity in the first mode (i.e., the baseboard management controller). The GPIO interface of the baseboard management controller is connected to the operating system's ERRO interface sequentially through the PLC's GPIO2 interface, link L2, the PLC's GPIO0 interface, and link L1.

[0189] Meanwhile, the baseboard management controller can also be connected to the programmable logic device via link L5.

[0190] S506: The basic input / output system sends an enable command for target detection mode to the board management controller.

[0191] In some embodiments of this application, the specific methods of sending the enable command have been illustrated in the preceding text and will not be repeated here.

[0192] S507: The baseboard management controller enables the target detection mode.

[0193] In some embodiments of this application, the baseboard management controller enables the target detection mode through a preset interface associated with the target detection mode in response to receiving an enable command.

[0194] S508: The first mode of operation for the substrate management controller.

[0195] In some embodiments of this application, when the first mode is used as the target detection mode, the baseboard management controller can run the first mode to detect whether the server is faulty.

[0196] S509: Operating system running in second mode.

[0197] In some embodiments of this application, when the second mode is used as the target detection mode, the operating system can run the second mode to detect whether the server is faulty.

[0198] Furthermore, in some embodiments of this application, the services initialized above can determine the method for handling RAS events. For example, an interrupt-based approach can be used, i.e., the ERRO signal can be detected, and the ERRO of each communication protocol can be connected to the GPIO of the BMC and configured as level-triggered. Alternatively, RAS events can be handled by polling. For example, some types of error sources can be polled, such as at intervals of 1 minute, 30 seconds, 2 minutes, etc. The specific details can be configured through attribute settings.

[0199] When an error (such as a correctable error) or anomaly is detected on the server, the RAS library can process the error event and construct an error message. The BMC can then transmit the error information to the operating system via MMBI, so that the operating system can be aware of the error.

[0200] The working principles of the first and second modes of this application are illustrated below with examples.

[0201] When the first mode is used as the target detection mode, the first mode detects whether there are correctable errors on the server through error interruption handling or polling; in response to the existence of correctable errors, the first mode constructs an error message through a preset error handling library and passes it to the operating system.

[0202] When the second mode is used as the target detection mode, the second mode configures the backup interrupt as the fault signal terminal option and enables it, triggering the detection server to detect whether there is a correctable error through the system interrupt; in response to the existence of a correctable error, the second mode executes the error handling logic through the message processor and feeds it back to the baseboard management controller.

[0203] Specifically, in response to the number of correctable errors reaching an error threshold, the central processing unit changes the state of the fault signal terminal option, causing a system interrupt to trigger the execution of a pre-stored method within the basic input / output system; wherein, the pre-stored method is used to pass the identifier of the message processor to the operating system;

[0204] The operating system can query the memory location of the message processor that matches the identifier, integrate the buffer information of correctable errors, and pass it to the message processor so that the message processor can process the correctable errors.

[0205] In detail, when the Power Management Rectifier (PRM) handles Error Correction (CE), it can connect Error#0 (a fault signal pin) to a GPIO (General-purpose input / output) that supports triggering System Interrupt (SCI) via hardware. The BIOS configures the GPIO's register to 1, meaning "Routing can cause SCI." Alternatively, the Spare Interrupt option can be set to Error Pin, and IIO Error Pin0 (the bus fault signal pin) can be enabled, while CSMI (SMI-related settings) is disabled. By calling the source language that defines the advanced configuration and power management interface objects through SCI, the PRM is triggered to perform CE processing, thus preventing the transmission of error signals via SMI.

[0206] When a CE (Error Detection) generated by PCIe or Memory reaches a preset error threshold, the CPU pulls Error#0Pin to change the state of the GPIO that supports SCI (Signal Message Processing) trigger. SCI is then triggered, executing a pre-written pre-stored method in the BIOS code. This pre-stored method passes the GUID of the Prm Handler (PRM message handler) to the OS's ACPI subsystem (the driver that handles the OS's Advanced Configuration and Power Management Interface). The ACPI subsystem retrieves the Prm Handler's location in Memory from the PRMT Table (a structure used in driver development to define platform-related information) using the Prm Handler GUID, integrates information such as ContextBuffer / ParamBuffer, and passes it to the Prm Handler for CE processing. The ContextBuffer / ParamBuffer contains error-correcting buffer information.

[0207] In summary, in some embodiments of this application, the overall architecture of the server configuration system can primarily focus on RAS error handling, especially the correctable errors mentioned above. The BIOS can perform hardware initialization to enable RAS functionality and can also pass the required policies and data values ​​to the BMC RAS ​​handler. During this process, a handshake can be performed via platform management interface commands to exchange information. RAS processing functionality can be added to the ras-manager (RAS management) service within the OpenBMC (open-source BMC) service. The ras-manager service enables interaction with multiple components, such as RasLib, the platform management interface service, and MMBI. RasLib contains most of the RAS processing logic, enabling RAS processing and register read / write access using the platform-environmental control interface.

[0208] This application's server enables flexible switching of fault detection modes to adapt to the needs of different business scenarios. Furthermore, through the coordinated work of hardware and firmware, the timeliness and accuracy of fault detection can be improved, meaning fault detection can be relatively efficient and accurate. In addition, BIOS settings and platform management interfaces are provided to facilitate the configuration and management of fault detection modes, improving management convenience.

[0209] It should be understood that although the steps in the flowcharts of Figures 1 to 4 are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in Figures 1 to 4 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0210] Please refer to Figure 7, which is a schematic diagram of the structure of a server according to an embodiment of this application.

[0211] In some embodiments of this application, the server includes a server body 11 and a control system 12.

[0212] The server body 11 is a collection of basic components that enable server functions. It may include the server's graphics processing unit module, central processing unit module, heat dissipation module, power supply module, chassis, etc., which will not be described in detail here. For example, the server body 11 may include the BMC, CPLD, BIOS, PRW SW, etc., as described above.

[0213] The BIOS enables fault detection mode selection and configuration information distribution during the initial server startup phase. The CPLD, based on the BIOS configuration, controls the routing of fault signals (i.e., establishing a path between the fault signal terminal and the CPU or BMC). The BMC is responsible for collecting and reporting error information in RAS offload mode. The PRM SW can monitor and handle errors in PRM mode.

[0214] The control system 12 is located on the server body 11 and is used to implement the server configuration method as described in the above embodiments; or, to implement the server management method as described in the above embodiments.

[0215] Specifically, the control system 12 may include a basic input / output system, a baseboard management controller, an operating system, and a programmable logic device, as described above.

[0216] Thus, when the control system implements the server configuration method described above, it can be further refined as follows: In response to server startup, the basic input / output system initializes the first mode and the second mode, as well as the data space shared by the two, so that the first mode and the second mode call the initialization data in the data space for initialization; wherein, the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input / output system or the baseboard management controller takes one of the first mode and the second mode as the target detection mode and feeds back the target detection mode to the programmable logic device (PLD); in response to acquiring the target detection mode, the PLD selects the path between the fault signal terminal and the operating entity of the target detection mode, so as to use the target detection mode to perform fault detection on the server; wherein, the fault signal terminal is used to transmit correctable errors, and the operating entity is the baseboard management controller or the operating system.

[0217] For specific limitations regarding the control system, please refer to the limitations on server configuration and management methods described above, which will not be repeated here. Each module in the aforementioned control system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0218] In some embodiments of this application, a computer-readable instruction product is provided, including computer-readable instructions. When executed by a processor, the computer-readable instructions implement the steps of the server configuration method or server management method described above, which will not be repeated here.

[0219] In other words, when the computer-readable instructions are executed by the processor, the following steps are achieved: In response to server startup, the basic input / output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein, the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input / output system or the baseboard management controller takes one of the first mode and the second mode as the target detection mode and feeds back the target detection mode to the programmable logic device (PLD); in response to acquiring the target detection mode, the PLD selects the path between the fault signal terminal and the running entity of the target detection mode, so as to use the target detection mode to perform fault detection on the server; wherein, the fault signal terminal is used to transmit correctable errors, and the running entity is the baseboard management controller or the operating system.

[0220] Alternatively, when the computer-readable instructions are executed by the processor, the following steps can be achieved: monitoring the server's operating parameters; assessing whether the server is faulty based on the operating parameters using a target detection mode; wherein the target detection mode is activated using the server configuration method as described in any of the above embodiments.

[0221] Please refer to Figure 8, which is a schematic diagram of the structure of a computer device according to an embodiment of this application.

[0222] In some embodiments of this application, the computer device may be a server, and its internal structure diagram may be illustrated in Figure 8.

[0223] The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer-readable instructions, and database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions stored in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When the computer-readable instructions are executed by the processor, they implement a server configuration method or a server management method.

[0224] Those skilled in the art will understand that the structure illustrated in Figure 8 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.

[0225] In some embodiments of this application, a computer device is provided, including a memory and one or more processors. The memory stores computer-readable instructions, which, when executed by the one or more processors, cause the one or more processors to perform the following steps: In response to server startup, a basic input / output system initializes a data space shared by a first mode and a second mode, and calls initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on a baseboard management controller, and the second mode runs on an operating system; the basic input / output system or the baseboard management controller uses one of the first mode and the second mode as a target detection mode and feeds back the target detection mode to a programmable logic device (PLD); in response to acquiring the target detection mode, the PLD selects a path between a fault signal terminal and the operating entity of the target detection mode to perform fault detection on the server using the target detection mode; wherein the fault signal terminal is used to transmit correctable errors, and the operating entity is the baseboard management controller or the operating system.

[0226] In some embodiments of this application, when the processor executes computer-readable instructions, it may also perform the following steps: monitoring the server's operating parameters; assessing whether the server is faulty based on the operating parameters using a target detection mode; wherein the target detection mode is activated using the server configuration method as described in any of the above embodiments.

[0227] Please refer to Figure 9, which is a schematic diagram of the structure of a non-transitory computer-readable storage medium in one embodiment of this application.

[0228] In some embodiments of this application, the non-transitory computer-readable storage medium 20 is used to store instruction / program data 21, which can be executed to implement the server configuration method as described in any of the above embodiments, or the server management method as described in the above embodiments, and will not be described again here.

[0229] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are illustrative; for instance, the division of modules or units is a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or units may be electrical, mechanical, or other forms.

[0230] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0231] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0232] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium 20 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned computer-readable storage medium 20 includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks or optical disks, and servers.

[0233] In some embodiments of this application, a non-transitory computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed, the following steps are implemented: In response to server startup, the basic input / output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein, the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input / output system or the baseboard management controller takes one of the first mode and the second mode as the target detection mode and feeds back the target detection mode to the programmable logic device; in response to acquiring the target detection mode, the programmable logic device selects the path between the fault signal terminal and the running entity of the target detection mode to perform fault detection on the server using the target detection mode; wherein, the fault signal terminal is used to transmit correctable errors, and the running entity is the baseboard management controller or the operating system.

[0234] In some embodiments of this application, when the computer-readable instructions are executed by the processor, the following steps can also be implemented: monitoring the operating parameters of the server; evaluating whether the server is faulty based on the operating parameters using a target detection mode; wherein the target detection mode is activated using the server configuration method as described in any of the above embodiments.

[0235] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile computer-readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0236] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. The embodiments described above only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the concept of this application, and these all fall within the protection scope of this application.

Claims

1. A server configuration method, characterized in that, The server configuration method includes: In response to server startup, the basic input / output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein, the first mode runs on the baseboard management controller, and the second mode runs on the operating system; The basic input / output system or the baseboard management controller uses one of the first mode and the second mode as the target detection mode and feeds back the target detection mode to the programmable logic device; and In response to acquiring the target detection mode, the programmable logic device selects a path between the fault signal terminal and the operating entity of the target detection mode to perform fault detection on the server using the target detection mode; wherein, the fault signal terminal is used to transmit correctable errors, and the operating entity is the baseboard management controller or the operating system.

2. The server configuration method according to claim 1, characterized in that, The basic input / output system initializes the shared data space for both the first and second modes, including: The basic input / output system initializes the data to be shared; and The basic input / output system transmits the data to be shared to the preset data space within the baseboard management controller, and maps the data space to the memory setting information of the basic input / output system, so that the baseboard management controller and the basic input / output system can access the data space based on the memory setting information and call the data to be shared to initialize the first mode and the second mode.

3. The server configuration method according to claim 2, characterized in that, The data to be shared includes at least one of the error correction handling strategy parameters and the memory topology.

4. The server configuration method according to claim 1, characterized in that, The initialization of the first mode and the second mode includes: The basic input / output system queries the baseboard management controller to see if it supports the first mode; In response to feedback from the substrate management controller that it supports the first mode, the basic input / output system sends a first combination of parameters for the first mode to the substrate management controller; and The basic input / output system creates the operating conditions for the second mode in response to the operating system supporting the second mode.

5. The server configuration method according to claim 4, characterized in that, The basic input / output system sends the first parameter combination of the first mode to the baseboard management controller, including: The basic input / output system unloads the first mode and installs the first mode onto the baseboard management controller; and sends the first parameter combination to the baseboard management controller; wherein, the first parameter combination includes functional support capability parameters, correctable error handling strategy parameters, bus topology, and memory topology; and In response to the first mode being the target detection mode, the basic input / output system sends an enable command to the baseboard management controller to enable the execution of a correctable error handling function; wherein the enable command carries a function mask, the function mask being used to identify the first mode function to be activated.

6. The server configuration method according to claim 5, characterized in that, The basic input / output system unloads the first mode and installs the first mode to the baseboard management controller, including: The basic input / output system retains target configuration control; wherein, the target configuration control is used to control enabling or disabling the first mode; and The basic input / output system unloads the first mode function and sends the associated data of the first mode function to the baseboard management controller.

7. The server configuration method according to claim 5, characterized in that, The substrate management controller enables corrective error handling functions, including: The substrate management controller initializes the library initialization routine for the first mode; and The baseboard management controller identifies the first mode function to be activated carried by the function mask and initializes the calling interface of the first mode function.

8. The server configuration method according to claim 4, characterized in that, The operating conditions for creating the second mode include: The basic input / output system initializes the message processor for the second mode; and The basic input / output system registers the message processor with the configuration protocol of the second mode, so that when the server is under the control of the operating system, it can invoke the message processor based on the registration information in the configuration protocol.

9. The server configuration method according to claim 8, characterized in that, The conditions for creating the second mode also include: The basic input / output system is configured with a transmission port associated with the second mode so that a corresponding code command is executed when the associated register signal is triggered; and The basic input / output system is configured with a second parameter combination for initialization of the second mode; wherein the second parameter combination includes correctable error handling strategy parameters, bus topology, and memory topology.

10. The server configuration method according to claim 1, characterized in that, The step of using either the first mode or the second mode as the target detection mode includes: In response to the initial phase of server startup, the basic input / output system reads the fault detection settings options; obtains the preset fault detection mode identified by the settings options, and uses it as the target detection mode; and / or, In response to the initial phase of startup via the server, the baseboard management controller acquires a switching instruction; parses the fault detection mode to be switched identified by the switching instruction, and uses it as the target detection mode.

11. The server configuration method according to claim 1, characterized in that, The substrate management controller uses one of the first mode and the second mode as the target detection mode, including: The baseboard management controller acquires the status parameters of the server; The baseboard management controller analyzes the status parameters to assess whether to switch to the fault detection mode of the target detection mode; and In response to determining the fault detection mode to switch to the target detection mode, the baseboard management controller controls the programmable logic device to select other fault detection modes as the new target detection mode; wherein, the other fault detection modes refer to fault detection modes that are not currently used as the target detection mode.

12. The server configuration method according to claim 11, characterized in that, The baseboard management controller analyzes the status parameters to evaluate whether to switch to the fault detection mode of the target detection mode, including: In response to the first mode being used as the target detection mode, the baseboard management controller obtains the resource utilization parameters of the baseboard management controller; compares the resource utilization parameters with the resource utilization threshold; and in response to the resource utilization parameters reaching the resource utilization threshold, determines to switch to the fault detection mode of the target detection mode. The programmable logic device selects other fault detection modes as new target detection modes, including: The programmable logic device will select the path between the fault signal terminal and the operating system, and switch to the second mode as the target detection mode.

13. The server configuration method according to claim 1, characterized in that, The path between the fault selection signal terminal and the main body of the target detection mode also includes: the basic input / output system triggering a system management interrupt, and providing fault feedback through the interrupt signal; The path between the fault signal terminal and the main body of the target detection mode also includes: The basic input / output system sends an enable command for the target detection mode to the baseboard management controller; in response to receiving the enable command, the baseboard management controller enables the target detection mode through a preset interface associated with the target detection mode.

14. The server configuration method according to claim 1, characterized in that, When the first mode is used as the target detection mode, the step of using the target detection mode to perform fault detection on the server includes: The first mode detects whether the server has the correctable error by means of error interruption handling or polling; in response to the existence of the correctable error, the first mode constructs an error message by means of a preset error handling library and passes it to the operating system; And / or, When the second mode is used as the target detection mode, the fault detection of the server using the target detection mode includes: The second mode configures the backup interrupt as a fault signal option and enables it, and detects whether the server has the correctable error through a system interrupt trigger; in response to the existence of the correctable error, the second mode executes error handling logic through a message processor and feeds it back to the baseboard management controller.

15. The server configuration method according to claim 14, characterized in that, In response to the existence of the correctable error, the second mode executes error handling logic via a message processor, including: In response to the number of correctable errors reaching an error threshold, the central processing unit changes the state of the fault signal terminal option, causing the system interrupt to trigger execution of the pre-stored method within the basic input / output system; wherein, the pre-stored method is used to pass the identifier of the message processor to the operating system; and The operating system queries the memory location of the message processor that matches the identifier, integrates the correctable error buffer information, and passes it to the message processor so that the message processor can process the correctable error.

16. A server management method, characterized in that, The server management method includes: Monitor server operating parameters; and The server is assessed for malfunction based on the operating parameters using a target detection mode; wherein the target detection mode is activated using the server configuration method as described in any one of claims 1-15.

17. A server, characterized in that, The server includes: Server body; A control system, located on the server body, is used to implement the server configuration method as described in any one of claims 1-15; or to implement the server management method as described in claim 16.

18. A computer-readable instruction product, comprising computer-readable instructions, characterized in that, When executed by a processor, the computer-readable instructions implement the steps of the server configuration method of any one of claims 1-15; or, implement the server management method of claim 16.

19. A computer device comprising a memory and one or more processors, the memory storing computer-readable instructions which, when executed by the one or more processors, cause the one or more processors to perform the steps of the server configuration method as claimed in any one of claims 1 to 15; or to implement the server management method as claimed in claim 16.

20. A non-transitory computer-readable storage medium storing computer-readable instructions thereon, characterized in that, When the computer-readable instructions are executed by a processor, they implement the steps of the server configuration method as described in any one of claims 1 to 15; or, implement the server management method as described in claim 16.