Server configuration and management method, server, program product, device and medium
By initializing multiple fault detection modes and selecting target detection modes when the server is started, and using programmable logic to perform fault detection, the problem that traditional fault detection strategies are difficult to adapt to efficient and accurate fault detection is solved, the accuracy and timeliness of server fault detection are achieved, and the server performance is improved.
Patent Information
- Application Number
- CN202510230755.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-28
AI Technical Summary
Traditional static fault detection strategies are difficult to adapt to the current needs of efficient and accurate fault detection, especially when server applications are complex and diversified.
A server configuration method and a server management method are provided. By initializing multiple fault detection modes at the server startup and selecting one of them as the target detection mode, a programmable logic gates the path between the fault signal end and the operating body of the target detection mode to realize fault detection.
It realizes the accuracy and timeliness of server fault detection, improves server performance, and can flexibly select suitable fault detection modes according to current fault detection requirements.
Smart Images

Figure CN119718473B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a server configuration method, a server management method, a server, a computer program product, a computer device, and a computer-readable storage medium. Background Art
[0002] A server is an Internet technology device that can provide computing power and run software applications in a network environment. It can usually provide computing or application services to other clients such as personal computers and smart phones in the network. For example, a server can have the ability to respond to service requests, provide services, and guarantee services.
[0003] In order to improve the stability of server operation, server fault detection can be performed through static fault detection strategies. However, with the complexity and diversification of server applications, traditional static fault detection strategies are difficult to adapt to the current needs of efficient and accurate fault detection. Summary of the invention
[0004] Based on this, it is necessary to provide a server configuration method, a server management method, a server, a computer program product, a computer device and a computer-readable storage medium to address the above technical problems, which are compatible with multiple server fault detection modes to help optimize server performance.
[0005] On the one hand, a server configuration method is provided, and the server configuration method includes: in response to server startup, a basic input and output system initializes a data space shared by a first mode and a second mode, and calls initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on a baseboard management controller, and the second mode runs on an operating system; the basic input and output system or the baseboard management controller uses one of the first mode and the second mode as a target detection mode, and feeds back the target detection mode to a programmable logic device; in response to obtaining the target detection mode, the programmable logic device selects a path between a fault signal terminal and an operating subject of the target detection mode, so as to perform fault detection on the server using the target detection mode; wherein the fault signal terminal is used to transmit correctable errors, and the operating subject is the baseboard management controller or the operating system.
[0006] On the other hand, a server management method is provided, which includes: monitoring operating parameters of the server; evaluating whether the server is faulty based on the operating parameters using a target detection mode; wherein the target detection mode is activated using a server configuration method as in any of the above embodiments.
[0007] On the other hand, a server is provided, the server comprising: a server body and a control system; the control system is arranged in the server body, and is used to implement the following steps: in response to the server startup, the basic input and output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input and output system or the baseboard management controller uses one of the first mode and the second mode as the target detection mode, and feeds back the target detection mode to the programmable logic device; the programmable logic device selects the path between the fault signal end and the operating subject of the target detection mode in response to obtaining the target detection mode, so as to perform fault detection on the server using the target detection mode; wherein the fault signal end is used to transmit correctable errors, and the operating subject is the baseboard management controller or the operating system. Or, the following steps are implemented: monitoring the operating parameters of the server; based on the operating parameters, using the target detection mode to evaluate whether the server is faulty; wherein the target detection mode is activated using the server configuration method in any of the above embodiments.
[0008] In another aspect, a computer program product is provided, comprising a computer program, which implements the following steps when executed by a processor: in response to the server startup, the basic input and output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input and output system or the baseboard management controller uses one of the first mode and the second mode as the target detection mode, and feeds back the target detection mode to the programmable logic device; the programmable logic device selects the path between the fault signal end and the operating subject of the target detection mode in response to obtaining the target detection mode, so as to use the target detection mode to perform fault detection on the server; wherein the fault signal end is used to transmit correctable errors, and the operating subject is the baseboard management controller or the operating system. Or, implement the following steps: monitoring the operating parameters of the server; based on the operating parameters, using the target detection mode to evaluate whether the server is faulty; wherein the target detection mode is activated using the server configuration method in any of the above embodiments.
[0009] In another aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: in response to the server startup, the basic input and output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input and output system or the baseboard management controller uses one of the first mode and the second mode as the target detection mode, and feeds back the target detection mode to the programmable logic device; the programmable logic device selects the path between the fault signal terminal and the operating subject of the target detection mode in response to obtaining the target detection mode, so as to perform fault detection on the server using the target detection mode; wherein the fault signal terminal is used to transmit correctable errors, and the operating subject is the baseboard management controller or the operating system. Or, the following steps are implemented: monitoring the operating parameters of the server; and evaluating whether the server is faulty using the target detection mode based on the operating parameters; wherein the target detection mode is activated using the server configuration method in any of the above embodiments.
[0010] In another aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented: in response to the server startup, multiple fault detection modes are initialized respectively; one of the multiple fault detection modes is used as a target detection mode; a fault detection environment of the target detection mode is configured to activate the target detection mode. Or, the following steps are implemented: monitoring the operating parameters of the server; based on the operating parameters, using the target detection mode to evaluate whether the server is faulty; wherein the target detection mode is activated using the server configuration method in any of the above embodiments.
[0011] The above-mentioned server configuration method, server management method, server, computer program product, computer device and computer-readable storage medium can configure the server to be compatible with multiple fault detection modes, that is, at least including a first mode running on a baseboard management controller and a second mode running on an operating system. After the server is started, the multiple fault detection modes can be initialized respectively, so that the fault detection mode has a basic application environment. In this way, one of the first mode and the second mode can be selected as the current target detection mode of the server, and the target detection mode can be configured to achieve activation, and the target detection mode can be used to perform fault detection on the server. At the same time, the data space shared by the first mode and the second mode in the present application is equivalent to realizing the normalization of the first mode and the second mode, which can be beneficial to ensure the synchronization of the data obtained by the two. In addition, in the present embodiment, the first mode and the second mode are hardware isolated by a programmable logic device, so as to facilitate the independent operation and reliability of the first mode and the second mode. It can be seen that the present application is compatible with multiple fault detection modes, so according to the current fault detection requirements of the server, the fault detection mode can be flexibly selected from a variety of fault detection modes as the target detection mode, thereby improving the accuracy and timeliness of server fault detection by improving the adaptability of the target detection mode to the actual business scenarios, thereby improving the server performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a flowchart of an embodiment of the application server management method;
[0013] Figure 2 It is a flowchart of the first embodiment of the application server configuration method;
[0014] Figure 3 It is a flowchart of the second embodiment of the application server configuration method;
[0015] Figure 4 It is a flowchart of the third embodiment of the application server configuration method;
[0016] Figure 5 It is a structural diagram of an embodiment of the server management system of the present application;
[0017] Figure 6 It is a flowchart of the fourth embodiment of the application server configuration method;
[0018] Figure 7 It is a structural diagram of an embodiment of the server of the present application;
[0019] Figure 8 It is a structural schematic diagram of an embodiment of a computer device of the present application;
[0020] Fig. 9 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0022] In order to solve the technical problem that the fault detection strategies existing in the related technologies are difficult to adapt to the current demand for efficient and accurate fault detection, the present application provides a server configuration method, a server management method, a server, a computer program product, a computer device and a computer-readable storage medium. The server configuration method includes: in response to the server startup, initializing a plurality of fault detection modes respectively; using one of the plurality of fault detection modes as the target detection mode; configuring the fault detection environment of the target detection mode to activate the target detection mode. In this way, the present application can be compatible with a plurality of server fault detection modes, so as to improve the accuracy and timeliness of server fault detection, thereby improving server performance. The specific working principle of the present application is described below with examples.
[0023] See also Figure 1 , Figure 1 It is a flow chart of an embodiment of the server management method of the present application.
[0024] S101: Monitor the operating parameters of the server.
[0025] In this embodiment, the operating parameters include relevant data generated during the operation of the server, etc.
[0026] For example, the server can be sensed by sensors and the sensing data can be used as one of the operating parameters; resource usage information can also be used as one of the operating parameters; or the computing efficiency, response efficiency, etc. of the server can be monitored and computing power-related data can be used as one of the operating parameters, without limitation here.
[0027] S102: Evaluate whether the server is faulty using a target detection mode based on operating parameters; wherein the target detection mode is activated using a server configuration method.
[0028] In this embodiment, the target detection mode may be used to detect whether a server currently has a fault; and / or to analyze whether the server has a fault trend.
[0029] In this embodiment, the server can be compatible with multiple fault detection modes, and the target detection mode is one of the multiple fault detection modes.
[0030] Among them, the server can be compatible with multiple fault detection modes, and one of them can be used as the target detection mode. In other words, the fault detection mode used as the target detection mode can be flexibly selected, that is, the fault detection mode that adapts to the current business scenario and application scenario can be selected as the target detection mode, and the fault detection mode used as the target detection mode can be switched when the fault detection effect of the current target detection mode is not ideal, which is conducive to improving the adaptability of the target detection mode to the current scenario, and is also conducive to improving the timeliness and accuracy of fault detection, thereby ensuring the operational reliability and stability of the server, and improving the server performance.
[0031] See also Figure 2 , Figure 2 It is a flowchart of the first embodiment of the server configuration method of the present application.
[0032] S201: In response to server startup, the basic input and output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on the baseboard management controller, and the second mode runs on the operating system.
[0033] In this embodiment, the first mode and the second mode are two fault detection modes, in other words, the fault detection mode of the server includes at least the first mode and the second mode, wherein the first mode runs on the baseboard management controller, and the second mode runs on the operating system.
[0034] As the name implies, the fault detection mode has a fault detection function. In this embodiment, multiple fault detection modes may be pre-set, and the fault detection principles of the multiple fault detection modes may be different, that is, the fault detection principles of the first mode and the second mode may be different. The detailed working principles of the first mode and the second mode will be described in detail later, and will not be repeated here.
[0035] When the server starts, various fault detection modes including the first mode and the second mode can be initialized so that the server can basically support the fault detection mode. In other words, the basic configuration can be performed for the subsequent use of the fault detection mode as a target detection mode, which is conducive to simplifying the subsequent configuration work when setting it as a target detection mode, and is also conducive to improving the switching efficiency of the fault detection mode as a target detection mode.
[0036] At the same time, in this embodiment, a data space is also created. As the name implies, the data space can be used to store data. The data space created in this embodiment can be in a shared state, that is, both the first mode and the second mode can obtain data in the data space. In this embodiment, the basic input and output system can also initialize the data space shared by the first mode and the second mode, so that when in the first mode or the second mode, the data space can be accessed to obtain data therein.
[0037] S202: The basic input / output system or the baseboard management controller uses one of the first mode and the second mode as a target detection mode, and feeds back the target detection mode to the programmable logic device.
[0038] In this embodiment, the basic input / output system can use one of the first mode and the second mode as the target detection mode. Or, the baseboard management controller can use one of the first mode and the second mode as the target detection mode. That is to say, in this embodiment, when selecting / controlling the fault detection mode as the target detection mode, it can be controlled by the basic input / output system or the baseboard management controller. For example, in the initial stage of server startup, the basic input / output system can select the fault detection mode as the target detection mode; after the server startup is completed, the baseboard management controller can select the fault detection mode as the target detection mode.
[0039] When one of the first mode and the second mode is selected as the target detection mode, the target detection mode can be fed back to the programmable logic device so that the programmable logic device can adaptively control the enabling of the target detection mode.
[0040] S203: In response to acquiring the target detection mode, the programmable logic device selects the path between the fault signal end and the operating subject of the target detection mode to perform fault detection on the server using the target detection mode; wherein the fault signal end is used to transmit correctable errors, and the operating subject is a baseboard management controller or an operating system.
[0041] In this embodiment, the programmable logic device can be connected to the operating entities of the first mode and the second mode respectively, that is, the programmable logic device can be connected to the baseboard management controller and the operating system respectively. In addition, the programmable logic device can also be connected to the fault signal terminal to transmit the error information fed back by the first mode and the second mode to the fault signal terminal.
[0042] The programmable logic device responds to the target detection mode of obtaining feedback from the baseboard management controller or the basic input and output system, and can select the path between the fault signal terminal and the operating subject of the target detection mode. That is, when the first mode is used as the target detection mode, the programmable logic device can select the path between the fault signal terminal and the baseboard management controller; when the second mode is used as the target detection mode, the programmable logic device can select the path between the fault signal terminal and the operating system.
[0043] It can be seen that in this embodiment, the server can be configured to be compatible with multiple fault detection modes. That is, the server can be compatible with at least the first mode running on the baseboard management controller and the second mode running on the operating system. After the server is started, the multiple fault detection modes can be initialized respectively, so that the fault detection mode has a basic application environment. The first mode and the second mode can share the data space, which is equivalent to realizing the normalization of the first mode and the second mode, which can be beneficial to ensure the synchronization of the data obtained by the two. In addition, in this embodiment, the first mode and the second mode are hardware isolated by a programmable logic device to facilitate the independent operation and reliability of the first mode and the second mode. In this way, one of the first mode and the second mode can be selected as the current target detection mode of the server, and the target detection mode can be configured to achieve activation, and the target detection mode can be used to perform fault detection on the server. In this way, according to the current fault detection requirements of the server, the fault detection mode as the target detection mode can be flexibly selected from a variety of fault detection modes, so that by improving the adaptability of the target detection mode to the actual business scenario, the accuracy and timeliness of server fault detection can be improved, and the server performance can be improved.
[0044] In other words, after the server is started, multiple fault detection modes can be initialized respectively, so that the fault detection mode has a basic application environment. In this way, one of the multiple fault detection modes can be selected as the current target detection mode of the server, and the target detection mode can be configured to activate it, and the target detection mode can be used to perform fault detection on the server. It can be seen that multiple fault detection modes are compatible in this application, so according to the current fault detection needs of the server, a fault detection mode can be flexibly selected from multiple fault detection modes as the target detection mode, thereby improving the accuracy and timeliness of server fault detection by improving the adaptability of the target detection mode to the actual business scenario, and then improving server performance.
[0045] Furthermore, when initializing a plurality of fault detection modes respectively, the data and status of the initialization of the fault detection mode can be configured so that the server can support the fault detection function of the fault detection mode. Moreover, during the initialization process, when the first mode and the second mode can share the data in the data space, and the baseboard management controller and the operating system obtain data from the data space to initialize the first mode and the second mode, the data required for the fault detection mode they are running can be obtained. Compared with passive data reception, even a shared data space is conducive to isolating the transmission of initialization data between fault detection modes, and the first mode and the second mode can also be relatively isolated through isolation such as programmable logic devices. In this way, the reliable transmission of data in each fault detection mode can be guaranteed, and the reliability of each other's fault detection function configuration affected by data transmission interference can be reduced. Moreover, it is also conducive to achieving the relative independence of the fault detection functions between each fault detection mode and reducing the risk of mutual interference.
[0046] As described in the embodiment, the fault detection mode may include at least a first mode and a second mode. For example, the first mode performs correctable error processing in the baseboard management controller; the second mode performs error processing through the platform operation mechanism. Furthermore, the fault detection mode may also include other fault detection modes such as a static fault detection mode, which are not limited here.
[0047] For example, the first mode may be RAS offload (correctable error handling acceleration function), etc. RAS stands for Reliability Availability and Serviceability, which is one of the important functions of a server.
[0048] The second mode may be, for example, PRM (Platform Runtime Mechanism).
[0049] See also Figure 3 , Figure 3 It is a flow chart of the second embodiment of the application server configuration method.
[0050] S301: Read the device option for fault detection and provide fault feedback via an interrupt signal.
[0051] In this embodiment, in response to being in the initial stage of server startup, the basic input and output system can read the setting options of fault detection to obtain the preset fault detection mode identified by the setting options, which can be used as the target detection mode.
[0052] The server startup process may include an initial stage and a transition stage. The initial stage (i.e., the Post stage) indicates that the server system is controlled by the basic input and output system. For example, the basic input and output system may be a BIOS (Basic Input Output System). The transition stage (i.e., the Runtime stage) indicates that the server system is controlled by an operating system. For example, the operating system may be an OS (Operating System).
[0053] In other words, in this embodiment, setting options can be added in the server to support the pre-setting of the server's default fault detection mode, so that the server's target detection mode setting can be implemented even without external intervention and without adaptive matching strategies.
[0054] For example, a Fault Diagnosis Select option can be created in advance, through which the default fault detection mode, i.e., the preset fault detection mode, is selected. In this process, the first mode or the second mode can be set as the preset fault detection mode, i.e., RAS offload or PRM. In addition, faults in the Post stage can be reported through the traditional SMI (System Management Interrupt).
[0055] That is to say, in the initial stage of server startup, before using the target detection mode for fault detection, the basic input and output system can trigger a system management interrupt and provide fault feedback through an interrupt signal.
[0056] S302: Check whether the baseboard management controller supports the first mode.
[0057] In this embodiment, when it is determined that the baseboard management controller supports the first mode, step S303 is executed; when it is determined that the baseboard management controller does not support the first mode, step S304 is executed.
[0058] That is to say, in the process of initializing the fault detection mode, it is also possible to verify whether the drive control module of the fault detection mode supports the functional implementation of the fault detection mode, and in response to the drive control module supporting the functional implementation of the fault detection mode, send the initialization data of the fault detection mode to the drive control module. In response to the drive control module not supporting the fault detection mode, it is also unnecessary to initialize the unsupported fault detection mode, thereby improving the initialization efficiency of the fault detection mode and marking that the fault detection mode cannot be used as a target detection mode. In layman's terms, a handshake can be made with the BMC (baseboard management controller) to confirm whether the BMC supports the RASoffload function.
[0059] In this embodiment, a server configuration method with good generalization performance is provided, and it is expected that the server can be compatible with multiple fault detection functions, and ideally the server can be compatible with all fault detection modes. At the same time, in this embodiment, it is further considered that in some cases, the server may not support the corresponding fault detection function, so it is possible to pre-check whether the server can support various fault detection modes to reduce the situation where the unsupported fault detection mode is still enabled when the drive control module corresponding to the fault detection mode does not support its implementation, resulting in invalid fault detection. That is, it can reduce the risk that the target detection mode cannot detect server faults, thereby improving the reliability of the server configuration method, and ensuring that the target detection mode can run reliably to detect server faults.
[0060] Furthermore, whether each fault detection mode is supported by the corresponding drive control module can be determined in different stages, which will not be described in detail here.
[0061] In this embodiment, it is possible to verify whether the baseboard management controller supports the first mode at the initial stage of server startup; wherein the server system is controlled by the basic input and output system at the initial stage. In response to the baseboard management controller supporting the first mode and the first mode being the target detection mode, an enable instruction is sent to the baseboard management controller to enable the baseboard management controller to enable the correctable error handling function.
[0062] S303: Sending a first parameter combination initialized in the first mode to a baseboard management controller.
[0063] In this embodiment, a first parameter combination for initialization of the first mode may be configured.
[0064] Among them, the first parameter combination includes function support capability parameters, correctable error handling strategy parameters, bus topology, and memory topology. The first parameter combination is sent to the baseboard management controller so that the baseboard management controller can initialize the functions and support configurations related to the first mode. In layman's terms, RAS offload related data can be prepared and sent to the BMC. In this embodiment, the RAS function can be offloaded to the OS for processing, which is beneficial to improving the stability and reliability of the server system. It is also possible to further perform real-time dynamic updates on SMM drivers (a firmware driver) to ensure that the server system remains up to date and reduce security issues caused by software omissions, outdated drivers, and other factors.
[0065] S304: Initialize the message processor of the second mode.
[0066] In this embodiment, the message processor of the second mode may be initialized.
[0067] The transmission port associated with the second mode can be configured so that the corresponding code command is executed when the register signal associated with it is triggered. A second parameter combination for initializing the second mode can also be configured; wherein the second parameter combination includes a correctable error handling strategy parameter, a bus topology, and a memory topology. In other words, a PRM handler can be initialized. Wherein, a handler is a message processing device, that is, the message processor in this embodiment can include a PRM handler.
[0068] S305: Send the configuration information of the target detection mode to the programmable logic device, so as to configure the control module connected to the fault signal terminal based on the configuration information.
[0069] In this embodiment, the target detection mode is taken as the second mode as an example for explanation.
[0070] That is, in this embodiment, the configuration information of the target detection mode can be obtained, and the configuration information can be sent to the programmable logic device, so that the programmable logic device configures the control module connected to the fault signal terminal based on the configuration information; wherein the control module includes a central processing unit and a baseboard management controller.
[0071] In other words, the fault detection mode of BIOS setup can be obtained as the target detection mode, and the corresponding configuration information can be sent to the CPLD through VGPIO, etc. In this way, the CPLD can be set according to the fault detection mode transmitted by the BIOS, that is, the target detection mode configuration fault signal end is connected to the CPU or BMC.
[0072] Among them, VGPIO stands for voltage-controlled GPIO (General Purpose Input / Output).
[0073] CPLD (Complex Programmable logic device) is a programmable logic device with high density characteristics.
[0074] The CPU (Central Processing Unit) serves as the computing and control core of the server system.
[0075] The fault signal terminal can be defined as an "error pin", that is, the CPLD can configure error pin #0 to the CPU or BMC.
[0076] Optionally, step S301 to step S305 may occur in the post stage of server startup, that is, the initial stage of server startup. The post stage / initial stage means that the server BIOS is controlled during the startup stage.
[0077] When the first mode is used as the target detection mode as described above, an enable command may be sent to the baseboard management controller. Specifically, the BIOS may send a command such as "Send Start Cmd w / Offload Features" to the BMC, which indicates enabling features related to the first function. In this way, when the first mode is activated, the BMC may poll the errorpin to collect error information and report it.
[0078] S306: Construct management software of the second mode.
[0079] In this embodiment, the management software is associated with the operating system, and the management software is also used to monitor and handle server failures when the second mode is used as the target detection mode. The management software can be PRM SW (software), etc., which is not limited here.
[0080] The management software of the second mode may be constructed to trigger the management software to obtain the function support information.
[0081] As explained above, the verification of whether the driver control module of the first mode supports it can be done in the post stage. The verification of whether the driver control module of the second mode supports it can be done in the runtime stage in this step, and the runtime stage refers to the stage in which the OS (Operating System) controls the server system during the server startup process. Among them, the driver control modules of the first mode and the second mode are different. In other words, the driver control modules of at least some of the fault detection modes in the multiple fault detection modes can be the same, and the driver control modules of the fault detection modes are allowed to be different.
[0082] S307: Evaluate whether the operating system supports the second mode.
[0083] In this embodiment, when the evaluation determines that the operating system supports the second mode, step S308 is executed; when the evaluation determines that the operating system does not support the second mode, step S309 is executed.
[0084] Specifically, in response to the second mode being used as the target detection mode, the function support information of the operating system can be obtained before entering the server system to evaluate whether the operating system supports the second mode. In response to determining that the operating system supports the second mode, the second mode setting is further performed. The second mode setting includes controlling the correctable error to be fed back through the second mode.
[0085] In other words, the PRM SW can be triggered before booting into the system to obtain the OS support for PRM. Boot can mean the process of loading the operating system from a storage device when the server power is turned on, that is, the system is in the booting stage.
[0086] It should be noted that, in this embodiment, the target detection mode preset in the BIOS is used as the second mode for example. When the preset target detection mode is the first mode, after executing step S305, an enable signal can be sent to the baseboard management controller to activate the first mode, and the baseboard management controller polls the error pin to collect fault information.
[0087] S308: Setting a second mode, and performing fault detection through the second mode.
[0088] In this embodiment, if the system supports PRM, PRM settings may be performed to control ce (correctable error) to be reported through PRM.
[0089] S309: Feedback to the programmable logic device that the operating system does not support the second mode, so as to update the target detection mode to the first mode.
[0090] In this embodiment, in response to determining that the operating system does not support the second mode, a path between the fault signal terminal and the baseboard management controller can be enabled; wherein the fault signal terminal is used to transmit a correctable error. The correctable error is controlled to be fed back through the first mode, an enable signal is sent to the baseboard management controller, and the target detection mode is updated to the first mode.
[0091] In other words, if the OS does not support PRM, a command can be sent to the CPLD to control the error pin to switch to the BMC, and the CE can report through the RAS offload.
[0092] If so, in this embodiment, when the operating system does not support the second mode and the second mode is used as the target detection mode, it can be discovered in time that the second mode is not supported, and it can be considered that it cannot be reliably implemented or even cannot implement the fault detection function. It should be noted that here it is considered that the second mode may not be implemented or cannot be reliably implemented, which does not mean that the second mode must not be implemented. At this time, the target detection mode can be switched in time, so that the first mode that can be supported by its drive control module is used as the target detection mode, which is beneficial to ensure the reliable implementation of the server fault detection function.
[0093] In an alternative embodiment, if both the first mode and the second mode are not supported, static fault detection and other methods may be used to perform fault detection on the server and / or report related situations.
[0094] S310: Sending an enable signal to a baseboard management controller to enable the baseboard management controller to poll the state of the fault signal terminal to collect fault information.
[0095] In this embodiment, the BIOS may also send an enable signal to the BMC to enable the RAS offload function. When the RAS offload is activated, the BMC may poll the error pin status as described above to collect and report error information.
[0096] In this case, the server fault detection is performed by using the first mode or the second mode as the target detection mode until the reporting is completed, which may mean that the server is shut down or stopped being used, etc., which is not limited here.
[0097] See also Figure 4 , Figure 4 It is a flowchart of the third embodiment of the server configuration method of the present application.
[0098] S401: Monitoring platform management interface.
[0099] In this embodiment, the platform management interface is used to communicate with the outside of the server. In layman's terms, a user can input control instructions through the platform management interface.
[0100] For example, the platform management interface may be an intelligent platform management interface or the like.
[0101] S402: Obtain a switching instruction.
[0102] In this embodiment, in response to the initial stage of being started by the server, that is, when the operating system takes over the control authority, the server, host, operating system, etc. can obtain the external input switching instruction so that the baseboard management controller can obtain the switching instruction. The switching instruction can come from outside the server, such as the platform management interface described in step S401.
[0103] In this way, the baseboard management controller can parse the fault detection mode to be switched identified by the switching instruction and use it as the target detection mode. In this way, the present embodiment can also support external input switching instructions to switch the fault detection mode currently used as the target detection mode of the server, so that the target detection mode of the server can adapt to the real-time needs of the user, and in this embodiment, the configuration is performed in the form of instructions such as pre-related intelligent platform management interface instructions, that is, in response to passing the initial stage, the fault detection mode used as the target detection mode is allowed to be switched through the platform management interface, so as to achieve simple and free switching between multiple fault detection modes.
[0104] In other words, it is possible to provide users with a solution for actively switching fault detection modes, which can help improve the efficiency of identifying server faults. It allows users to actively switch to the corresponding fault detection mode as the target detection mode according to the business usage scenario and customized strategy, thereby reducing the fault handling time and thus helping to improve business continuity. It is also possible to improve the effectiveness of server resource utilization and reduce the risk of resource waste through timely adjustment of hardware registers, which is conducive to improving overall operational efficiency.
[0105] For example, the PRM may be enabled via a predefined interface.
[0106] S403: Analyze the fault detection mode to be switched, which is identified by the switching instruction, and use it as the target detection mode.
[0107] S404: Specify a transmission port, and trigger a system management interrupt through the specified transmission port.
[0108] In this embodiment, VGPIO may be designated to allow SMI to be triggered through VGPIO.
[0109] S405: Configure the target detection mode indicated by the switching instruction.
[0110] In this embodiment, taking the indicated target detection mode as PRM as an example, a PRM enable handler may be entered.
[0111] S406: Check whether the drive control module supports the target detection mode.
[0112] In this embodiment, when it is determined that the driving control module supports the target detection mode indicated by the switching instruction, step S407 is executed; when it is determined that the driving control module does not support the target detection mode indicated by the switching instruction, step S409 is executed.
[0113] S407: Feedback the target detection mode to the programmable logic device to enable the path between the fault signal terminal and the drive control module.
[0114] S408: Perform fault detection using the new target detection mode.
[0115] In this embodiment, it can be determined whether the OS supports PRM. If supported, a command can be sent to the CPLD to switch the errorpin to connect to the CPU and execute PRM to report ce.
[0116] S409: Exit the system management interruption and restore the original target detection mode to perform fault detection.
[0117] In this embodiment, for example, it can be determined whether the OS supports PRM. If the OS does not support PRM, the SMI can be exited and an error can be returned, and ce reporting can be performed through the original target detection mode.
[0118] In this embodiment, the original target detection mode is the first mode or other fault detection mode as an example. Optionally, the functionality of the switching instruction can be simplified, that is, the switching instruction only has a switching function to reduce the data volume of the switching instruction. That is, when the switching instruction is received, the fault detection mode as the target detection mode is switched. For example, taking the fault detection mode including the first mode and the second mode, and the original target detection mode being the first mode as an example, when the switching instruction is received, the target detection mode is switched to the second mode.
[0119] In layman's terms, when an SMI is triggered by an interface such as VGPIO, no matter the current target detection mode is the first mode or the second mode, it can be switched to the other mode. That is, if the current target detection mode is PRM, it can be switched to RAS offload; if the current target detection mode is RAS offload, it can be switched to PRM.
[0120] Optionally, it can also support adaptive identification of server status to spontaneously switch to a fault detection mode that better matches the current server status. Active switching of fault detection modes is beneficial to effectively reduce server operation risks and reduce the workload of operation and maintenance personnel. Through real-time monitoring and data analysis, it is beneficial to pre-detect potential server failures for corresponding preventive maintenance, reduce the risk of actual server failures, and thus help reduce server downtime caused by failures and improve business continuity.
[0121] Specifically, the baseboard management controller may obtain the status parameters of the server. In response to obtaining the status parameters of the server, the baseboard management controller may analyze the status parameters to evaluate whether to switch the fault detection mode as the target detection mode. In response to determining to switch the fault detection mode as the target detection mode, the baseboard management controller may control the programmable logic device to select other fault detection modes as new target detection modes; wherein the other fault detection modes represent fault detection modes that are not currently the target detection mode.
[0122] For example, as described above, the fault detection mode may include a first mode and a second mode.
[0123] In response to the target detection mode being the first mode, the baseboard management controller or other modules such as the central processing unit can obtain the resource utilization parameters of the baseboard management controller. To compare the resource utilization parameters with the resource utilization threshold. In response to the resource utilization parameters reaching the resource utilization threshold, it can be considered that there is a certain burden on the baseboard management controller to perform fault detection, and it is determined to switch to the fault detection mode as the target detection mode, so it can be switched to the second mode as the target detection mode. That is, the programmable logic switches the path between the fault signal terminal and the operating system to the second mode as the target detection mode.
[0124] That is to say, in this embodiment, considering that the baseboard management controller is one of the important control modules of the server, it can usually use sensors to monitor the server. In order to ensure that the basic functions of the baseboard management control, namely the monitoring function of the server, can be reliably realized, when the resource utilization rate of the baseboard management controller reaches a certain level, the fault detection function can be transferred to reduce the burden of the baseboard management controller, which is conducive to ensuring the reliability of the fault detection function while ensuring the stability of the baseboard management controller, and further conducive to ensuring the reliability of the server.
[0125] Furthermore, the fault detection mode can be preset to be associated with several control options to preset whether the fault detection mode is available or not. For example, the fault detection mode can be associated with a control option, and whether it is available is switched by different values of the control option. Alternatively, it can be set by two or more control options, which is not limited here.
[0126] For example, a first control option may be defined for the first mode, and a second control option may be defined for the second mode.
[0127] When both the first control option and the second control option are turned on, one of the first mode and the second mode, which is the default mode, can be used as the target detection mode. For example, when the first mode is the default mode, the first mode can be used as the target detection mode.
[0128] When the first control option is turned on and the second control option is turned off, the first mode can be used as the target detection mode.
[0129] When the first control option is turned off and the second control option is turned on, the second mode can be used as the target detection mode.
[0130] When the first control option and the second control option are both turned off, the system management interrupt can be used as the target detection mode.
[0131] Taking the specifically defined control options as an example, the control option of PRM can be defined as PRMsuppot; the control option of RAS offload can be defined as oobrassupport.
[0132] When oobrassupport and PRMsupport are both turned on, you can use PRM or RAS offload as the default mode as the target detection mode. For example, when PRM is the default mode, you can use PRM as the target detection mode.
[0133] When oobrassupport is turned on and PRMsupport is turned off, RAS offload can be used as the target detection mode.
[0134] When oobrassupport is turned off and PRMsupport is turned on, PRM can be used as the target detection mode.
[0135] When oobrassupport and PRMsupport are both turned off, SMI can be used as the target detection mode.
[0136] It is also possible to set whether oobrassupport (the first control option) is turned on and whether initialization data is transmitted to isolate it. That is, whether oobrassupport is turned off does not affect the transmission of RasOffLoad initialization data to BMC. In this way, even if rasoffload is turned off, the risk of PRM mode affecting address resolution on the BMC side is reduced.
[0137] In this embodiment, the flexible switching of fault diagnosis strategies in the Post stage and the Runtime stage can be achieved through the collaboration between BIOS settings, hardware interfaces such as VGPIO and CPLD, and firmware such as BMC and PRM SW, so as to adapt to the needs of different business scenarios, that is, it is possible to implement a fault diagnosis strategy that better matches the current usage scenario of the server, so as to improve the timeliness and reliability of fault detection. The compatibility design of the PRM solution and the RAS Offload solution in the Runtime stage can meet the matching of appropriate fault detection modes for different business scenarios to ensure stable business operation. In the Post stage, the traditional SMI can be used for fault reporting, and fault detection and processing can be performed through PRM and RAS Offload under the OS, and switching between the two can also be performed through platform management interface commands.
[0138] Please refer to Figure 5 as well as Figure 6 , Figure 5 It is a structural diagram of an embodiment of the server management system of the present application. Figure 6 It is a flowchart of the fourth embodiment of the server configuration method of the present application.
[0139] S501: The basic input and output system is initialized to the first mode.
[0140] In this embodiment, as described above, the first mode may be run on the baseboard management controller.
[0141] When initializing the first mode, the basic input and output system may inquire whether the baseboard management controller supports the first mode, so as to verify in advance whether the baseboard management controller can reliably run the first mode, so as to reliably detect faults of the server when the first mode is used as the target detection mode. In response to the baseboard management controller feeding back that it supports the first mode, the basic input and output system may send a first parameter combination of the first mode to the baseboard management controller, so as to initialize the first mode using the first parameter combination.
[0142] Specifically, the basic input and output system can uninstall the first mode and install the first mode to the baseboard management controller, and send the first parameter combination to the baseboard management controller. In this way, the baseboard management controller can load the first mode using the obtained first parameter combination to run the first mode to perform fault detection on the server.
[0143] For example, the basic input and output system may retain the target configuration control; wherein the target configuration control is used to control enabling or disabling of the first mode.
[0144] The basic input and output system uninstalls the first mode function and sends the associated data of the first mode function to the baseboard management controller.
[0145] The first parameter combination includes function support capability parameters, correctable error handling strategy parameters, bus topology, and memory topology.
[0146] In response to the first mode being the target detection mode, the basic input and output system sends an enable instruction to the baseboard management controller to enable the baseboard management controller to enable the correctable error handling function. The enable instruction carries a function mask, which is used to identify the first mode function to be started.
[0147] That is, BIOS can enable or disable RASOffload through a specific BIOS knob (i.e., target configuration control). Other RAS configuration knob values are sent to BMC during the handshake process, and the BMC's processing of RAS functions is affected by the feedback configuration knob values. For example, various error thresholds, backup strategies, memory-related configurations (such as mirror mode, etc.) and other configuration knob values can be sent. During the POST (i.e., initial stage) process, BIOS only performs corresponding hardware programming (such as setting a preset level of leaky bucket threshold, etc.) according to the knobs related to RAS configuration. At the same time, BMC can decide the RAS handler or function it executes according to the configuration transmitted by BIOS. For example, it can decide whether to install a backup firmware handler based on the technical knob for correcting errors in storage devices.
[0148] It can be seen that during the server startup phase, BIOS and BMC can determine the RAS functions to be unloaded to BMC through the handshake process, which involves the transfer of one or more data structures such as memory topology (MEM_TOPOLOGY), RAS policy (RAS_POLICIES) and IIO topology (IIO_TOPOLOGY). In this process, data interaction can be achieved through platform management interface commands, such as commands such as obtaining capabilities, sending data and starting RAS. At the same time, BIOS and BMC can respectively set corresponding processing APIs to implement operations such as data request, sending and function start notification in the handshake process. At runtime, BMC and host communicate through MMBI to support various RAS functions. As a transport and protocol layer, MMBI protocol can carry RAS-related request and response messages, including PCODE Mailbox Passthru, Runtime sPPR, WHEA Log Passthru, EDPCPort Info and other commands (all RAS-related commands), and can also set the sending and response data format for each command.
[0149] Further, the baseboard management controller can initialize the library initialization routine of the first mode. The baseboard management controller identifies the first mode function to be started carried by the function mask, and initializes the calling interface of the first mode function to enable the correctable error handling function.
[0150] Taking the first mode as RAS offload as an example, when initializing services such as RAS-manager, basic event registration and initialization calls can be performed. For example, platform management interface command handlers, DBus (Desktop Bus) and other inter-process communication interface registrations can be registered. A series of RAS library initialization routines can also be called, and rasLibInit (a RAS library initialization routine) can be executed as the main initialization routine.
[0151] If so, the BIOS can perform a handshake with the BMC through a platform management interface command to control the uninstallation of the RAS function, transfer related data and policies, and notify the BMC to start processing RAS events.
[0152] For example, Get Capabilities CMD* (0x22): BIOS queries the RAS capability of the BMC, the BMC returns a capability list, and the BIOS determines the RAS function (ie, the first mode function) to be finally enabled according to the capability list returned by the BMC.
[0153] Send Data CMD* (0x23): BIOS sends data and policies related to the offload function to the BMC, such as memory topology (MEM_TOPOLOGY), RAS policies (RAS_POLICIES), and IIO topology (IIO_TOPOLOGY). The data types are distinguished by different subtypes, such as 0x3 for memory topology, 0x5 for RAS policies, and 0x6 for IIO topology. Each type of data has a corresponding "Start" subcommand (such as 0xFD) before sending, and a "Transfer Complete" subcommand (such as 0xFE) after sending.
[0154] Start RAS command (Start RAS CMD*, 0x24): BIOS notifies BMC to start processing RAS events and passes the RAS function startup mask to indicate the started RAS function.
[0155] Furthermore, as the first mode is the RAS offload mode as exemplified above, when the basic input and output system initializes the first mode, the basic input and output system may initialize the RAS library, register the command processing program of the platform management interface, and so on.
[0156] S502: The basic input and output system initializes the second mode.
[0157] In this embodiment, as described above, the second mode can be run on the operating system.
[0158] If so, before initializing the second mode, the basic input and output system can determine whether the operating system supports the second mode, and in response to supporting the second mode, create the operating conditions of the second mode.
[0159] Specifically, the basic input and output system initializes the message processor of the second mode.
[0160] The basic input and output system may register the message processor with the configuration protocol of the second mode, so that the server, when under the control of the operating system, calls the message processor based on the registration information in the configuration protocol.
[0161] Furthermore, as described above, the basic input and output system can also configure a transmission port associated with the second mode so that the corresponding code command is executed when the register signal associated with it is triggered. The basic input and output system configures the second parameter combination for initialization of the second mode, which will not be described in detail here.
[0162] Furthermore, as the second mode explained in the previous article, it can be PRM. When the basic input and output system initializes the second mode, it can register macros such as PRM Handler through PRM_MODULE_EXPORT (a macro used to define the module export interface), install relevant information of PRM Module to the corresponding environment (for example, PRMConfigProtocol (a PRM used to configure parameters and settings of network protocols)), etc.
[0163] It should be noted that, in this embodiment, the first mode can be initialized first, and the second mode can be initialized after the first mode is initialized. Through a reasonably planned fault detection mode initialization sequence, it can be helpful to ensure the compatibility of the configurations of the first mode and the second mode, and it can be helpful to reduce the risk of configuration conflicts between the two, so that the functions of the subsequent first mode and the second mode can be reliable. In addition, in this embodiment, the first mode is run on the baseboard management controller, and the second mode is run on the operating system on the host side. While facilitating data interaction between the two, it can also be beneficial for the first mode and the second mode to realize data interaction, error feedback, etc. through MMBI (memory map BMC interface, BMC memory mapping). The following is a detailed description of the initialization process of the data space shared by the first mode and the second mode.
[0164] S503: The basic input and output system initializes the shared data space.
[0165] In this embodiment, considering that the first mode and the second mode may involve some of the same data, data structures, etc. during the implementation of functions and / or initialization stages, a data space that can be shared by the first mode and the second mode is created in this embodiment to store data that may be shared by the two modes and need to be synchronized.
[0166] That is, in response to the server startup, the basic input and output system can initialize both the first mode and the second mode, and the basic input and output system can also initialize the data space shared by the first mode and the second mode. Optionally, initializing the data space and initializing the first mode and the second mode can be performed in parallel, or interspersed, or the data space can be initialized first, so that the first mode and the second mode can call the initialization data in the data space for initialization when they are initialized.
[0167] Specifically, the basic input and output system initializes the data to be shared.
[0168] The basic input / output system transmits the data to be shared to a preset data space in the baseboard management controller, and maps the data space to the memory setting information of the basic input / output system, so that the baseboard management controller and the basic input / output system access the data space based on the memory setting information, and call the first mode and the second mode to initialize the data to be shared.
[0169] For example, the data to be shared may include at least one of a correctable error handling strategy parameter and a memory topology.
[0170] That is to say, during the startup phase of the server, the memory topology structure can be initialized, and the RAS configuration can be initialized, and the initialized memory topology and RAS configuration can be transmitted to the shared data space, so that the data in the data space can meet the requirements of the first mode while also meeting the requirements of the second mode. This makes it convenient for the first mode and the second mode to call the required RAS data to complete the initialization of their own modules, and can reduce the risk of repeated initialization due to data inconsistency.
[0171] In detail, when the correctable error handling strategy parameters are initialized, it can be considered that the RAS library is initialized. Specifically, the RAS library provides multiple initialization APIs (Application Programming Interface). As mentioned above, rasLibInit must be called before calling other library APIs. Therefore, rasLibInit can be considered as the main library initialization routine, which is used to initialize common components in the library, such as setting default cache property values.
[0172] In addition, each RAS function also has its own initialization API, which must be called for API initialization before calling the "process event" API of the corresponding function.
[0173] Other information related to the RAS library is described below. For example, RAS function support. That is to say, in this embodiment, multiple RAS functions can be supported, such as memory authentication report, post-packaging repair, detection and correction of errors in memory, mirror failover, multiple protocol reports, low-rate communication recovery, etc. Each function can be triggered by a specific associated API, and the triggering conditions of each other can be different. For example, the memory authentication report triggers ERRO (fault signal end) through a preset level of leaky bucket configuration.
[0174] This embodiment may also involve event reporting and recording. The RAS library uses a preset format to report errors and events, and supports multiple types of preset format records, including standard records (such as platform memory errors, PCIe (peripheralcomponent interconnect express, a high-speed serial computer expansion bus standard) errors, etc.) and non-standard records (such as memory backup events, cloud server errors, etc.).
[0175] Two internal monitors can be automatically registered in the RAS library, and the internal monitors can be WheaLogger and JournalLogger respectively. Among them, WheaLogger and JournalLogger are the names of the internal monitors.
[0176] WheaLogger can be responsible for adding standard preset format records to the memory area used to store hardware error records in the hardware error architecture inside the RAS library, and can configure the buffer in the hardware error source notification mechanism of the OS to send the preset format records. JournalLogger can decode all received preset format records into readable strings and output them to standard output.
[0177] Furthermore, the RAS library can also play and record functions, that is, the RAS library can record external inputs in the error handling process, such as register access, topology and strategy data, MMBI interaction, etc. And the RAS library can obtain data from pre-recorded files in the playback mode for simulation execution.
[0178] S504: The basic input and output system feeds back the target detection mode to the programmable logic device.
[0179] In this embodiment, taking the fault detection mode controlled by the basic input / output system as the target detection mode in the initial stage of server startup as an example, the basic input / output system can select the fault detection mode as the target detection mode based on whether the operating subject supports the fault detection mode it needs to run, and / or based on the preset fault detection mode. The specific selection method has been described in the previous text and will not be repeated here.
[0180] In response to selecting the fault detection mode as the target detection mode, the basic input output system may feed back the target detection mode to the programmable logic device, so as to enable the programmable logic device to adaptively enable the fault detection mode as the target detection mode.
[0181] For example, when the first mode is selected as the target detection mode, the basic input and output system may feed back relevant information that the target detection mode is the first mode to the programmable logic device.
[0182] S505: The programmable logic device enables the path between the fault signal terminal and the operating subject of the target detection mode.
[0183] In this embodiment, the programmable logic device can select the path between the fault signal end and the operating subject of the target detection mode in response to obtaining the target detection mode, so as to perform fault detection on the server using the target detection mode; wherein the fault signal end is used to transmit correctable errors, and the operating subject is a baseboard management controller or an operating system.
[0184] like Figure 5 As shown in the example, the operating system may include an ESPI interface, a GPIO0 interface, and an ERRO interface. The baseboard management controller includes an ESPI interface and a GPIO interface. The programmable logic unit includes a GPIO0 interface, a GPIO1 interface, and a GPIO2 interface.
[0185] The ESPI interface of the operating system may be considered as an Enhanced Serial Peripheral Interface (ESPI), and is used to connect to the ESPI interface of the baseboard management controller through the link L3.
[0186] As described above, the second mode runs on the operating system, and the GPIO0 interface of the operating system is connected to the GPIO1 interface of the programmable logic device, that is, when the programmable logic device selects the link L4 between its GPIO0 and GPIO1 interfaces, the path between the fault signal end and the operating subject of the second mode (i.e., the operating system) is selected. The GPIO0 interface of the operating system is connected to the ERRO interface of the operating system through the GPIO1 interface of the programmable logic device, the link L4, the GPIO0 interface of the programmable logic device, and the link L1. Among them, the ERRO interface can represent the fault signal end, which can be set in the CPU and does not conflict with the "error pin" mentioned above.
[0187] The GPIO interface of the baseboard management controller can be connected to the GPIO2 interface of the programmable logic device. Similarly, when the programmable logic device selects the link L2 between its GPIO0 interface and the GPIO2 interface, the path between the fault signal end and the operating subject of the first mode (i.e., the baseboard management controller) is selected. The GPIO interface of the baseboard management controller is connected to the ERRO interface of the operating system through the GPIO2 interface of the programmable logic device, the link L2, the GPIO0 interface of the programmable logic device, and the link L1.
[0188] At the same time, the baseboard management controller can also be connected to the programmable logic device through the link L5.
[0189] S506: The basic input and output system sends a target detection mode enabling command to the baseboard management controller.
[0190] In this embodiment, the specific method of sending the enable command has been explained in the previous text with examples, which will not be repeated here.
[0191] S507: The baseboard management controller enables the target detection mode.
[0192] In this embodiment, the baseboard management controller enables the target detection mode through its preset interface associated with the target detection mode in response to the acquisition enable command.
[0193] S508: The baseboard management controller runs in the first mode.
[0194] In this embodiment, when the first mode is used as the target detection mode, the baseboard management controller can run the first mode and use the first mode to detect whether the server is faulty.
[0195] S509: The operating system runs in the second mode.
[0196] In this embodiment, when the second mode is used as the target detection mode, the operating system can run the second mode and use the second mode to detect whether the server is faulty.
[0197] Furthermore, in this embodiment, the service initialized in the foregoing text can determine the way to handle RAS events. For example, an interrupt-based method can be adopted, that is, the ERRO signal can be detected, and the ERRO of each communication protocol can be connected to the GPIO of the BMC and configured as a level trigger. Alternatively, RAS events can be processed by polling. For example, some types of error sources can currently be polled, such as polling every 1 minute, 30 seconds, 2 minutes, etc. The specific content can be configured through attribute configuration.
[0198] When an error (such as a correctable error) or an anomaly is detected in the server, the RAS library can process the error event and construct an error message, so that the BMC can pass the error information to the operating information through MMBI to make the operating system aware of the error.
[0199] The following is an illustrative description of the detailed working principles of the first mode and the second mode of the present application.
[0200] When the first mode is used as the target detection mode, the first mode detects whether there is a correctable error in the server through error interrupt processing or polling; in response to the existence of a correctable error, the first mode constructs an error message through a preset error handling library and passes it to the operating system.
[0201] When the second mode is used as the target detection mode, the second mode configures the backup interrupt as a fault signal terminal option and enables it, and triggers the detection of whether there is a correctable error in the server through the system interrupt; in response to the existence of a correctable error, the second mode executes the error handling logic through the message processor and feeds back to the baseboard management controller.
[0202] Specifically, in response to the number of correctable errors reaching an error threshold, the central processing unit changes the state of the fault signal terminal option, causing a system interrupt to trigger the execution of a pre-stored method in the basic input and output system; wherein the pre-stored method is used to pass an identifier of the message processor to the operating system;
[0203] The operating system may query the storage location in the memory of the message processor matching the identifier, integrate the buffer information of the correctable error and transmit it to the message processor, so that the message processor processes the correctable error.
[0204] In detail, when the PRM processes CE (correctable error), it can connect Error#0 (a fault signal terminal) to a GPIO (General-purpose input / output) that supports triggering SCI (system interrupt) through hardware. The BIOS configures the register of the GPIO to 1 = Routing can cause SCI (this PIO routing can trigger SCI). You can also set the Spare Interrupt option to Error Pin (the pin that transmits the error) and enable IIOError Pin0 (the fault signal terminal of the bus), while turning off CSMI (the related setting option of SMI). The source language that defines the advanced configuration and power management interface object is called through SCI to trigger the PRM to perform CE processing so that the error signal is no longer transmitted using SMI.
[0205] When the CE generated by PCIe or Memory reaches the preset error threshold, the CPU pulls the Error#0 Pin to change the state of the GPIO configured above to support triggering SCI, and SCI is triggered to execute the pre-written pre-stored method in the BIOS code. This pre-stored method can pass the Guid (identifier) of the Prm Handler (PRM message processor) to be called to the ACPI subsystem of the OS (i.e., the driver for handling the advanced configuration and power management interface of the OS). The ACPI subsystem queries the location of the Prm Handler in the Memory through the Prm HandlerGuid from the PRMTTable (a structure used in driver development to define platform-related information), and integrates the content such as ContextBuffer / ParamBuffer, and passes it to the Prm Handler to process the CE. Among them, ContextBuffer / ParamBuffer is the buffer information that can correct errors.
[0206] In summary, the overall architecture of the server configuration system in this embodiment can focus mainly on RAS error handling, especially the correctable errors mentioned above. The BIOS can perform hardware initialization to enable the RAS function, and can also pass the required policies and data values to the BMC RAS handler, during which a handshake can be performed through the platform management interface command to transfer information. The RAS processing function can be added to the ras-manager (RAS management) service in the OpenBMC (open source BMC) service, and the ras-manager service can be used to interact with multiple components, such as RasLib, platform management interface service and MMBI, among which RasLib contains most of the RAS processing logic, can perform RAS processing and use the platform environment control interface to read and write registers.
[0207] The server of this application realizes flexible switching of fault detection modes to meet the needs of different business scenarios. It can also improve the timeliness and accuracy of fault detection through the collaborative work of hardware and firmware, that is, fault detection can be relatively efficient and accurate. In addition, BIOS settings and platform management interfaces are provided to facilitate the configuration and management of fault detection modes, thereby improving management convenience.
[0208] It should be understood that although Figures 1 to 4 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 1 to 4At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0209] See also Figure 7 , Figure 7 It is a structural diagram of an embodiment of the server of this application.
[0210] In one embodiment, the server includes a server body 11 and a control system 12 .
[0211] The server body 11 is a collection of basic components for realizing the server functions, which may include a graphics processor module, a central processing unit module, a heat dissipation module, a power module, a chassis, etc. of the server, which will not be described in detail here. For example, the server body 11 may include a BMC, a CPLD, a BIOS, a PRW SW, etc. as described above.
[0212] Among them, BIOS can implement fault detection mode selection and configuration information delivery in the initial stage of server startup. CPLD can control the routing of fault signals (i.e., conduct the path between the fault signal end and the CPU or BMC) according to BIOS configuration. BMC can be responsible for collecting and reporting error information in RAS offload mode. PRM SW can monitor and handle errors in PRM mode.
[0213] The control system 12 is disposed in the server body 11 and is used to implement the server configuration method in the above embodiment; or, to implement the server management method in the above embodiment.
[0214] Specifically, the control system 12 may include a basic input / output system, a baseboard management controller, an operating system, and a programmable logic device as described above.
[0215] As such, when the control system implements the server configuration method as described above, it can be further refined as follows: in response to server startup, the basic input and output system initializes the first mode and the second mode and the data space shared by the two, so that the first mode and the second mode call the initialization data in the data space for initialization; wherein, the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input and output system or the baseboard management controller uses one of the first mode and the second mode as the target detection mode, and feeds back the target detection mode to the programmable logic device; in response to obtaining the target detection mode, the programmable logic device selects the path between the fault signal end and the operating subject of the target detection mode, so as to use the target detection mode to perform fault detection on the server; wherein, the fault signal end is used to transmit correctable errors, and the operating subject is the baseboard management controller or the operating system.
[0216] For the specific definition of the control system, please refer to the definition of the server configuration method and the server management method above, which will not be repeated here. Each module in the above control system can be implemented in whole or in part by software, hardware and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0217] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the server configuration method or server management method described above are implemented, which will not be described in detail here.
[0218] That is, when the computer program is executed by the processor, the following steps are implemented: in response to the server startup, the basic input and output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the data space to initialize the first mode and the second mode; wherein, the first mode runs on the baseboard management controller, and the second mode runs on the operating system; the basic input and output system or the baseboard management controller uses one of the first mode and the second mode as the target detection mode, and feeds back the target detection mode to the programmable logic device; in response to obtaining the target detection mode, the programmable logic device selects the path between the fault signal end and the operating subject of the target detection mode, so as to use the target detection mode to perform fault detection on the server; wherein, the fault signal end is used to transmit correctable errors, and the operating subject is the baseboard management controller or the operating system.
[0219] Alternatively, when the computer program is executed by the processor, the following steps can be implemented: monitoring the operating parameters of the server; evaluating whether the server is faulty based on the operating parameters using a target detection mode; wherein the target detection mode is activated using a server configuration method such as in any of the above embodiments.
[0220] See also Figure 8 , Figure 8 It is a structural diagram of an embodiment of a computer device of the present application.
[0221] In one embodiment, the computer device may be a server, and its internal structure diagram may be as follows: Figure 8 As shown in the example.
[0222] The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a server configuration method or a server management method is implemented.
[0223] Those skilled in the art will understand that Figure 8 The structure shown as an example is merely a block diagram of a portion of the structure related to the present application solution, and does not constitute a limitation on the computer device to which the present application solution is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0224] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps can be implemented: in response to server startup, a basic input / output system initializes a data space shared by a first mode and a second mode, and calls initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on a baseboard management controller, and the second mode runs on an operating system; the basic input / output system or the baseboard management controller uses one of the first mode and the second mode as a target detection mode, and feeds back the target detection mode to a programmable logic device; in response to obtaining the target detection mode, the programmable logic device selects a path between a fault signal terminal and an operating subject of the target detection mode, so as to perform fault detection on the server using the target detection mode; wherein the fault signal terminal is used to transmit correctable errors, and the operating subject is the baseboard management controller or the operating system.
[0225] In one embodiment, the processor may also implement the following steps when executing the computer program: monitor the operating parameters of the server; and evaluate whether the server is faulty based on the operating parameters using a target detection mode; wherein the target detection mode is activated using a server configuration method as in any of the above embodiments.
[0226] See also Fig. 9 , Fig. 9 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application.
[0227] In one embodiment, the computer-readable storage medium 20 is used to store instruction / program data 21, and the instruction / program data 21 can be executed to implement the server configuration method as described in any of the above embodiments, or the server management method as described in the above embodiments, which will not be repeated here.
[0228] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device implementation described above is schematic, for example, the division of modules or units is a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0229] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of this embodiment.
[0230] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0231] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a computer-readable storage medium 20, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned computer-readable storage medium 20 includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, server and other media that can store program codes.
[0232] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps can be implemented: in response to server startup, a basic input / output system initializes a data space shared by a first mode and a second mode, and calls initialization data in the data space to initialize the first mode and the second mode; wherein the first mode runs on a baseboard management controller, and the second mode runs on an operating system; the basic input / output system or the baseboard management controller uses one of the first mode and the second mode as a target detection mode, and feeds back the target detection mode to a programmable logic device; in response to obtaining the target detection mode, the programmable logic device selects a path between a fault signal terminal and an operating subject of the target detection mode, so as to perform fault detection on the server using the target detection mode; wherein the fault signal terminal is used to transmit correctable errors, and the operating subject is the baseboard management controller or the operating system.
[0233] In one embodiment, when the computer program is executed by the processor, the following steps can also be implemented: monitoring the operating parameters of the server; using the target detection mode to evaluate whether the server is faulty based on the operating parameters; wherein the target detection mode is activated using the server configuration method as in any of the above embodiments.
[0234] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0235] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0236] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A server configuration method, characterized in that: The server configuration method comprises: In response to the server starting, the basic input and output system initializes the data space shared by the first mode and the second mode, and calls the initialization data in the shared data space to initialize the first mode and the second mode; wherein the first mode runs on the baseboard management controller, and the second mode runs on the operating system; The basic input and output system or the baseboard management controller uses one of the first mode and the second mode as a target detection mode, and feeds back the target detection mode to the programmable logic device; In response to obtaining the target detection mode, the programmable logic device enables a path between a fault signal terminal and an operating entity of the target detection mode, so as to perform fault detection on the server using the target detection mode; wherein the fault signal terminal is used to transmit correctable errors, and the operating entity is the baseboard management controller or the operating system.
2. The server configuration method according to claim 1, characterized in that: The data space shared by the first mode and the second mode of the basic input and output system initialization includes: The basic input and output system initializes the data to be shared; The basic input / output system transmits the data to be shared to the shared data space preset in the baseboard management controller, and maps the shared data space to the memory setting information of the basic input / output system, so that the baseboard management controller and the basic input / output system access the shared data space based on the memory setting information, and call the data to be shared to initialize the first mode and the second mode.
3. The server configuration method according to claim 2, characterized in that: The data to be shared includes at least one of a correctable error handling strategy parameter and a memory topology.
4. The server configuration method according to claim 1, characterized in that: Initializing the first mode and the second mode includes: The basic input and output system inquires whether the baseboard management controller supports the first mode; In response to the baseboard management controller feeding back that it supports the first mode, the basic input and output system sends a first parameter combination of the first mode to the baseboard management controller; The basic input and output system determines whether the operating system supports the second mode, and in response to supporting the second mode, creates an operating condition for the second mode.
5. The server configuration method according to claim 4, characterized in that: The basic input and output system sending the first parameter combination of the first mode to the baseboard management controller includes: The basic input and output system uninstalls the first mode and installs the first mode to the baseboard management controller; and sends the first parameter combination to the baseboard management controller; wherein the first parameter combination includes a function support capability parameter, a correctable error handling strategy parameter, a bus topology, and a memory topology; In response to the first mode being the target detection mode, the basic input and output system sends an enable instruction to the baseboard management controller to enable the baseboard management controller to enable the correctable error handling function; wherein the enable instruction carries a function mask, and the function mask is used to identify the first mode function to be started.
6. The server configuration method according to claim 5, characterized in that: The basic input and output system uninstalls the first mode and installs the first mode to the baseboard management controller, comprising: The basic input and output system retains a target configuration control; wherein the target configuration control is used to control enabling or disabling of the first mode; The basic input and output system uninstalls the first mode function and sends the associated data of the first mode function to the baseboard management controller.
7. The server configuration method according to claim 5, characterized in that: The baseboard management controller enables execution of correctable error handling functions including: The baseboard management controller initializes a library initialization routine of the first mode; The baseboard management controller identifies the first mode function to be started carried by the function mask, and initializes a calling interface of the first mode function.
8. The server configuration method according to claim 4, characterized in that: The operating conditions for creating the second mode include: The basic input and output system initializes the message processor of the second mode; The basic input and output system registers the message processor with the configuration protocol of the second mode, so that the server calls the message processor based on the registration information in the configuration protocol when it is under the control of the operating system.
9. The server configuration method according to claim 8, characterized in that: The operating conditions for creating the second mode also include: The basic input and output system configures the transmission port associated with the second mode so that a corresponding code command is executed when a register signal associated with the second mode is triggered; The basic input and output system configures a second parameter combination for initializing the second mode; wherein the second parameter combination includes correctable error handling strategy parameters, bus topology, and memory topology.
10. The server configuration method according to claim 1, characterized in that: The using one of the first mode and the second mode as the target detection mode includes: In response to being in the initial stage of server startup, the basic input and output system reads a setting option for fault detection; obtains a preset fault detection mode identified by the setting option and uses it as the target detection mode; and / or, In response to the initial stage of starting up through the server, the baseboard management controller obtains a switching instruction; and parses the fault detection mode to be switched identified by the switching instruction, and uses it as the target detection mode.
11. The server configuration method according to claim 1, characterized in that: The baseboard management controller uses one of the first mode and the second mode as a target detection mode, including: The baseboard management controller obtains the status parameters of the server; The baseboard management controller analyzes the state parameter to evaluate whether to switch to a fault detection mode as the target detection mode; In response to determining to switch the fault detection mode as the target detection mode, the baseboard management controller controls the programmable logic device to enable other fault detection modes as new target detection modes; wherein the other fault detection modes represent fault detection modes that are not currently the target detection mode.
12. The server configuration method according to claim 11, characterized in that: The analyzing the state parameter to evaluate whether to switch to the fault detection mode as the target detection mode includes: In response to the first mode being the target detection mode, the baseboard management controller acquires a resource utilization parameter of the baseboard management controller; compares the resource utilization parameter with a resource utilization threshold; in response to the resource utilization parameter reaching the resource utilization threshold, determines to switch to a fault detection mode as the target detection mode; The programmable logic device selects other fault detection modes as new target detection modes including: The programmable logic device switches the path between the fault signal terminal and the operating system to the second mode as the target detection mode.
13. The server configuration method according to claim 1, characterized in that: The path between the strobe fault signal terminal and the target detection mode operation subject also includes: the basic input and output system triggers a system management interrupt, and performs fault feedback through an interrupt signal; The path between the strobe fault signal terminal and the operating body of the target detection mode further includes: The basic input and output system sends an enabling command of the target detection mode to the baseboard management controller; in response to obtaining the enabling command, the baseboard management controller enables the target detection mode through a preset interface associated with the target detection mode.
14. The server configuration method according to claim 1, characterized in that: When the first mode is used as the target detection mode, performing fault detection on the server by using the target detection mode includes: The first mode detects whether the server has the correctable error by error interrupt processing or polling; in response to the correctable error, the first mode constructs an error message by a preset error processing library and transmits it to the operating system; and / or, When the second mode is used as the target detection mode, performing fault detection on the server by using the target detection mode includes: The second mode configures the backup interrupt as a fault signal terminal option and enables it, and detects whether the server has the correctable error through a system interrupt trigger; in response to the existence of the correctable error, the second mode executes error handling logic through a message processor and feeds back to the baseboard management controller.
15. The server configuration method according to claim 14, characterized in that: In response to the presence of the correctable error, the second mode executing error handling logic by the message processor includes: In response to the number of correctable errors reaching an error threshold, the central processing unit changes the state of the fault signal terminal option, causing the system interrupt to trigger the execution of a pre-stored method in the basic input and output system; wherein the pre-stored method is used to pass the identifier of the message processor to the operating system; The operating system queries the storage location of the message processor matching the identifier in the memory, integrates the buffer information of the correctable error and transmits it to the message processor, so that the message processor processes the correctable error.
16. A server management method, characterized in that: The server management method comprises: Monitor the server's operating parameters; Based on the operating parameters, a target detection mode is used to evaluate whether the server is faulty; wherein the target detection mode is activated using the server configuration method according to any one of claims 1-15.
17. A server, characterized in that: The server comprises: Server body; A control system, arranged in the server body, is used to implement the server configuration method as described in any one of claims 1 to 15; or, to implement the server management method as described in claim 16.
18. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the server configuration method described in any one of claims 1 to 15 are implemented; or, the server management method described in claim 16 is implemented.
19. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the processor implements the steps of the server configuration method described in any one of claims 1 to 15; or, implements the server management method described in claim 16.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the server configuration method described in any one of claims 1 to 15 are implemented; or, the server management method described in claim 16 is implemented.
Citation Information
Patent Citations
Method and device for reporting error information and medium
CN113064745A
Memory fault processing method and device, computer equipment and storage medium
CN116483612A