Server warning function management method, server, storage medium and program product

By disabling the warning function when the server starts, obtaining the memory status and restoring the faulty memory before enabling the warning function, the problem of server downtime caused by memory failure during the development phase was solved, the workload of R&D and testing personnel was reduced, and the server's fault alarm capability was maintained.

CN120386694BActive Publication Date: 2025-09-19INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510884211.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-19
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

During the server development phase, downtime problems caused by memory failures occur frequently, and turning off the warning function will sacrifice the fault alarm capability and increase the workload of R&D and testing personnel.

Method used

Disable the warning function when the server starts, obtain the memory status, and enable the warning function after the faulty memory is restored. Manage the warning function switch through interaction between the BMC and BIOS to ensure that the warning function is enabled only when the memory status is normal.

Benefits of technology

It avoids server downtime caused by memory failure, reduces the workload of R&D and testing personnel, and at the same time maintains the server's fault alarm capability, achieving a balance between rapid resolution of downtime problems and fault alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386694B_ABST
    Figure CN120386694B_ABST
Patent Text Reader

Abstract

The present invention discloses a server warning function management method, a server, a storage medium and a program product, which relate to the field of computer technology. When the server is turned on, the warning function of the server can be turned off first, and the memory in the first state, that is, the faulty memory, is restored, and then the server is restarted. When all the memories of the server are in the second state, that is, the memory is in a normal state, the warning function of the server is turned on, so as to avoid the server downtime caused by memory failure, reduce the time spent by R&D or testing personnel to find the downtime problem, and at the same time ensure the server's warning capability for failure, and strike a balance between avoiding the server downtime problem and improving the warning capability when the server fails. Therefore, the technical problems in the related technology of server downtime caused by memory failure, thereby increasing the time to find the downtime problem, and reducing the server's warning capability by turning off the warning function can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a server warning function management method, a server, a storage medium and a program product. Background Art

[0002] The server's warning function, Promote Warnings, is a system warning switch provided by the server's CPU (Central Processing Unit). When this function is enabled, it can provide more detailed and comprehensive system-level warnings when a server failure occurs. Therefore, this system-level warning function is generally enabled in the BIOS (Basic Input Output System) and CRB (Customer Reference Board) program code provided by the CPU manufacturer. However, this design is based on the CPU manufacturer's public motherboard and CRB code. All configurations use the minimized boot configuration and one or two memory sticks, so memory failures rarely occur.

[0003] However, during the server product development phase, OEMs (Original Equipment Manufacturers) often test server configurations with multiple memory modules, even fully configured. Furthermore, the memory used during product development is often prototypes, often with numerous issues. In both cases, the likelihood of a memory failure is significantly increased. If the Promote Warnings feature is enabled, a faulty memory module could cause the server to crash during the initialization process, preventing it from booting properly. This is especially true during the product development phase, where failures can lead to numerous server downtimes, forcing developers to constantly identify the cause of the downtime, consuming significant time and effort. Consequently, OEMs often disable the Promote Warnings feature. However, this often sacrifices the ability to warn of server failures in order to mitigate downtime caused by memory failures. Summary of the Invention

[0004] The present invention provides a server warning function management method, server, storage medium and program product to at least solve the problem in the related art that server downtime due to memory failure increases the time to find the downtime problem, and that disabling the warning function reduces the server's alarm capability.

[0005] The present invention provides a server warning function management method, comprising the following steps: turning off the warning function of the server at the time of server startup; acquiring memory data of the server, extracting the operating status of all memories of the server in the memory data, wherein the operating status includes a first state and a second state, the first state indicating a fault state of the memory, and the second state indicating a normal state of the memory; if the first state exists in the operating status of all memories in the server, performing a fault recovery action on the memory in the first state, and restarting the server after recognizing that the fault recovery action is completed; if the operating status of all memories in the server is the second state, starting the warning function of the server, and using the warning function to issue a fault alarm to the server.

[0006] The present invention also provides a server warning function management device, including: a shutdown module, used to shut down the server's warning function when the server is started; an extraction module, used to obtain the server's memory data, and extract the operating status of all the server's memories in the memory data, wherein the operating status includes a first state and a second state, the first state represents a fault state of the memory, and the second state represents a normal state of the memory; an execution module, used to perform a fault recovery action on the memory in the first state if the first state exists in the operating status of all the memories in the server, and restart the server after recognizing that the fault recovery action is completed; a startup module, used to start the server's warning function if the operating status of all the memories in the server is the second state, and use the warning function to issue a fault alarm to the server.

[0007] The present invention also provides a server, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server warning function management methods when executing the computer program.

[0008] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server warning function management methods are implemented.

[0009] The present invention also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned server warning function management methods when executed by a processor.

[0010] Through the present invention, since during the product development stage of the server, the server machine configuration will adopt multiple memories or even full memory testing, and the memories used in the product development stage are basically samples, there are many problems with the memory, so the possibility of memory failure will be greatly increased. When the warning function of the server is turned on, if there is a faulty memory, the server will encounter a memory failure during the startup memory initialization process, which will cause the server to crash and be unable to start normally, and then the R&D personnel will need to find the server crash problem, which will cost the R&D and testing personnel a lot of time and energy, or directly turn off the server's warning function to avoid the server crash problem caused by memory failure, but it will sacrifice the server's alarm capability for other faults except memory failure. Therefore, the present invention can first turn off the server's warning function when the server is turned on, and restore the memory in the first state, that is, the faulty memory, and then restart the server. The server, when all the memory of the server is in the second state, that is, when the memory is in a normal state, the warning function of the server is turned on to avoid the server downtime caused by memory failure, reduce the time spent by R&D or testers to find the downtime problem, and at the same time ensure the server's warning ability for failure, and strike a balance between avoiding the server downtime problem and improving the warning ability of the server when a failure occurs. Therefore, it can solve the technical problems of server downtime caused by memory failure in related technologies, thereby increasing the time to find the downtime problem, and turning off the warning function to reduce the server's warning ability, so as to achieve the technical effect of avoiding the server downtime problem caused by memory failure, reducing the time spent by R&D or testers to find the downtime problem, and at the same time ensuring the server's warning ability for failure, and avoiding the server downtime problem and improving the warning ability of the server when a failure occurs. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 Flowchart of a server warning function management method according to an embodiment of the present invention;

[0013] Figure 2 An execution diagram of a server warning function management method provided according to an embodiment of the present invention;

[0014] Figure 3 A schematic diagram of fault recovery actions corresponding to fault types provided in an embodiment of the present invention;

[0015] Figure 4A schematic diagram of a server warning function management device according to an embodiment of the present invention;

[0016] Figure 5 A schematic diagram of the structure of a server provided according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0018] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0019] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] An embodiment of the present invention provides a server warning function management method, and the method is described in detail in conjunction with the execution flow of the server warning function management method.

[0021] Figure 1 The figure is a flow chart of a method for managing a server warning function according to an embodiment of the present invention.

[0022] like Figure 1 As shown, the server warning function management method includes the following steps:

[0023] In step S101, the warning function of the server is turned off when the server is started.

[0024] The warning function can be the server's Promote Warnings function. Promote Warnings is a system warning function provided by the server CPU to control whether the warning is elevated to the system level when a fault occurs. This allows for more detailed and comprehensive warning processing when a server fault occurs at the system level. The server's BMC (Baseboard Manager Controller) controls whether the warning function is enabled or disabled.

[0025] It should be noted that the server warning function management method of the present invention is mainly used in the product development stage of the server.

[0026] In step S102, memory data of the server is obtained, and the operating status of all memories of the server in the memory data is extracted, wherein the operating status includes a first state and a second state, the first state indicates a faulty state of the memory, and the second state indicates a normal state of the memory.

[0027] Among them, memory refers to the memory bar in the server hardware, that is, the physical hardware component used by the computer to temporarily store data. Memory data can be obtained through the BIOS program.

[0028] It is understandable that the embodiment of the present invention can obtain the memory data of the server and extract the operating status of all memories of the server from the memory data to determine the status of each memory in the server.

[0029] In an embodiment of the present invention, obtaining memory data of a server includes: identifying the amount of memory in the server; setting a memory status array according to the amount of memory; and recording the memory data of the server using the memory status array.

[0030] The memory status array is an array used to record the status of each memory, which can be defined as A[n], where n is the memory index.

[0031] It can be understood that the embodiment of the present invention can identify the amount of memory in the server, set the memory status array according to the memory amount, and use the memory status array to record the memory data of the server, and uniformly manage all memory data through the array structure to facilitate subsequent memory fault judgment.

[0032] For example, if the server has four memory sticks installed, create arrays A[0]-A[3] to record the status of each memory stick.

[0033] In an embodiment of the present invention, recording memory data of a server using a memory status array includes: identifying the amount of recorded data in the memory status array; and controlling recording of the memory data of the server according to the amount of recorded data.

[0034] It is understandable that the embodiment of the present invention can utilize the ability to identify the memory status and control the recording of memory data of the server according to the amount of recorded data.

[0035] In an embodiment of the present invention, the recording of the memory data of the server is controlled according to the amount of recorded data, including: if the amount of recorded data is less than the memory amount, the memory data of the server is controlled to be recorded until the amount of recorded data reaches the memory amount, and the recording is stopped.

[0036] The memory size is represented by M, and the number of recorded data is represented by n.

[0037] It can be understood that the embodiment of the present invention can control the server's memory data to record when the number of recorded data is less than the memory number, that is, when n is less than M, until the number of recorded data reaches the memory number M, stop recording, and perform subsequent steps, thereby avoiding missing or repeated recording of memory status, and ensuring the accuracy of subsequent memory fault judgment.

[0038] For example, if the server has four memory sticks, record n ​​starting from 0 and stop recording when n=3 to ensure that the status of all four memory sticks are recorded.

[0039] In step S103, if the first state exists in the running states of all memories in the server, a fault recovery action is performed on the memory in the first state, and the server is restarted after it is identified that the fault recovery action is completed.

[0040] The memory in the first state may be referred to as a faulty memory.

[0041] It can be understood that the embodiment of the present invention can perform a fault recovery action on the memory in the first state, i.e., the faulty memory, when the operating status of all memories in the server is in the first state, and restart the server after identifying that the fault recovery action has been completed, thereby eliminating and recovering the faulty memory.

[0042] In an embodiment of the present invention, performing a fault recovery action on a memory in a first state includes: obtaining a fault type of the memory in the first state; and determining a fault recovery action based on the fault type.

[0043] It is understandable that the embodiment of the present invention can obtain the fault type of the memory in the first state, and determine the fault recovery action based on the fault type, thereby accurately determining the fault recovery action of the normal memory.

[0044] In an embodiment of the present invention, obtaining the fault type of the memory in the first state includes: identifying fault information in memory data of the memory in the first state; and determining the fault type based on the fault information.

[0045] It is understandable that the embodiment of the present invention can determine the fault type based on the fault information in the memory data of the memory in the first state.

[0046] In an embodiment of the present invention, determining the fault type based on the fault information includes: acquiring a correspondence between the fault information and the fault type; and determining the fault type by querying the correspondence based on the fault information.

[0047] It is understandable that the embodiment of the present invention can determine the fault type of the faulty memory according to the correspondence between the fault information and the fault type, so as to improve the accuracy of the faulty memory diagnosis.

[0048] For example, the fault information in the memory data of the fault memory is A. By querying the corresponding relationship, it is determined that the fault type is A1.

[0049] In the embodiment of the present invention, the fault types include first to fifth types, and the fault recovery actions include first to fifth actions.

[0050] It is understandable that the embodiments of the present invention include multiple types of faults, and a corresponding fault recovery action can be accurately matched according to each fault type to achieve fault recovery of the faulty memory.

[0051] In an embodiment of the present invention, the first type is a hardware damage abnormal failure, and the first action is to replace the memory in the first state; the second type is a poor contact abnormal failure, and the second action is to re-plug the memory in the first state, clean the dust in the motherboard slot corresponding to the memory in the first state, or change the memory slot position in the first state; the third type is a compatibility abnormal failure, and the third action is to replace the memory with the same model as the memory in the first state; the fourth type is a configuration abnormal failure, and the fourth action is to restore the default parameters of the memory in the first state; the fifth type is an error abnormal failure, and the fifth action is to detect the memory in the first state and repair or replace the memory in the first state.

[0052] Among them, hardware damage abnormal failures may be damaged memory sticks; poor contact abnormal failures may be poor contact between the memory stick and the motherboard slot; compatibility abnormal failures may be incompatibility between the installed memory of different models or brands and the server; configuration abnormal failures may be improper configuration of memory frequency or parameters; and error abnormal failures may be memory ECC (Error Checking and Correction) error failures recorded in the log.

[0053] It is understandable that the embodiment of the present invention can perform different fault recovery actions according to the specific type of faulty memory, specifically:

[0054] If the fault type is type 1, which is a hardware damage problem, the first recovery action is to replace the memory to completely resolve the physical fault and prevent subsequent server downtime;

[0055] If the fault type is type 2, which is a physical problem such as poor contact between the memory module and the motherboard slot, the fault recovery action is the second action, which involves reseating the memory module, cleaning the motherboard slot dust, or relocating the slot to fix the physical connection problem. There is no need to replace the memory.

[0056] If the fault type is Type 3, it is a memory compatibility fault caused by installing different models or brands of memory, resulting in system instability or partial memory unrecognition. The fault recovery action is Action 3, which is to replace the memory with the same model to resolve the system instability caused by inconsistent memory specifications.

[0057] If the fault type is Type 4, which is a configuration fault caused by improper memory frequency or parameter configuration, resulting in an abnormal display of system memory frequency or capacity, the fault recovery action is Action 4. This involves clearing custom settings in the BIOS option by using a motherboard jumper or clearing the CMOS (Complementary Metal-Oxide-Semiconductor) battery, restoring default parameters, and repairing the memory fault caused by the suspected configuration error.

[0058] If the fault type is type 5, which means a memory ECC error is recorded in the log, the fault recovery action is action 5. Use the corresponding diagnostic tool to detect the memory module or replace the memory module with the ECC fault to locate and repair the memory data error.

[0059] In step S104, if the operating status of all memories in the server is the second status, the warning function of the server is started, and a fault alarm is issued to the server using the warning function.

[0060] During the product development phase of the server, the server configuration will use multiple memory sticks or even full memory testing, and the memory used in the product development phase is basically a sample, which has many problems. Therefore, the possibility of memory failure will be greatly increased. When the server warning function is turned on, if there is a faulty memory, the server will encounter a memory failure during the memory initialization process, which will cause the server to crash and be unable to start normally. R&D personnel will then need to find the server crash problem, which will cost R&D and test personnel a lot of time and energy. Alternatively, the server warning function can be turned off directly to avoid the server crash caused by memory failure, but this will sacrifice The server's ability to warn of other faults except memory failure is sacrificed. Therefore, an embodiment of the present invention first turns off the server's warning function when the server is turned on, recovers the memory in the first state, that is, the faulty memory, and then restarts the server. When all the memories of the server are in the second state, that is, the memory is in a normal state, the server's warning function is turned on to avoid server downtime caused by memory failure, reduce the time spent by R&D or testing personnel to find downtime problems, and at the same time ensure the server's ability to warn of faults, and strike a balance between avoiding server downtime problems and improving the ability to warn when a server failure occurs.

[0061] In summary, the server warning function management method proposed in the embodiment of the present invention can intelligently adjust the server warning function. During the server startup process, the BMC and BIOS will exchange information and place the function switch option that controls the warning function in the BMC. The initial value is set to turn off the warning function. Each time the server is started, the BIOS program obtains the function status of the warning function from the BMC. At this time, if there is a faulty memory, the downtime problem can also be avoided. During the startup process, the memory status is captured. If there is no memory abnormality, the memory status is sent to the BMC. The BMC sets the function switch option to turn off the function. When the server is started next time, the status obtained from the BMC is turned on, thereby improving the system's warning level capability and avoiding the waste of resources caused by no abnormal memory. If there is a memory fault, the faulty memory problem is solved and the machine is in a memory fault-free state. In this way, a balance is struck between the server downtime problem caused by the memory fault caused by turning on the function and the improvement of the system's warning level capability when an overall fault occurs, thereby avoiding the downtime problem and improving the system's warning level capability when an overall fault occurs.

[0062] Specifically, the server warning function management method of the embodiment of the present invention is described with the Promote Warnings warning function. The specific process is as follows: Figure 2 Shown, including:

[0063] 1. Power on the computer and set the Promote Warnings switch on the BMC to off.

[0064] 2. Check the current machine memory, record it as M, and define the memory state as A[n] (n=n+1, n=0, 1, 2…).

[0065] 3. Record the status of each memory as A[n] (n=n+1, n=0, 1, 2…).

[0066] 4. When n is not greater than M, it means that there is still unrecorded memory. Return to step 3 and continue recording the unrecorded memory.

[0067] 5. When n reaches M, it means that all memory states have been recorded, and the memory state recording stops at this time.

[0068] 6. Determine the status of all memories and whether there is any abnormal faulty memory in A[n].

[0069] 7. There are many types of memory failures. If there is a faulty memory, relevant processing will be performed according to the memory failure type information recorded by A[n]. The specific processing actions for different failure types are as follows: Figure 3 Shown, including:

[0070] (1) If the memory is physically damaged, replace the memory.

[0071] (2) If the problem is physical, such as poor contact between the memory module and the motherboard slot, then re-plug the memory module, clean the motherboard slot dust, or change the slot position;

[0072] (3) If a memory compatibility failure occurs due to the installation of memory of different models or brands, causing system instability or partial memory unrecognition, replace the memory with the same model;

[0073] (4) If the memory frequency or parameters are improperly configured, resulting in an abnormal display of the system memory frequency or capacity, the BIOS option can be cleared by using the motherboard jumper or clearing the CMOS battery to restore the default parameters.

[0074] (5) If the system log records a memory ECC error, use the corresponding diagnostic tool to detect the memory module or replace the memory module with the ECC failure.

[0075] 8. After resolving the memory failure issue, power on the computer and start again from step 1.

[0076] 9. If all memory is normal and there is no abnormal faulty memory, a memory normal signal is sent to the BMC, triggering the BMC to set the control function switch option Promote Warnings to the on state to increase the alarm level of the system.

[0077] The following describes a method for managing a server warning function according to an embodiment of the present invention using a specific embodiment.

[0078] An OEM manufacturer is developing a new server with the following configuration: CPU (Intel Xeon processor supporting the Promote Warnings function), memory (8 memory sticks, all samples), BIOS (customized version based on the CRB reference board), and BMC.

[0079] During the testing phase, the server frequently fails to start up, so the server warning function management method of the present invention is applied to solve the problem. The specific steps are as follows:

[0080] 1. Server startup and warning functions are turned off.

[0081] When the server is powered on, the BMC automatically turns off the Promote Warnings function. Even if a memory failure occurs, the system will not immediately shut down due to the warning upgrade, allowing subsequent detection and repair.

[0082] 2. Memory status detection and recording.

[0083] (1) BIOS detects the number of memory sticks: Through hardware scanning, it is identified that the server has 8 memory sticks installed (M=8).

[0084] (2) Create a memory status array: define arrays A[0]-A[7] to record the status of each memory.

[0085] (3) Detect memory status in a loop: counter n starts at 0 and detects each memory stick in turn. When n=3, an ECC check error is detected in the fourth memory stick (A[3]) and recorded as a faulty state. Continue to detect the remaining memory sticks and finally stop the detection when n=8, confirming that A[3] is the only faulty memory stick.

[0086] 3. Identify and handle fault types.

[0087] (1) Obtaining fault information: The system log shows that the ECC error code of A [3] is 0x502 (for example, indicating that a single-bit error is unrecoverable).

[0088] (2) Determine the fault type: Through the preset correspondence, the error code 0x502 is mapped to the fifth type of fault (ECC error exception).

[0089] (3) Execute fault recovery action: automatically call the memory diagnostic tool to detect A [3], confirm that the memory bar is physically damaged, and prompt the operation and maintenance personnel to replace the memory bar corresponding to A [3].

[0090] 4. Restart the server and enable the warning function.

[0091] (1) Restart after replacing the faulty memory: After the operation and maintenance personnel replaced A [3], the server was powered on again.

[0092] (2) Recheck the memory status: Repeat step 2 to confirm that the status of all memories (A[0]-A[7]) are normal.

[0093] (3) Enable the warning function: When the BMC receives the signal that "all memory is normal", it automatically turns on the PromoteWarnings function.

[0094] As a result, the server boots up normally, and if a memory failure occurs during subsequent operation, the system will issue a timely alarm through the PromoteWarnings function.

[0095] In this embodiment, although there is a fault memory when the server is turned on for the first time, the server does not crash immediately because the warning function is turned off, allowing the fault detection and repair to be completed; by mapping the error code and the fault type, A [3] is quickly determined to be hardware damage, avoiding blind attempts at other repair methods; the warning function is automatically turned on after the repair, ensuring that the system has complete alarm capabilities in subsequent operations; compared with traditional methods (which require manual repeated troubleshooting and manual switching functions), this embodiment shortens the fault handling time, thereby realizing automatic shutdown function-fault detection-repair-restart-automatic activation function, forming a closed-loop management, and no manual intervention is required to switch the warning function on and off.

[0096] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0097] According to the server warning function management method proposed in an embodiment of the present invention, the server warning function can be turned off first when the server is turned on, and the memory in the first state, that is, the faulty memory, can be restored before restarting the server. When all the memories of the server are in the second state, that is, the memory is in a normal state, the server warning function is turned on to avoid server downtime problems caused by memory failures, reduce the time spent by R&D or testing personnel to find downtime problems, and at the same time ensure the server's warning capability for faults, and strike a balance between avoiding server downtime problems and improving the warning capability when server failures occur.

[0098] An embodiment of the present invention also provides a server warning function management device.

[0099] Figure 4 Schematic diagram of a server warning function management device according to an embodiment of the present invention.

[0100] like Figure 4 As shown, the server warning function management device 10 includes: a closing module 100 , an extracting module 200 , an executing module 300 and a starting module 400 .

[0101] Among them, the shutdown module 100 is used to turn off the server's warning function when the server is started; the extraction module 200 is used to obtain the server's memory data and extract the operating status of all the server's memories in the memory data, wherein the operating status includes a first state and a second state, the first state indicates a fault state of the memory, and the second state indicates a normal state of the memory; the execution module 300 is used to perform a fault recovery action on the memory in the first state if the first state exists in the operating status of all the memories in the server, and restart the server after recognizing that the fault recovery action is completed; the startup module 400 is used to start the server's warning function if the operating status of all the memories in the server is the second state, and use the warning function to issue a fault alarm to the server.

[0102] In the embodiment of the present invention, the extraction module 200 is further used to: identify the amount of memory in the server; set a memory status array according to the amount of memory; and record the memory data of the server using the memory status array.

[0103] In the embodiment of the present invention, the extraction module 200 is further configured to: identify the number of recorded data in the memory status array; and control the recording of memory data of the server according to the number of recorded data.

[0104] In an embodiment of the present invention, the extraction module 200 is further configured to: if the amount of recorded data is less than the memory amount, control the memory data of the server to record until the amount of recorded data reaches the memory amount, and then stop recording.

[0105] In the embodiment of the present invention, the execution module 300 is further configured to: obtain the fault type of the memory in the first state; and determine a fault recovery action based on the fault type.

[0106] In the embodiment of the present invention, the execution module 300 is further configured to: identify fault information in the memory data of the memory in the first state; and determine the fault type based on the fault information.

[0107] In the embodiment of the present invention, the fault types include first to fifth types, and the fault recovery actions include first to fifth actions.

[0108] In the embodiment of the present invention, the first type is a hardware damage abnormal fault, and the first action is to replace the memory in the first state.

[0109] In an embodiment of the present invention, the second type is an abnormal fault due to poor contact, and the second action is to re-plug the memory in the first state, clean the dust in the motherboard slot corresponding to the memory in the first state, or replace the memory slot position in the first state.

[0110] In the embodiment of the present invention, the third type is a compatibility abnormality failure, and the third action is to replace the memory with a memory of the same model as the memory in the first state.

[0111] In the embodiment of the present invention, the fourth type is a configuration abnormality fault, and the fourth action is to restore the default parameters of the memory in the first state.

[0112] In the embodiment of the present invention, the fifth type is an abnormal fault with an error report, and the fifth action is to detect the memory in the first state and repair or replace the memory in the first state.

[0113] It should be noted that, for the description of the features in the embodiment corresponding to the server warning function management device, reference can be made to the relevant description of the embodiment corresponding to the server warning function management method, which will not be repeated here.

[0114] According to the server warning function management device provided by the embodiment of the present invention, the server warning function can be turned off first when the server is turned on, and the memory in the first state, that is, the faulty memory, can be restored, and then the server can be restarted. When all the memory of the server is in the second state, that is, the memory is in a normal state, the server warning function is turned on to avoid server downtime problems caused by memory failures, reduce the time spent by R&D or testing personnel to find downtime problems, and at the same time ensure the server's warning capability for faults, and strike a balance between avoiding server downtime problems and improving the warning capability when the server fails.

[0115] The embodiment of the present invention further provides a server, such as Figure 5 As shown, it includes a memory 501 and a processor 502, the memory 501 stores a computer program, and the processor 502 is configured to run the computer program to execute the steps in any of the above server warning function management method embodiments.

[0116] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned server warning function management method embodiments when running.

[0117] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0118] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server warning function management method embodiments are implemented.

[0119] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0120] The above is a detailed introduction to the server warning function management provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A server warning function management method, characterized in that: The following steps are involved: disabling a warning function of the server at the time of server startup, wherein the warning function is a PromoteWarnings function of the server, and the PromoteWarnings function is a system-level warning prompt function of the server; Obtaining memory data of the server, and extracting the operating status of all memories of the server from the memory data, wherein the operating status includes a first state and a second state, the first state indicating a fault state of the memory, and the second state indicating a normal state of the memory; If the first state exists in the running states of all memories in the server, performing a fault recovery action on the memory in the first state, and restarting the server after recognizing that the fault recovery action is completed; If the operating status of all memories in the server is the second state, the warning function of the server is activated, and the warning function is used to issue a fault alarm to the server, wherein the activation or deactivation of the warning function of the server is controlled by the baseboard management controller of the server.

2. The server warning function management method according to claim 1, characterized in that: The acquiring of the memory data of the server includes: Identifying the amount of memory in the server; Setting a memory state array according to the memory quantity; The memory status array is used to record the memory data of the server.

3. The server warning function management method according to claim 2, characterized in that: The utilizing the memory status array to record the memory data of the server includes: Identifying the number of recorded data in the memory status array; The recording of the memory data of the server is controlled according to the amount of the recorded data.

4. The server warning function management method according to claim 3, characterized in that: Controlling the recording of memory data of the server according to the amount of the recorded data includes: If the amount of the recorded data is less than the memory amount, the memory data of the server is controlled to be recorded until the amount of the recorded data reaches the memory amount, and then the recording is stopped.

5. The server warning function management method according to claim 1, characterized in that: The performing of a fault recovery action on the memory in the first state includes: Obtaining a fault type of the memory in the first state; The fault recovery action is determined based on the fault type.

6. The server warning function management method according to claim 5, characterized in that: The obtaining of the fault type of the memory in the first state includes: identifying fault information in memory data of the memory in the first state; The fault type is determined based on the fault information.

7. The server warning function management method according to claim 5, characterized in that: The fault types include the first to fifth types, and the fault recovery actions include the first to fifth actions, wherein: The first type is a hardware damage abnormal fault, and the first action is to replace the memory in the first state; The second type is a poor contact abnormal fault, and the second action is to re-plug the memory in the first state, clean the dust of the motherboard slot corresponding to the memory in the first state, or change the position of the memory slot in the first state; The third type is a compatibility abnormality fault, and the third action is to replace the memory with the same model as the memory in the first state; The fourth type is a configuration abnormality fault, and the fourth action is to restore the default parameters of the memory in the first state; The fifth type is an abnormal fault with an error report, and the fifth action is to detect the memory in the first state and repair or replace the memory in the first state.

8. A server, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the server warning function management method according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server warning function management method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the server warning function management method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Power supply control method and device of server, storage medium and electronic equipment

    CN118642584A

  • Fault processing method and device, equipment, storage medium and computer program product

    CN119396613A