Server warning function management method, server, storage medium and program product
By turning off the warning function when the server is started, obtaining and restoring the fault memory status, and enabling the warning function after all memory is normal, the downtime caused by memory failure is solved, reducing R&D time and maintaining the server's fault alarm capability.
Patent Information
- Application Number
- CN202510884211.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
During the server development phase, downtime caused by memory failures increases the workload of R&D and testers, and turning off the warning function will reduce the server's fault alarm capability.
Turn off the warning function when the server starts, obtain the memory status, restart after the failed memory, wait until all memory is normal, and then enable the warning function through the BMC and BIOS interactive control of the warning function switch.
It avoids server downtime caused by memory failure, reduces the time for R&D personnel to find out the downtime problem, and maintains the server's fault alarm capability, achieving a balance between avoiding downtime problems and alarm capabilities.
Smart Images

Figure CN120386694A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method for managing a server warning function, a server, a storage medium, and a program product. Background Art
[0002] The warning function Promote Warnings of a server is a function switch for system prompt warnings provided by the CPU (Central Processing Unit) in the server field. When this function is enabled, starting from the system level, it can provide more refined and comprehensive warnings when the server fails. Therefore, in the BIOS (Basic Input Output System) CRB (Customer Reference Board) program code provided by the CPU manufacturer, this system-level warning function is generally enabled. However, the starting point of this design is based on the public board and CRB code of the CPU manufacturer. The various configurations use the minimum startup configuration, and the memory is also configured with 1 or 2 memory modules. Memory failure problems rarely occur.
[0003] However, during the server product development stage by each OEM (Original Equipment Manufacturer), on the one hand, multiple memory modules or even full-configured memory are used for testing in the server machine configuration; on the other hand, the memory used during the product development stage is basically samples, and there are many problems with the memory. In both cases, the possibility of memory failure is greatly increased. When the Promote Warnings function is enabled, if there is a faulty memory, during the memory initialization process when the server boots, encountering a memory failure will cause the server to crash and fail to boot normally. Especially during the product development stage, there are various faults, resulting in more server crashing problems, which causes R & D personnel to constantly search for crashing problems and requires a large amount of time and effort from R & D and testing personnel. Therefore, the common operation of OEM manufacturers is to disable this Promote Warnings function, but this often sacrifices the warning ability of the server when a failure occurs in order to solve the crashing problem caused by memory failure. Summary of the Invention
[0004] The present invention provides a method for managing a server warning function, a server, a storage medium, and a program product, so as to at least solve the problems in the related technologies that the server crashes due to memory failure, thereby increasing the time for finding the crashing problem, and the warning ability of the server is reduced when the warning function is turned off.
[0005] The present invention provides a method for managing the server warning function, comprising the following steps: closing the warning function of the server at the moment of server startup; obtaining the memory data of the server, and extracting the operating states of all memories in the server from the memory data, wherein the operating states include a first state and a second state, the first state represents the fault state of the memory, and the second state represents the normal state of the memory; if there is a first state among the operating states of all memories in the server, performing a fault recovery action on the memory in the first state, and restarting the server after it is recognized that the fault recovery action is completed; if the operating states of all memories in the server are all in the second state, starting the warning function of the server, and using the warning function to perform fault warning on the server.
[0006] The present invention also provides a device for managing the server warning function, comprising: a closing module, configured to close the warning function of the server at the moment of server startup; an extracting module, configured to obtain the memory data of the server and extract the operating states of all memories in the server from the memory data, wherein the operating states include a first state and a second state, the first state represents the fault state of the memory, and the second state represents the normal state of the memory; an executing module, configured to, if there is a first state among the operating states of all memories in the server, perform a fault recovery action on the memory in the first state, and restart the server after it is recognized that the fault recovery action is completed; a starting module, configured to, if the operating states of all memories in the server are all in the second state, start the warning function of the server, and use the warning function to perform fault warning on the server.
[0007] The present invention also provides a server, comprising: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above-mentioned methods for managing the server warning function when executing the computer program.
[0008] The present invention also provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of any one of the above-mentioned methods for managing the server warning function when being executed by a processor.
[0009] The present invention also provides a computer program product, comprising a computer program, which implements the steps of any one of the above-mentioned methods for managing the server warning function when being executed by a processor.
[0010] Through the present invention, during the product development stage of the server, multiple memory modules or even fully configured memory modules are used for testing in the server's machine configuration. Moreover, the memory used during the product development stage is basically samples, and there are many problems with the memory. Therefore, the possibility of memory failures will increase significantly. When the warning function of the server is enabled, if there is a faulty memory module, during the power-on memory initialization process of the server, encountering a memory failure will cause the server to crash and be unable to boot normally. As a result, R & D personnel need to find the problem of the server crashing, which consumes a lot of time and effort of R & D and testing personnel. Or directly turn off the warning function of the server to avoid the server crashing problem caused by memory failures, but this will sacrifice the warning ability of the server for other failures except memory failures. Therefore, the present invention can first turn off the warning function of the server when the server is powered on, restore the faults of the memory in the first state, that is, the faulty memory, and then restart the server. When all the memories of the server are in the second state, that is, when the memory is in a normal state, turn on the warning function of the server, avoiding the server crashing problem caused by memory failures, reducing the time spent by R & D or testing personnel to find the crashing problem, and at the same time ensuring the warning ability of the server for failures, achieving a balance between avoiding the server crashing problem and enhancing the warning ability of the server when a failure occurs. Therefore, it can solve the technical problems in the related art that the server crashes due to memory failures, thereby increasing the time to find the crashing problem, and reducing the warning ability of the server by turning off the warning function, achieving the technical effect of avoiding the server crashing problem caused by memory failures, reducing the time spent by R & D or testing personnel to find the crashing problem, and at the same time ensuring the warning ability of the server for failures, achieving a balance between avoiding the server crashing problem and enhancing the warning ability of the server when a failure occurs. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 It is a flowchart of the server warning function management method provided according to an embodiment of the present invention; Figure 2 It is an execution diagram of the server warning function management method provided according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the fault recovery actions corresponding to the fault types provided according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the server warning function management device provided according to an embodiment of the present invention; Figure 5 The figure is a schematic structural diagram of a server provided according to an embodiment of the present invention. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0014] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and not to describe a specific order or sequence.
[0015] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0016] An embodiment of the present invention provides a method for managing a server warning function. In combination with the execution process of the method for managing the server warning function, the method will be described in detail.
[0017] Figure 1 The figure is a schematic flow diagram of a method for managing a server warning function provided according to an embodiment of the present invention. [[ID=2,1]]
[0018] As Figure 1 shown, the method for managing the server warning function includes the following steps: In step S101, the warning function of the server is turned off at the moment when the server starts up.
[0019] Among them, the warning function can be the Promote Warnings function of the server. Promote Warnings is a function switch provided by the CPU in the server field to provide system prompt warnings, which is used to control whether to promote warnings to the system level when a failure occurs. Thus, starting from the system level, it is possible to make a more refined and comprehensive alarm processing for the server when a failure occurs; the opening or closing of the warning function is controlled by the BMC (Baseboard Manager Controller) of the server.
[0020] It should be noted that the server warning function management method of the present invention is mainly used in the product development stage of the server.
[0021] In step S102, the memory data of the server is obtained, and the operating states of all memories of the server in the memory data are extracted, where the operating states include a first state and a second state, the first state indicates the fault state of the memory, and the second state indicates the normal state of the memory.
[0022] Among them, the memory refers to the memory module in the server hardware, that is, the physical hardware component used by the computer to temporarily store data, and the memory data can be obtained through the BIOS program.
[0023] It can be understood that the embodiments of the present invention can obtain the memory data of the server and extract the operating states of all memories of the server in the memory data to judge the state of each memory in the server.
[0024] In the embodiments of the present invention, obtaining the memory data of the server includes: identifying the number of memories in the server; setting a memory status array according to the number of memories; and using the memory status array to record the memory data of the server.
[0025] Among them, the memory status array is an array used to record the status of each memory, which can be defined as A[n], where n is the memory index.
[0026] It can be understood that the embodiments of the present invention can identify the number of memories in the server, set a memory status array according to the number of memories, and use the memory status array to record the memory data of the server, and uniformly manage all memory data through the array structure, which is convenient for subsequent memory fault judgment.
[0027] For example, if the server is installed with 4 memories, then create arrays A[0]-A[3] to record the status of each memory respectively.
[0028] In the embodiments of the present invention, using the memory status array to record the memory data of the server includes: identifying the number of recorded data in the memory status array; and controlling the recording of the memory data of the server according to the number of recorded data.
[0029] It can be understood that the embodiments of the present invention can identify the memory status and control the recording of the memory data of the server according to the number of recorded data.
[0030] In the embodiments of the present invention, controlling the recording of the memory data of the server according to the number of recorded data includes: if the number of recorded data is less than the number of memories, then control the memory data of the server to be recorded until the number of recorded data reaches the number of memories, and stop recording.
[0031] Among them, the number of memories can be represented by M, and the number of recorded data can be represented by n.
[0032] It can be understood that in the embodiments of the present invention, when the number of recorded data is less than the number of memories, that is, when n is less than M, the memory data of the server can be controlled for recording until the number of recorded data reaches the number of memories M, and then the recording is stopped and subsequent steps are performed, so as to avoid missing or repeating the recording of the memory state and ensure the accuracy of subsequent memory fault judgment.
[0033] For example, the server has 4 memories, and n starts recording from 0 and stops recording when n = 3 to ensure that the states of all 4 memories are recorded.
[0034] In step S103, if there is a first state among the operating states of all memories in the server, a fault recovery action is performed on the memory in the first state, and after it is recognized that the fault recovery action is completed, the server is restarted.
[0035] Among them, the memory in the first state can be called a faulty memory.
[0036] It can be understood that in the embodiments of the present invention, when there is a first state among the operating states of all memories in the server, a fault recovery action is performed on the memory in the first state, that is, the faulty memory, and after it is recognized that the fault recovery action is completed, the server is restarted, so as to achieve the elimination and recovery of the faulty memory.
[0037] In the embodiments of the present invention, performing a fault recovery action on the memory in the first state includes: obtaining the fault type of the memory in the first state; determining the fault recovery action based on the fault type.
[0038] It can be understood that in the embodiments of the present invention, the fault type of the memory in the first state can be obtained, and the fault recovery action can be determined based on the fault type to accurately determine the fault recovery action of the faulty memory.
[0039] In the embodiments of the present invention, obtaining the fault type of the memory in the first state includes: identifying the fault information in the memory data of the memory in the first state; determining the fault type based on the fault information.
[0040] It can be understood that in the embodiments of the present invention, the fault type can be determined based on the fault information in the memory data of the memory in the first state.
[0041] In the embodiments of the present invention, determining the fault type based on the fault information includes: obtaining the correspondence between the fault information and the fault type; querying the correspondence based on the fault information to determine the fault type.
[0042] It can be understood that the embodiments of the present invention can determine the fault type of the faulty memory according to the correspondence between the fault information and the fault type, so as to improve the accuracy of faulty memory diagnosis.
[0043] For example, the fault information in the memory data of the faulty memory is A, and by querying the corresponding relationship, the fault type is determined to be A1.
[0044] In the embodiments of the present invention, the fault types include the first to fifth types, and the fault recovery actions include the first to fifth actions.
[0045] It can be understood that the fault types in the embodiments of the present invention include multiple types, and the corresponding fault recovery actions can be accurately matched according to each fault type to achieve the fault recovery of the faulty memory.
[0046] In the embodiments of the present invention, the first type is a hardware damage abnormal fault, and the first action is to replace the memory in the first state; the second type is a poor contact abnormal fault, and the second action is to re-plug and unplug the memory in the first state, clean the dust of the motherboard slot corresponding to the memory in the first state, or change the position of the memory slot of the first state; the third type is a compatibility abnormal fault, and the third action is to replace the memory of the same model as the memory in the first state; the fourth type is a configuration abnormal fault, and the fourth action is to restore the default parameters of the memory in the first state; the fifth type is an error reporting abnormal fault, and the fifth action is to detect the memory in the first state and repair or replace the memory in the first state.
[0047] Among them, the hardware damage abnormal fault may be a damaged memory module; the poor contact abnormal fault may be poor contact between the memory module and the motherboard slot, etc.; the compatibility abnormal fault may be incompatibility between the installed different models or different brands of memory and the server; the configuration abnormal fault may be improper configuration of memory frequency or parameters; the error reporting abnormal fault may be an ECC (Error Checking and Correction) error reporting fault of the memory recorded in the log.
[0048] It can be understood that the embodiments of the present invention can perform different fault recovery actions according to the specific type of the faulty memory, specifically: If the fault type is the first type, which is a hardware damage problem, the fault recovery action is the first action, that is, to replace the memory to completely solve the physical fault and prevent the subsequent repeated downtime of the server; If the fault type is the second type, which is a physical problem such as poor contact between the memory module and the motherboard slot, the fault recovery action is the second action, that is, to re-plug and unplug the memory module, clean the dust of the motherboard slot, or change the slot position to repair the physical connection problem without replacing the memory; If the fault type is the third type, which is a memory compatibility fault where different models or brands of memory are installed, causing system instability or some memory not being recognized, then the fault recovery action is the third action, which is to replace the memory with the same model to solve the system instability caused by inconsistent memory specifications; If the fault type is the fourth type, which is a configuration fault where the memory frequency or parameter configuration is improper, resulting in abnormal display of the system memory frequency or capacity, then the fault recovery action is the fourth action, which is to clear the custom settings in the BIOS options by using the motherboard jumper or clearing the CMOS (Complementary Metal-Oxide-Semiconductor) battery to restore the default parameters and repair the memory fault caused by incorrect manual configuration; If the fault type is the fifth type, which is a memory ECC error reporting fault recorded in the log, then the fault recovery action is the fifth action, which is to use the corresponding diagnostic tool to detect the memory module or replace the memory module with an ECC fault to locate and repair the memory data error problem.
[0049] In step S104, if the operating status of all the memory in the server is the second status, then the warning function of the server is started to perform fault warning on the server by using the warning function.
[0050] Since in the product development stage of the server, multiple memory modules or even full-configured memory are used for testing in the machine configuration of the server, and the memory used in the product development stage is basically samples with many problems, the possibility of memory faults will increase greatly. When the warning function of the server is turned on, if there is a faulty memory, during the memory initialization process when the server boots up, the memory fault will cause the server to crash and fail to boot normally. As a result, the R & D personnel need to find the problem of the server crashing, which takes a lot of time and effort of the R & D and testing personnel, or directly turn off the warning function of the server to avoid the server crashing problem caused by memory faults, but this will sacrifice the warning ability of the server for other faults except memory faults. Therefore, in the embodiment of the present invention, the warning function of the server is first turned off when the server is started. After the fault of the memory in the first status, that is, the faulty memory, is recovered, the server is restarted. When all the memory in the server is in the second status, that is, when the memory is in the normal status, the warning function of the server is turned on, avoiding the server crashing problem caused by memory faults, reducing the time for the R & D or testing personnel to find the crashing problem, and at the same time ensuring the warning ability of the server for faults, achieving a good balance between avoiding the server crashing problem and improving the warning ability of the server when a fault occurs.
[0051] In summary, the server warning function management method proposed in the embodiments of the present invention can intelligently adjust the warning function of the server. During the server startup process, the BMC and the BIOS will perform information interaction, place the function switch option for controlling this warning function in the BMC, and set the initial value to turn off this warning function. Each time the server is powered on, the BIOS program obtains the function status of this warning function from the BMC. At this time, if there is a faulty memory, the problem of system downtime can also be avoided. During the startup process, the memory status is captured. If there is no memory abnormality, the memory status is sent to the BMC, and the BMC sets the function switch option to turn off this function. When the server is powered on next time, the status obtained from the BMC is open, so as to improve the alarm level ability of the system and avoid resource waste caused by memory without abnormalities. If there is a memory fault, the problem of the faulty memory is solved to make the machine in a state without memory faults. In this way, a balance is achieved between the server downtime problem caused by the memory fault due to the opening of this function and the improvement of the alarm level ability of the system when a fault occurs as a whole, avoiding the problem of system downtime and improving the alarm level ability of the system when a fault occurs as a whole.
[0052] Specifically, the server warning function management method in the embodiments of the present invention is described by taking the Promote Warnings warning function as an example. The specific process is as Figure 2 shown and includes: 1. Power on the server. The BMC sets the function switch option Promote Warnings for controlling this function to the off state.
[0053] 2. Detect the number of memories in the current machine, denoted as M, and define the status of the memory as A[n] (n = n + 1, n = 0, 1, 2...).
[0054] 3. Record the status of each memory as A[n] (n = n + 1, n = 0, 1, 2...).
[0055] 4. When n is not greater than M, it means that there are still memories whose status has not been recorded. Return to step 3 and continue to record the memories whose status has not been recorded.
[0056] 5. When n reaches M, it means that the status of all memories has been recorded. At this time, stop recording the memory status.
[0057] 6. Judge whether there is a faulty memory in the status of all memories, A[n].
[0058] 7. There are various types of memory faults. If there is a faulty memory, relevant processing is performed according to the memory fault type information recorded in A[n]. The specific processing actions for different fault types are as Figure 3 shown and include: (1) If the memory has a physical hardware damage problem, replace the memory; (2) If there are physical problems such as poor contact between the memory module and the motherboard slot, re-plug the memory module, clean the dust in the motherboard slot, or change the slot position. (3) If memory compatibility failures occur due to installing memory modules of different models or brands, resulting in system instability or some memory not being recognized, replace the memory with the same model. (4) If configuration failures occur due to improper memory frequency or parameter configuration, resulting in abnormal display of system memory frequency or capacity, clear the custom settings in the BIOS options by using the motherboard jumper or removing the CMOS battery to restore the default parameters. (5) If memory ECC error reports are recorded in the system log, use the corresponding diagnostic tool to detect the memory module or replace the memory module with ECC failures.
[0059] 8. After dealing with the memory failure problem, power on the computer and start from step 1 again.
[0060] 9. If all memory is normal and there is no abnormal memory failure, send a memory normal signal to the BMC, and trigger the BMC side to set the control function switch option Promote Warnings to the on state to improve the warning level made by the system.
[0061] The following uses a specific embodiment to describe the server warning function management method of the embodiments of the present invention.
[0062] An OEM manufacturer is developing a new server with the following configuration: CPU (Intel Xeon processor supporting the Promote Warnings function), memory (8 pieces, all samples), BIOS (a customized version based on the CRB reference board), and BMC.
[0063] During the testing phase, the server frequently failed to boot. Therefore, the server warning function management method of the present invention is applied to solve the problem, and the specific steps are as follows: 1. Server startup and warning function off.
[0064] When the server is powered on and booted, the BMC automatically sets the Promote Warnings function switch to the off state. Even if there is a memory failure, the system will not immediately crash due to warning escalation, allowing subsequent detection and repair.
[0065] 2. Memory status detection and recording.
[0066] (1) The BIOS detects the number of memory: Through hardware scanning, it is identified that the server has installed 8 memory modules (M = 8).
[0067] (2)Create a memory status array: Define an array A[0] - A[7] to record the status of each memory.
[0068] (3)Loop to detect the memory status: The counter n starts from 0 and detects each memory in turn. When n = 3, it is detected that the 4th memory (A[3]) has an ECC check error, which is recorded as a fault status. Continue to detect the remaining memories, and finally stop detecting when n = 8, and confirm that A[3] is the only faulty memory.
[0069] 3. Fault type identification and handling.
[0070] (1)Obtain fault information: The system log shows that the ECC error code of A[3] is 0x502 (for example, indicating that a single-bit error is irrecoverable).
[0071] (2)Determine the fault type: Through the preset corresponding relationship, map the error code 0x502 to the fifth type of fault (ECC error reporting exception).
[0072] (3)Execute the fault recovery action: Automatically call the memory diagnostic tool to detect A[3], confirm that the memory module is physically damaged, and prompt the operation and maintenance personnel to replace the memory module corresponding to A[3].
[0073] 4. Restart the server and enable the warning function.
[0074] (1)Restart after replacing the faulty memory: After the operation and maintenance personnel replace A[3], the server is powered on again.
[0075] (2)Redetect the memory status: Repeat step 2 to confirm that the status of all memories (A[0] - A[7]) is normal.
[0076] (3)Enable the warning function: The BMC receives the signal "All memories are normal" and automatically enables the PromoteWarnings function.
[0077] Thus, the server boots up normally, and if a memory fault occurs during subsequent operation, the system will give a timely warning through the PromoteWarnings function.
[0078] In this embodiment, although there is a faulty memory when the server is powered on for the first time, since the warning function is turned off, the server does not crash immediately and is allowed to complete fault detection and repair. Through the mapping of error codes to fault types, it is quickly determined that A [3] is a hardware damage, avoiding blind attempts at other repair methods. After the repair, the warning function is automatically turned on to ensure that the system has a complete warning capability during subsequent operation. Compared with traditional methods (which require manual repeated troubleshooting and manual switching of functions), this embodiment shortens the fault handling time, thus realizing the automatic shutdown function - detecting faults - repairing - restarting - automatically enabling the function, forming a closed-loop management without the need for manual intervention in the switching of the warning function.
[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0080] The server warning function management method proposed according to the embodiments of the present invention can turn off the warning function of the server when the server is started, restore the faults of the memory in the first state, that is, the faulty memory, and then restart the server. When all the memories of the server are in the second state, that is, when the memory is in a normal state, the warning function of the server is turned on, avoiding the problem of server crashing caused by memory faults, reducing the time spent by R & D or test personnel in finding the crashing problem, and at the same time ensuring the warning capability of the server for faults, achieving a balance between avoiding server crashing problems and improving the warning capability of the server when a fault occurs.
[0081] An embodiment of the present invention also provides a server warning function management device.
[0082] Figure 4 It is a block diagram of the server warning function management device provided according to the embodiments of the present invention.
[0083] As Figure 4 shown, the server warning function management device 10 includes: a shutdown module 100, an extraction module 200, an execution module 300, and a startup module 400.
[0084] Among them, the shutdown module 100 is used to shut down the warning function of the server at the moment when the server starts; the extraction module 200 is used to obtain the memory data of the server and extract the running status of all memories in the server from the memory data, where the running status includes a first status and a second status, the first status indicates the fault status of the memory, and the second status indicates the normal status of the memory; the execution module 300 is used to perform a fault recovery action on the memory in the first status if there is a first status among the running statuses of all memories in the server, and restart the server after recognizing that the fault recovery action is completed; the startup module 400 is used to start the warning function of the server if the running statuses of all memories in the server are all in the second status, and perform a fault warning on the server by using the warning function.
[0085] In an embodiment of the present invention, the extraction module 200 is further used to: identify the number of memories in the server; set a memory status array according to the number of memories; and record the memory data of the server by using the memory status array.
[0086] In an embodiment of the present invention, the extraction module 200 is further used to: identify the number of recorded data in the memory status array; and control the recording of the memory data of the server according to the number of recorded data.
[0087] In an embodiment of the present invention, the extraction module 200 is further used to: if the number of recorded data is less than the number of memories, control the memory data of the server to be recorded until the number of recorded data reaches the number of memories, and then stop recording.
[0088] In an embodiment of the present invention, the execution module 300 is further used to: obtain the fault type of the memory in the first status; and determine a fault recovery action based on the fault type.
[0089] In an embodiment of the present invention, the execution module 300 is further used to: identify the fault information in the memory data of the memory in the first status; and determine the fault type based on the fault information.
[0090] In an embodiment of the present invention, the fault type includes a first to a fifth type, and the fault recovery action includes a first to a fifth action.
[0091] In an embodiment of the present invention, the first type is a hardware damage abnormal fault, and the first action is to replace the memory in the first status.
[0092] In an embodiment of the present invention, the second type is a poor contact abnormal fault, and the second action is to re-plug the memory in the first status, clean the dust in the motherboard slot corresponding to the memory in the first status, or change the position of the memory slot of the memory in the first status.
[0093] In an embodiment of the present invention, the third type is a compatibility abnormal fault, and the third action is to replace the memory with the same model as the memory in the first status.
[0094] In an embodiment of the present invention, the fourth type is a configuration anomaly fault, and the fourth action is to restore the default parameters of the memory in the first state.
[0095] In an embodiment of the present invention, the fifth type is an error anomaly fault, and the fifth action is to detect the memory in the first state and repair or replace the memory in the first state.
[0096] It should be noted that for the description of the features in the corresponding embodiment of the server warning function management device, reference can be made to the relevant description in the corresponding embodiment of the server warning function management method, which will not be elaborated here one by one.
[0097] According to the server warning function management device provided by the embodiment of the present invention, when the server is started, the warning function of the server can be first turned off. After the faults of the memories in the first state, that is, the faulty memories, are recovered, the server is restarted. When all the memories of the server are in the second state, that is, when the memories are in the normal state, the warning function of the server is turned on, avoiding the problem of server downtime caused by memory faults, reducing the time spent by R & D or test personnel to find the downtime problem, and at the same time ensuring the server's alarm ability for faults, achieving a good balance between avoiding server downtime problems and improving the alarm ability of the server when a fault occurs.
[0098] An embodiment of the present invention also provides a server, as Figure 5 shown, including a memory 501 and a processor 502. A computer program is stored in the memory 501, and the processor 502 is configured to run the computer program to execute the steps in any of the above embodiments of the server warning function management method.
[0099] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the server warning function management method when running.
[0100] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0101] An embodiment of the present invention also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the server warning function management method.
[0102] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0103] The above has introduced in detail a server warning function management provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A server warning function management method, characterized in that, It includes the following steps: Turn off the warning function of the server at the moment when the server starts up; Obtain the memory data of the server, and extract the running status of all memories of the server from the memory data, where the running status includes a first status and a second status, the first status represents the fault status of the memory, and the second status represents the normal status of the memory; If the first status exists in the running status of all memories in the server, perform a fault recovery action on the memory in the first status, and restart the server after it is recognized that the fault recovery action is completed; If the running status of all memories in the server is the second status, turn on the warning function of the server, and use the warning function to perform a fault warning on the server.
2. The server warning function management method according to claim 1, wherein The obtaining of the memory data of the server includes: Identify the number of memories in the server; Set a memory status array according to the number of memories; Use the memory status array to record the memory data of the server.
3. The server warning function management method according to claim 2, wherein, The using the memory status array to record the memory data of the server includes: Identify the number of recorded data in the memory status array; Control the recording of the memory data of the server according to the number of recorded data.
4. The server warning function management method according to claim 3, characterized in that Controlling the recording of the memory data of the server according to the number of recorded data includes: If the number of recorded data is less than the number of memories, control the memory data of the server to be recorded until the number of recorded data reaches the number of memories, and then stop recording.
5. The server warning function management method according to claim 1, wherein The performing a fault recovery action on the memory in the first status includes: Obtain the fault type of the memory in the first status; Determine the fault recovery action based on the fault type.
6. The server warning function management method according to claim 5, wherein The obtaining the fault type of the memory in the first status includes: Identify the fault information in the memory data of the memory in the first status; Determine the fault type based on the fault information.
7. The server warning function management method according to claim 5, characterized in that The fault type includes the first to fifth types, and the fault recovery action includes the first to fifth actions.
8. The server warning function management method according to claim 5, wherein The first type is a hardware damage abnormal fault, and the first action is to replace the memory in the first status.
9. The server warning function management method according to claim 5, characterized in that, The second type is a poor contact abnormal fault, and the second action is to re-plug the memory in the first status, clean the dust in the motherboard slot corresponding to the memory in the first status, or change the position of the memory slot of the memory in the first status.
10. The server warning function management method according to claim 5, characterized in that, The third type is a compatibility abnormal fault, and the third action is to replace the memory with the same model as the memory in the first status.
11. The server warning function management method according to claim 5, wherein The fourth type is a configuration abnormal fault, and the fourth action is to restore the default parameters of the memory in the first status.
12. The server warning function management method according to claim 5, wherein, The fifth type is an error reporting abnormal fault, and the fifth action is to detect and repair the memory in the first status or replace the memory in the first status.
13. A server, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the server warning function management method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the server warning function management method according to any one of claims 1 to 12 when executed by a processor.
15. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the server warning function management method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Memory fault processing method and computing device
CN115269245A
Server startup processing method and device, electronic equipment and storage medium
CN118152182A
Power supply control method and device of server, storage medium and electronic equipment
CN118642584A
Fault processing method and device, equipment, storage medium and computer program product
CN119396613A
Server network diagnostic system
US20110066895A1