BMC configuration management method, device and equipment and readable storage medium

By implementing the BMC configuration management method in the server, responding to runtime status change events, assessing health and automatically updating configuration restore points, the systemic risks caused by server configuration changes are resolved, and operational efficiency and business recovery capabilities are improved.

CN120803540APending Publication Date: 2025-10-17XINHUASAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510766191.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In existing technologies, systemic risks caused by server configuration changes, including BMC network failures, loss of management user information, firewall policy tampering, performance degradation and compatibility issues caused by firmware parameter configuration, cannot generate restore points in a timely and accurate manner, resulting in excessively long recovery times and recovery failures.

Method used

By implementing the BMC configuration management method in the server, it responds to runtime status change events, assesses the health of BMC operation, automatically updates the most recent correct or incorrect configuration restore point, and quickly calls the correct configuration restore point when the system is abnormal, thereby achieving intelligent and dynamic configuration management.

Benefits of technology

It enables timely updates of restore points when configuration changes occur, reducing the risk of data loss and service interruption, improving operational efficiency, ensuring rapid business recovery, and providing a basis for problem localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803540A_ABST
    Figure CN120803540A_ABST
Patent Text Reader

Abstract

The invention provides a BMC configuration management method, device and equipment and a readable storage medium, and the method comprises the steps: responding to a change event of an operation state, and evaluating the current BMC operation health degree according to a preset algorithm; updating the latest correct configuration restoration point according to the currently used BMC configuration; and updating the latest error configuration restoration point according to the currently used BMC configuration, and calling and restoring the BMC configuration according to the latest correct configuration restoration point. According to the technical scheme of the invention, the most recent correct and wrong configuration restoration points can be updated in time by responding to the operation state change event and evaluating the BMC operation health degree. When the system is abnormal, the most recent correct configuration restoration point can be quickly called to restore the BMC configuration, so that the risks of data loss and service interruption caused by wrong configuration are effectively reduced, the operation and maintenance efficiency of a server is improved, quick service restoration is assisted, and a basis is provided for problem positioning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of communication, and in particular, to a BMC configuration management method, device, equipment and readable storage medium. BACKGROUND

[0002] In practical applications, server configuration changes have become a key factor leading to systematic risks, resulting in a series of problems, including but not limited to out-of-band management interruption caused by BMC network failure, information loss of management users, tampering of firewall black and white list security policies, performance reduction or compatibility problems of components related to firmware parameter configuration, and data integrity damage caused by RAID card cache strategy damage. These problems are basically included in BMC configuration, BIOS configuration and RAID card configuration. Statistics show that the main server unplanned stop events are directly related to configuration changes, which seriously affect the continuity and stability of critical business.

[0003] The current industry generally adopts a user manual backup configuration file mechanism to archive server system historical configuration information, which has obvious defects. First, it is impossible to generate a reliable restore point, because the restore point created by the user may not be timely and accurate enough; second, the traditional rollback scheme requires manual intervention in the process, and the average recovery time often exceeds the industry-acceptable recovery time objective threshold. In addition, the process of manually exporting and importing configuration files is tedious, and recovery failure may occur due to damage to local configuration files, which cannot ensure effective configuration backup at critical time points. SUMMARY

[0004] Therefore, the present specification provides a BMC configuration management method, device, equipment and readable storage medium to improve the problem of inaccurate and timely creation of configuration restore points.

[0005] The specific technical solutions are as follows:

[0006] The present specification provides a BMC configuration management method applied to a server, which comprises: in response to a change event of a running state, evaluating a current BMC running health degree according to a preset algorithm; in response to an event that the current BMC running health degree is greater than a preset health threshold, updating a last correct configuration restore point according to a currently used BMC configuration; in response to an event that the current BMC running health degree is less than a preset error threshold, updating a last error configuration restore point according to the currently used BMC configuration, and calling and restoring the configuration of the BMC according to the last correct configuration restore point.

[0007] As a technical solution, the method further comprises: in response to the diagnosis signaling, calling and comparing the last correct configuration restoration point and the last error configuration restoration point, and obtaining and displaying the difference between the last correct configuration restoration point and the last error configuration restoration point.

[0008] As a technical solution, in response to the change event of the running state, the current BMC running health degree is evaluated according to a preset algorithm, which comprises: in response to a configuration change event and / or a load abnormal change event and / or a hardware change event, the current BMC running health degree is evaluated according to a preset weighting algorithm, and the input parameters of the preset weighting algorithm are associated with the current configuration and / or the current load and / or the current hardware.

[0009] As a technical solution, in response to the event that the current BMC running health degree is greater than a preset health threshold, the last correct configuration restoration point is updated according to the currently used BMC configuration, which comprises: in response to the event that the current BMC running health degree is greater than the preset health threshold, the existing last correct configuration restoration point is stored after being named according to a preset rule, and the last correct configuration restoration point is updated according to the currently used BMC configuration.

[0010] The specification also provides a BMC configuration management device applied to a server, which comprises: a first module for evaluating the current BMC running health degree according to a preset algorithm in response to a change event of a running state; a second module for updating the last correct configuration restoration point according to the currently used BMC configuration in response to an event that the current BMC running health degree is greater than a preset health threshold; and a third module for updating the last error configuration restoration point according to the currently used BMC configuration in response to an event that the current BMC running health degree is less than a preset error threshold, calling and restoring the configuration of the BMC according to the last correct configuration restoration point.

[0011] As a technical solution, the device further comprises: a fourth module for calling and comparing the last correct configuration restoration point and the last error configuration restoration point in response to diagnosis signaling, and obtaining and displaying the difference between the last correct configuration restoration point and the last error configuration restoration point.

[0012] As a technical solution, in response to the change event of the running state, the current BMC running health degree is evaluated according to a preset algorithm, which comprises: in response to a configuration change event and / or a load abnormal change event and / or a hardware change event, the current BMC running health degree is evaluated according to a preset weighting algorithm, and the input parameters of the preset weighting algorithm are associated with the current configuration and / or the current load and / or the current hardware.

[0013] As a technical solution, the method comprises the following steps: in response to an event that the current BMC running health degree is greater than a preset health threshold, updating a latest correct configuration restoration point according to a currently used BMC configuration, including: in response to the event that the current BMC running health degree is greater than the preset health threshold, storing an existing latest correct configuration restoration point after naming the existing latest correct configuration restoration point according to a preset rule, and updating the latest correct configuration restoration point according to the currently used BMC configuration.

[0014] The present specification also provides an electronic device, comprising a processor and a readable storage medium, wherein the readable storage medium stores machine executable instructions capable of being executed by the processor, and the processor executes the machine executable instructions to implement the aforementioned BMC configuration management method.

[0015] The present specification also provides a readable storage medium, wherein the readable storage medium stores machine executable instructions, and the machine executable instructions, when invoked and executed by a processor, cause the processor to implement the aforementioned BMC configuration management method.

[0016] The present specification provides the aforementioned technical solutions, which at least have the following beneficial effects:

[0017] By responding to the running state change event and evaluating the BMC running health degree, the latest correct and incorrect configuration restoration points can be updated in time. When the system is abnormal, the latest correct configuration restoration point can be quickly called to restore the BMC configuration, thereby effectively reducing the risk of data loss and service interruption caused by incorrect configuration, improving the server operation and maintenance efficiency, assisting the rapid recovery of business, and providing a basis for problem positioning. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the following will briefly introduce the drawings needed in the embodiments of the present specification or the prior art description. Obviously, the drawings in the following description are only some embodiments described in the present specification, and other drawings can also be obtained by those skilled in the art according to these drawings.

[0019] Figure 1 is a flowchart of the BMC configuration management method in an embodiment of the present specification;

[0020] Figure 2 is a structural diagram of the BMC configuration management device in an embodiment of the present specification;

[0021] Figure 3 is a structural diagram of the BMC configuration management device in an embodiment of the present specification;

[0022] Figure 4is a hardware structure diagram of an electronic device in an embodiment of the present specification.

[0023] Reference signs: first module 21, second module 22, third module 23, fourth module 24. DETAILED DESCRIPTION

[0024] The terminology used in the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the present specification. As used in the present specification and claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0025] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various information, the information should not be limited to these terms. These terms are only used to distinguish one piece of information from another. For example, a first information can also be termed a second information, and, similarly, a second information can also be termed a first information, without departing from the scope of the present specification. Furthermore, the word "if' can be interpreted as meaning "when" or "in response to determining" depending on the context.

[0026] In view of the above, the present specification provides a BMC configuration management method, device, equipment and readable storage medium to at least improve one of the above technical problems.

[0027] The specific technical solutions are described as follows.

[0028] In an embodiment, the present specification provides a BMC configuration management method applied to a server, the method comprising: in response to a change event of a running state, evaluating a current BMC running health degree according to a preset algorithm; in response to an event that the current BMC running health degree is greater than a preset health threshold, updating a last correct configuration restoration point according to a currently used BMC configuration; in response to an event that the current BMC running health degree is less than a preset error threshold, updating a last error configuration restoration point according to the currently used BMC configuration, and calling and restoring the configuration of the BMC according to the last correct configuration restoration point.

[0029] Specifically, as Figure 1 comprises the following steps:

[0030] Step S11, in response to a change event of a running state, evaluating a current BMC running health degree according to a preset algorithm.

[0031] During the running of the server, whenever a change event of the BMC configuration is detected, the system automatically starts a preset algorithm to evaluate the running health degree of the current BMC. The algorithm comprehensively considers multiple factors such as network connection status, user access response time, error log quantity, etc., and obtains a real-time health score through weighted calculation. This score can reflect the good or bad degree of the current running state of the BMC, and is an important basis for subsequent operations. For example, if after a change, the BMC has frequent network interruption phenomenon, which will cause the health score to decrease significantly; on the contrary, if the change does not cause any abnormality, the health score remains at a high level.

[0032] Step S12, in response to an event that the current BMC running health degree is greater than a preset health threshold, updating the latest correct configuration restoration point according to the currently used BMC configuration.

[0033] When the BMC running health degree is greater than the preset health threshold, it indicates that the current configuration is reliable and does not affect the normal operation of the system, and the latest correct configuration restoration point is updated according to the currently used BMC configuration, so as to facilitate quick recovery in the future when problems occur.

[0034] For example, after performing firmware upgrade, if all monitoring indicators show normal, the system will automatically generate an update of the "latest correct configuration" snapshot, saving this successful change. This not only records the latest stable configuration, but also provides a solution for future possible problems. At the same time, since these snapshots can be stored on non-volatile storage media such as BMC flash or embedded eMMC, even if power failure occurs, data will not be lost, further enhancing the safety and reliability of data.

[0035] Step S13, in response to an event that the current BMC running health degree is less than a preset error threshold, updating the latest error configuration restoration point according to the currently used BMC configuration, and calling and restoring the configuration of the BMC according to the latest correct configuration restoration point.

[0036] If the current BMC running health degree is lower than the preset error threshold, it means that the existing configuration has a problem, which may cause system instability or service interruption, and the system updates the latest error configuration restoration point according to the currently used BMC configuration. Unlike the correct configuration, this restoration point is only updated once when a problem is detected, and is mainly used to record the specific configuration situation that causes the fault.

[0037] For example, if a certain firmware parameter adjustment results in a serious performance degradation, the system will capture this change at the first time and generate an error configuration snapshot. This allows the administrator to easily backtrack to the state before the failure occurs, compare the differences between the two restore points, and quickly locate the root cause of the problem. In addition, by invoking and restoring the configuration of the BMC according to the last correct configuration restore point, normal service can be restored in the shortest time, reducing the loss caused by configuration errors.

[0038] In an embodiment, the method further comprises: in response to the diagnostic signaling, invoking and comparing the last correct configuration restore point and the last error configuration restore point, and obtaining and displaying the differences between the last correct configuration restore point and the last error configuration restore point.

[0039] In an embodiment, in response to the change event of the running state, the current BMC running health degree is evaluated according to a preset algorithm, which comprises: in response to the configuration change event and / or the abnormal change event of the load and / or the hardware change event, the current BMC running health degree is evaluated according to a preset weighting algorithm, and the input parameters of the preset weighting algorithm are associated with the current configuration and / or the current load and / or the current hardware.

[0040] In an embodiment, in response to the event that the current BMC running health degree is greater than the preset health threshold, the last correct configuration restore point is updated according to the currently used BMC configuration, which comprises: in response to the event that the current BMC running health degree is greater than the preset health threshold, the existing last correct configuration restore point is stored after being named according to a preset rule, and the last correct configuration restore point is updated according to the currently used BMC configuration.

[0041] In an embodiment, intelligent and dynamic management of the server BMC configuration is implemented, so that when the server configuration changes, the running health degree of the current BMC can be accurately evaluated in a timely manner, and the corresponding configuration restore point is dynamically updated according to the evaluation result, thereby providing strong support for quickly recovering the correct BMC configuration.

[0042] First, a set of algorithms specially used for evaluating the BMC running health degree is preset in the BMC of the server. The algorithm comprehensively considers multiple indicators closely related to the running state of the BMC, such as the connectivity of the BMC network, the response delay, the integrity and consistency of the management user information, the effectiveness of the firewall black and white list security policy, the rationality of the firmware parameter configuration, and the stability of the RAID card cache strategy, etc. Through real-time monitoring and data collection of these key indicators, and then according to the weights and calculation rules set in the preset algorithm for comprehensive processing, a quantitative value reflecting the current BMC running health condition, i.e. the BMC running health degree, is obtained.

[0043] When a change event occurs in the running state of the server, for example, the server performs an update operation of BMC configuration, system restart or detects a change in some key hardware parameters, etc., the BMC triggers the evaluation process of the current BMC running health degree. At this time, the preset health evaluation algorithm starts to play a role, and according to the real-time data of the above-mentioned various monitoring indicators, the current BMC running health degree value is calculated.

[0044] The calculated BMC running health degree is compared with the preset health threshold. This preset health threshold is a benchmark value set according to the ideal health state of the BMC when the server is running normally and the requirements of the business on the server performance and stability, which usually represents that the BMC runs at a relatively stable and reliable level.

[0045] If the current BMC running health degree is greater than or equal to the preset health threshold, it means that the configuration of the BMC at this time is in a good state and can support the normal and stable operation of the server. According to the currently used BMC configuration, the latest correct configuration restore point stored in the local flash memory or embedded eMMC of the BMC or other non-volatile storage medium is updated.

[0046] This update operation can not simply replace the original data, and a data backup and verification mechanism can be used to ensure that the new correct configuration restore point can completely and accurately save the configuration parameters of the current BMC, and will not interfere with the normal operation of the server during the storage process. For example, during the update process, the system can first generate a temporary snapshot file of the current BMC configuration data, then perform integrity check and format verification on the snapshot file, and only after confirming that there is no error, the snapshot file will be stored as the new latest correct configuration restore point, and at the same time, the previous several historical correct configuration restore points are retained, so that the user can have multiple choices for rollback operation when needed.

[0047] If the current BMC running health degree is less than the preset error threshold, it means that the BMC configuration of the server has a more serious problem, which may cause various abnormal situations of the server, such as network management function failure, security policy failure, etc. At this time, on the one hand, the latest error configuration restore point is updated according to the currently used BMC configuration, and this step is also based on the complete record of the current configuration state, which saves the configuration snapshot when the fault occurs, and provides a basis for subsequent problem analysis and diagnosis; on the other hand, the latest correct configuration restore point saved before is automatically called, and according to the configuration parameters contained in the restore point, the configuration of the BMC is restored.

[0048] In the reduction process, the system will strictly follow the configuration information in the reduction point, gradually restore the settings of the BMC, including network parameters, user permission configuration, security policy, etc. At the same time, during the execution of the reduction operation, the system will also monitor the running indicators of the server in real time to ensure the smooth progress of the reduction operation and the normal work of the BMC after reduction. For example, when restoring the BMC network parameters, the system will first disconnect the original network connection, then reconfigure the network interface according to the IP address, subnet mask, gateway, etc. information in the correct configuration reduction point, and immediately verify the network connectivity after the configuration is completed to ensure that the BMC can restore the external network management function.

[0049] In the user-interactive management interface, the user can manually trigger the creation of a configuration reduction point. For example, before a major BMC configuration upgrade, a manual correct configuration reduction point can be generated in advance to quickly roll back in case of problems during the upgrade. At the same time, the user can also delete, rename, etc. existing reduction points to maintain reasonable use of storage space and orderly management of reduction points. When the system automatically updates the correct or error configuration reduction point, the user will receive corresponding notification information on the management interface, including the time of update, the reason for update, and the comparison of differences between new and old reduction points, etc. so that the user can keep abreast of the changes in the server BMC configuration.

[0050] Taking an application scenario as an example, during the operation of a server in a data center, the network configuration of the BMC was modified due to a misoperation, resulting in a failure of the out-of-band management function of the BMC, which could not be remotely monitored and managed through the network. At this time, the BMC running health assessment algorithm of the server quickly detects the abnormal decrease of network connectivity indicators, and calculates that the current BMC running health is lower than the preset error threshold. The system responds immediately, on the one hand, updates the latest error configuration reduction point according to the current error BMC configuration, and records the error network parameter configuration; on the other hand, automatically calls the latest correct configuration reduction point saved before, and starts to restore the network configuration of the BMC. During the restoration process, the system reconfigures the IP address, subnet mask, gateway, etc. information of the BMC according to the network parameters in the correct reduction point, and verifies the network connectivity. After the restoration is completed, the out-of-band management function of the BMC returns to normal, and the administrator can again perform normal remote management operations on the server through the network, thereby shortening the business interruption time caused by configuration errors to the minimum.

[0051] In the process of firmware upgrade of the server, the new firmware version may have compatibility problems with the existing RAID card configuration, causing the storage performance of the server to decline, and even the risk of data loss. At this time, the running health degree evaluation algorithm of the BMC will calculate the current BMC running health degree to be lower than the error threshold according to the changes of the RAID card cache strategy stability index and the storage performance index. The system then updates the latest error configuration recovery point, records the configuration problems caused by firmware upgrade. At the same time, the latest correct configuration recovery point is called to restore the related configuration parameters of the RAID card to the stable state before the upgrade, so as to ensure the normal operation of the storage system of the server and protect the integrity and security of the data.

[0052] Through the above implementation, the server can realize automatic configuration state monitoring, health assessment and dynamic recovery point updating when facing various risks brought by configuration changes, greatly improving the efficiency and reliability of server operation and maintenance, reducing the risk of business interruption caused by configuration errors, and providing a more stable and secure operation environment for the key business of the enterprise.

[0053] In order to ensure the efficiency and accuracy of the BMC configuration management method, the following aspects can also be optimized and improved.

[0054] First of all, the optimization of the health assessment algorithm. In actual application, different server environments and business requirements may have different sensitivities and importance of each running index of BMC. Therefore, according to the specific use scene, the weight of each index in the health assessment algorithm needs to be dynamically adjusted. For example, in a data center with extremely high requirements for network security, the weight of the effectiveness index of the firewall black and white list security policy may be set higher; while in a computing cluster focusing on server performance optimization, the weight of the index of the impact of firmware parameter configuration on component performance will be relatively larger. Through continuous optimization and adjustment of the algorithm, the evaluation result of the BMC running health degree is more consistent with the actual running condition, so as to improve the timeliness and accuracy of the configuration recovery point update.

[0055] Secondly, the selection and management of storage media. Since the data of BMC configuration restore points need to be saved completely after the server power off, non-volatile storage media are usually selected for storage, such as BMC local flash, embedded eMMC or server SSD, etc. Different storage media have advantages and disadvantages in terms of storage capacity, read-write speed, data reliability, etc. For example, BMC local flash and embedded eMMC have the advantages of fast read-write speed and low latency, but the storage capacity is relatively small; while the server's SSD has a larger storage capacity and can save more historical configuration restore points, but the read-write speed is relatively slow, and frequent read-write operations may have some impact on the life of the SSD. Therefore, in actual application, the storage medium needs to be selected reasonably according to the configuration of the server and the business requirements, and the corresponding storage strategy needs to be formulated, such as automatically cleaning up expired restore points, dynamically allocating storage space, etc., to ensure that there is enough storage space for important configuration restore point data.

[0056] In addition, in order to improve the security of the system and the integrity of the data, encryption algorithms can be used to encrypt the data when storing and updating the configuration restore points, to prevent the configuration information from being tampered with or leaked maliciously. At the same time, before updating the restore point each time, the integrity of the current BMC configuration can be checked and virus scanned to ensure that the saved restore point data is safe and reliable.

[0057] In order to realize more comprehensive server configuration management, the running health degree evaluation results of BMC can be combined with the hardware state information of the server (such as CPU temperature, memory usage, hard disk health status, etc.) for comprehensive analysis, so as to understand the overall running status of the server more comprehensively. When the configuration change causes the running health degree of BMC to decrease, the hardware state can also be checked for abnormalities at the same time, so as to accurately judge whether the root cause of the problem is a pure configuration problem or a result of the interaction between configuration and hardware. For example, if the configuration change of BMC causes the fan speed control of the server to be abnormal, which in turn causes the CPU temperature to be too high, through the cooperation with the hardware monitoring system, this chain reaction can be discovered in time, and appropriate measures can be taken, such as restoring the correct configuration restore point of BMC first, and then checking and maintaining the hardware cooling system, so as to effectively avoid the hardware failure caused by a single configuration problem and improve the overall reliability of the server.

[0058] The BMC can also be integrated with the software management system of the server. When the application program or operating system on the server is updated, upgraded, or the like, the BMC can perceive these changes in advance and, according to the dependence of the application program on the BMC configuration, automatically adjust the related index weight in the running health evaluation algorithm of the BMC or generate a corresponding configuration restore point in advance. For example, after the key business software running on the server is upgraded, there can be higher requirements for the network bandwidth response and speed of the BMC. At this time, the BMC configuration management method can cooperate with the software management system to increase the weight of the network-related index in the health evaluation algorithm, and at the same time when the software upgrade is completed, a correct configuration restore point is generated according to the current new BMC configuration, so that when the BMC configuration compatibility problem caused by software upgrade occurs subsequently, the correct configuration state suitable for the new software can be quickly restored to ensure the normal operation of the business software.

[0059] In addition, in a multi-server cluster environment, the BMC configurations of all servers in the entire cluster can be uniformly managed and monitored. By integrating the BMC configuration management function in the cluster management software, the administrator can conveniently monitor the BMC running health of each server in the cluster in real time, create, update, and roll back the configuration restore point in batches, and can perform consistency checking and synchronous management on the BMC configurations of the servers in the cluster. For example, in a large data center with hundreds or thousands of servers, through the integration of the cluster management system, the administrator can generate the BMC configuration restore point of all servers with one key, or quickly locate the related error configuration restore point when a BMC configuration problem is found in a server, and restore the BMC configuration of the server to the correct state, and at the same time, the correct configuration can be promoted to other servers that may have similar risks, thereby greatly improving the operation and management efficiency and quality of the data center.

[0060] In terms of security protection, the security mechanism of the BMC configuration management method is further strengthened to prevent malicious attackers from damaging the server by using the BMC configuration management function. For example, multi-factor authentication technology is used to strictly verify the identity of users accessing the BMC configuration management interface and performing restore point operations, to ensure that only authorized administrators can perform related configuration management operations. At the same time, higher-level encryption protection is provided for the storage and transmission process of the BMC configuration restore point to prevent the configuration data from being stolen or tampered with. In addition, an access control policy can be set to limit the viewing and operation permissions of different users on the BMC configuration restore point, for example, ordinary operation and maintenance personnel can only view and roll back to specific restore points, while senior administrators have complete management permissions on all restore points, thereby forming a multi-level security protection system to ensure the security of the server BMC configuration management.

[0061] In an embodiment, as shown in FIG. 1, the BMC configuration management method comprises the following steps:Figure 2 This specification also provides a BMC configuration management device, which is applied to a server. The device includes: a first module, which is used to evaluate the current BMC operation health according to a preset algorithm in response to an operation status change event; a second module, which is used to update the last correct configuration restore point according to the currently used BMC configuration in response to an event that the current BMC operation health is greater than a preset health threshold; a third module, which is used to update the last incorrect configuration restore point according to the currently used BMC configuration in response to an event that the current BMC operation health is less than a preset error threshold, and call and restore the BMC configuration according to the last correct configuration restore point.

[0062] In one embodiment, Figure 3 The device also includes: a fourth module, which is used to call and compare the most recent correct configuration restore point and the most recent incorrect configuration restore point in response to the diagnostic signaling, and obtain and display the difference between the most recent correct configuration restore point and the most recent incorrect configuration restore point.

[0063] In one embodiment, responding to an operation status change event and evaluating the current BMC operation health according to a preset algorithm includes: responding to a configuration change event and / or an abnormal load change event and / or a hardware change event and evaluating the current BMC operation health according to a preset weighted algorithm, where input parameters of the preset weighted algorithm are associated with the current configuration and / or current load and / or current hardware.

[0064] In one embodiment, in response to an event that the current BMC operation health is greater than a preset health threshold, updating the last correct configuration restore point according to the currently used BMC configuration includes: in response to an event that the current BMC operation health is greater than the preset health threshold, naming the existing last correct configuration restore point according to a preset rule and storing it, and updating the last correct configuration restore point according to the currently used BMC configuration.

[0065] In one implementation, this relies on a continuously running monitoring process (often referred to as the HealthMonitor process) implemented within the server's BMC firmware. This process is capable of collecting a series of key performance indicators (KPIs) and status signals in real time or near real time (for example, at a preset interval, such as every second, every 5 seconds, or every minute, with the specific interval being configurable based on server load and accuracy requirements). These monitored indicators are the basis for determining the overall health of the BMC and the server, and they cover a wide range of key areas that affect system stability.

[0066] Typical monitoring indicators include but are not limited to: BMC network connectivity indicators (such as the link status of out-of-band management port, Ping success rate with upstream switch or management station, ARP table status, TCP connection establishment success rate, etc.), BMC core service process status (such as the running status and resource occupancy of Web service, Redfish / RESTful API service, IPMI service, SNMP agent, KVM over IP service, virtual media service), BMC hardware sensor readings (such as whether the temperature, voltage, fan speed of BMC itself and associated CPU / memory / PCH and other key components are within the safe range), user session and management activity status (such as the number of currently active administrator sessions, authentication failure frequency, frequency and type of configuration modification operations), BMC log information (such as the number and type of serious errors, warning events recorded in the system event log SEL, especially entries related to configuration changes or service failures), communication status with host system and key firmware (such as whether the communication heartbeat with host BIOS is normal, whether the BIOS configuration can be successfully read, communication status with RAID controller management module), BMC itself resource utilization (such as CPU occupancy, memory usage, file system space status), etc.

[0067] These indicators constitute the raw data pool for evaluating the running health of BMC. In order to convert these heterogeneous, dimensionless raw data into a unified, quantifiable health score, a health evaluation algorithm is preset, which adopts a weighted scoring model.

[0068] In implementation, a weight coefficient (Weight) is defined for each monitoring indicator, which reflects the importance of the indicator to the overall health score. For example, the weight of network connectivity can be very high (e.g. 0.4), as it is the basis of out-of-band management; the weight of core service process downtime can also be high (e.g. 0.3); and the weight of non-critical warnings in logs can be low (e.g. 0.05). Each indicator is mapped to a single score (Score) between 0 and 100 (or 0 and 1.0) according to its current state or value. The calculation rule of the single score can be diversified, for example: a Boolean state (such as network on / off) can be represented by 100 points (on) and 0 points (off); a range type indicator (such as CPU occupancy) can be designed as a linear or nonlinear deduction function (e.g. occupancy < 70% gets 100 points, 70%-80% gets 80 points, 80%-90% gets 60 points, > 90% gets 0 points); a counting type indicator (such as the number of error logs) can be set with thresholds for deduction (e.g. 0 errors gets 100 points, 1-2 errors gets 80 points, 3-5 errors gets 50 points, > 5 errors gets 0 points). Finally, the current BMC running health score (HealthScore) is calculated by the weighted sum formula: HealthScore = Σ (Weight_i * Score_i), where i iterates through all monitored indicators. This calculated HealthScore is a dynamic value that reflects the comprehensive health status of the BMC and related systems in real time. System administrators can pre-configure two key thresholds according to business needs and risk tolerance: a preset health threshold (HealthyThreshold) and a preset fault threshold (FaultThreshold). Usually, HealthyThreshold is set to a higher value (e.g. 85 points), indicating that the system is in good condition; FaultThreshold is set to a lower value (e.g. 50 points), indicating that the system has entered an unstable or faulty state. These two thresholds are the thresholds for triggering the creation and recovery of restore points.

[0069] The running of the whole method is driven by change events of running state. The "change event" here is a broad concept, which includes both the periodic calculation and discovery of HealthScore changes (regardless of the size of the change) by the monitoring process, and some specific discrete events that can significantly change the system state are triggered. The latter, for example: the administrator explicitly modifies any configuration item of the BMC (such as IP address, user password, SNMP Trap target, KVM encryption settings, power policy, etc.) through the Web interface, Redfish API or IPMI command; BMC detects that itself or associated hardware (such as fan failure, temperature overrun) triggers a serious alarm (SEL records Critical events); the key service process is terminated or restarted unexpectedly; communication with the host BIOS or RAID card is interrupted, etc. When such change events occur (whether the HealthScore changes are discovered by periodic monitoring or specific events are triggered), the monitoring process will immediately (or in the next monitoring period) start the health evaluation process, that is, calculate the current HealthScore according to the preset algorithm.

[0070] Next, the system will perform different operations according to the comparison result of the calculated HealthScore and the preset threshold.

[0071] In response to the event that the current BMC running health score is greater than the preset health threshold (i.e. HealthScore > HealthyThreshold). This situation indicates that the server, after experiencing possible configuration changes or events, its state is still assessed as healthy and stable, which is an ideal time to create or update the "golden standard" restore point. Therefore, the system will perform the operation of "updating the last correct configuration restore point according to the current BMC configuration used". The detailed steps of this operation are as follows: First, the BMC generates a snapshot of its complete configuration data currently in use. This snapshot contains all key and manageable configuration items of the BMC, such as: network configuration (IP address, subnet mask, gateway, VLAN ID, DNS server), user account information (username, permission, password hash or encrypted credential), security settings (SSL / TLS certificate, password complexity policy, access control list ACL, IP filtering rule), service configuration (enabled service port, session timeout, KVM video quality and compression settings), hardware monitoring threshold (alarm threshold of temperature, voltage, fan speed), firmware update settings (automatic update policy), power and restart policy, log settings (SEL storage policy, remote Syslog server), and configuration bridge information related to the host, etc. This configuration snapshot data needs to be serialized into a structured, easy-to-store and compare format, such as JSON, XML or custom binary format. Then, this serialized configuration data, together with the precise timestamp (e.g. UTC time, accurate to milliseconds) and optional health score, trigger reason (such as "periodic monitoring" or "after configuration modification") and other metadata, is written into a specific file or storage block identified by the system as LastKnownGoodConfig. Whether there is a LastKnownGoodConfig before or not, this operation will completely overwrite the old storage content with the current healthy configuration. This means that under this configuration, LastKnownGoodConfig always only saves one piece of data - the configuration snapshot at the time when the health threshold condition is met. It is not a history list, but a dynamically updated single-point restore benchmark that always represents the "known reliable latest good state". This configuration restore point needs to be stored in non-volatile memory to ensure that the data is not lost after the server power is off. The specific storage location can be a dedicated partition of the BMC on-board Flash chip, a reserved space of the embedded eMMC storage chip, or if the BMC has access, it can also be stored in a specific secure directory on the server host operating system's SSD / NVMe hard disk after encryption. When choosing to store in SSD, it needs to ensure that the BMC can independently access the restore point without being affected by the host OS or file system failure.

[0072] In response to an event that the current BMC running health score is less than the pre-set error threshold (i.e. HealthScore < FaultThreshold). This situation indicates that the system detects a serious problem, the BMC or the relevant part managed by it is in an unhealthy or unstable state, usually accompanied by functional failure (such as out-of-band management network interruption, web service inaccessible, frequent key alarms). At this time, update the last fault configuration restore point according to the current BMC configuration. This operation is similar to updating the correct restore point: first, capture the current complete configuration of the BMC (which may contain the error configuration that triggered the problem at this time), serialize and attach metadata such as timestamp, health score at the time of triggering, error event details (such as associated SEL event ID), etc. Then, write this data to another specific file or storage block identified as LastKnownFaultConfig. Similarly, "update" means to overwrite the old content, so LastKnownFaultConfig also only saves the configuration snapshot at the time when the last error threshold is triggered. The storage location requirements are the same as the non-volatile requirements of LastKnownGoodConfig. The core goal of creating this restore point is to accurately record the system configuration state at the time of problem occurrence, providing key evidence for subsequent problem diagnosis. Call and restore the configuration of the BMC according to the last correct configuration restore point. At the same time or after generating the error restore point (usually immediately to minimize the impact of failure), the system will automatically start the configuration recovery process. This operation first reads the latest data stored in the LastKnownGoodConfig location. Then, the BMC parses this serialized configuration data and compares it with the current running problematic configuration. The system designs a safe and atomic configuration restoration engine. This engine performs the restoration operation, which may include: first, according to the restore point data, restore network settings, user accounts, security policies, service parameters, etc. item by item. This process needs to handle dependencies (such as restoring the network first to connect external verification sources) and potential conflicts. The restoration engine needs to have rollback capability, if an error is encountered during the restoration process (for example, a configuration item in the restore point is illegal or fails to apply in the current environment), the engine should be able to abort the restoration operation and try to restore to the state before the restoration started (or at least record detailed error logs), avoid causing more serious system deadlock.

[0073] After reboot, BMC loads the "correct configuration" from LastKnownGoodConfig. The system (or monitoring process) performs health assessment again after BMC reboots and finishes initialization. The expected result is that HealthScore goes back up to above FaultThreshold, and possibly close to or even above HealthyThreshold, indicating that the system has recovered from the error state and business continuity is guaranteed. The core goal of this operation is to automatically and quickly roll back the system to a known good configuration state upon detecting a serious fault, minimizing the mean time to recovery (MTTR) without human intervention.

[0074] In one scenario, a misconfigured network causes out-of-band management outage. An administrator enters an incorrect subnet mask 255.255.0.0 instead of 255.255.255.0 when modifying the BMC's IP address in the web interface. After saving, the BMC cannot communicate with the gateway due to the mismatched IP address and subnet, and the out-of-band management network is interrupted. The web interface is inaccessible, and the SSH connection is disconnected.

[0075] The health assessment is triggered by the configuration modification event. The network connectivity indicator (Ping gateway failure) scores a sharp drop, causing HealthScore to drop sharply below FaultThreshold (e.g., 30 points). The system immediately creates LastKnownFaultConfig, recording the current configuration containing the incorrect subnet mask. The system immediately reads LastKnownGoodConfig (whose HealthScore was 95 points when saved, containing the correct network configuration). The restoration engine applies the correct configuration (especially the subnet mask). BMC automatically reboots. After reboot, BMC loads the correct network configuration, and the out-of-band management network is restored. After the administrator logs in again, the timestamps of LastKnownFaultConfig and LastKnownGoodConfig are seen in the restore point management interface. Using the Diff function, it is immediately discovered that the subnet mask is the only changed and incorrect configuration item.

[0076] The management outage time is shortened from hours required for possible manual recovery (on-site operation or serial port connection) to a few minutes (BMC reboot time). The problem configuration is accurately locked.

[0077] In one embodiment, the present specification provides an electronic device comprising a processor and a readable storage medium, the readable storage medium storing machine executable instructions capable of being executed by the processor, the processor executing the machine executable instructions to implement the aforementioned BMC configuration management method. From the hardware layer, the hardware architecture schematic diagram can be referred to Figure 4 as shown.

[0078] In one embodiment, the present specification provides a readable storage medium, which stores machine executable instructions, when the machine executable instructions are invoked and executed by a processor, the machine executable instructions cause the processor to implement the aforementioned BMC configuration management method.

[0079] Here, the readable storage medium can be any electronic, magnetic, optical, or other physical storage apparatus, which can contain or store information such as executable instructions, data, etc. For example, the readable storage medium can be a RAM (Random Access Memory), a volatile memory, a non-volatile memory, a flash memory, a storage drive (such as a hard drive), a solid state drive, any type of storage disk (such as an optical disk, a DVD, etc.), or similar storage medium, or a combination thereof.

[0080] The system, apparatus, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0081] For the convenience of description, the above apparatus is described in various units by function respectively. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware in the implementation of the present specification.

[0082] Those skilled in the art will understand that the embodiments of the present specification can be provided as a method, a system, or a computer program product. Therefore, the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0083] The embodiments of the present specification are described with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems) and computer program products according to the embodiments of the present specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The flowcharts and / or block diagrams can include one or more flows and / or blocks that can represent a computer system or a specific part of it and / or a specific combination of these. Figure 1 The flowcharts and / or block diagrams can include one or more flows and / or blocks that can represent a computer system or a specific part of it and / or a specific combination of these.

[0084] Furthermore, these computer program instructions can also be stored in a computer readable memory that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The flowcharts and / or block diagrams can include one or more flows and / or blocks that can represent a computer system or a specific part of it and / or a specific combination of these. Figure 1 The flowcharts and / or block diagrams can include one or more flows and / or blocks that can represent a computer system or a specific part of it and / or a specific combination of these.

[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The flowcharts and / or block diagrams can include one or more flows and / or blocks that can represent a computer system or a specific part of it and / or a specific combination of these. Figure 1 The flowcharts and / or block diagrams can include one or more flows and / or blocks that can represent a computer system or a specific part of it and / or a specific combination of these.

[0086] Those skilled in the art should understand that the embodiments of the present specification can be provided as a method, a system or a computer program product. Therefore, the present specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present specification can take the form of a computer program product implemented on one or more computer usable storage media (which can include, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0087] The above only describes the embodiments of the present specification and is not intended to limit the present specification. Those skilled in the art can make various changes and modifications to the present specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present specification shall be included in the scope of claims of the present specification.

Claims

1. A BMC configuration management method, characterized in that: Applied to a server, the method includes: In response to the change event of the operating status, the current BMC operating health is evaluated according to the preset algorithm; In response to an event that the current BMC operation health is greater than a preset health threshold, updating the most recently correctly configured restore point based on the currently used BMC configuration; In response to an event that the current BMC operation health is less than a preset error threshold, the most recent incorrect configuration restore point is updated according to the currently used BMC configuration, and the BMC configuration is called and restored according to the most recent correct configuration restore point.

2. The method according to claim 1, characterized in that The method further comprises: In response to the diagnostic signaling, the most recent correct configuration restoration point and the most recent incorrect configuration restoration point are called and compared, and the difference between the most recent correct configuration restoration point and the most recent incorrect configuration restoration point is obtained and displayed.

3. The method according to claim 1, characterized in that The step of evaluating the current BMC operational health according to a preset algorithm in response to an operational status change event includes: In response to a configuration change event and / or an abnormal load change event and / or a hardware change event, the current BMC operational health is evaluated according to a preset weighted algorithm, where input parameters of the preset weighted algorithm are associated with the current configuration and / or the current load and / or the current hardware.

4. The method according to claim 1, wherein In response to the event that the current BMC operation health is greater than a preset health threshold, updating the most recently correctly configured restore point according to the currently used BMC configuration includes: In response to an event that the current BMC operation health is greater than a preset health threshold, the existing most recent correct configuration restore point is named and stored according to a preset rule, and the most recent correct configuration restore point is updated according to the currently used BMC configuration.

5. A BMC configuration management device, characterized in that: Applied to a server, the device includes: The first module is used to respond to the change event of the operating status and evaluate the current BMC operating health according to the preset algorithm; The second module is configured to update the most recently correctly configured restore point according to the currently used BMC configuration in response to an event that the current BMC operation health is greater than a preset health threshold; The third module is used to respond to the event that the current BMC operation health is less than a preset error threshold, update the last incorrect configuration restore point according to the currently used BMC configuration, and call and restore the BMC configuration according to the last correct configuration restore point.

6. The device according to claim 5, characterized in that The device further comprises: The fourth module is used to call and compare the most recent correct configuration restoration point and the most recent incorrect configuration restoration point in response to the diagnostic signaling, and obtain and display the difference between the most recent correct configuration restoration point and the most recent incorrect configuration restoration point.

7. The device according to claim 5, characterized in that The step of evaluating the current BMC operational health according to a preset algorithm in response to an operational status change event includes: In response to a configuration change event and / or an abnormal load change event and / or a hardware change event, the current BMC operational health is evaluated according to a preset weighted algorithm, where input parameters of the preset weighted algorithm are associated with the current configuration and / or the current load and / or the current hardware.

8. The device according to claim 5, characterized in that In response to the event that the current BMC operation health is greater than a preset health threshold, updating the most recently correctly configured restore point according to the currently used BMC configuration includes: In response to an event that the current BMC operation health is greater than a preset health threshold, the existing most recent correct configuration restore point is named and stored according to a preset rule, and the most recent correct configuration restore point is updated according to the currently used BMC configuration.

9. An electronic device, characterized in that: include: A processor and a readable storage medium, wherein the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1 to 4.

10. A readable storage medium, characterized in that: The readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the method according to any one of claims 1 to 4.