Parameter adjustment method, device, electronic device and storage medium
The RAS tasks are independently handled by the substrate management controller, which solves the problem of large-scale use of CPU resources, improves CPU performance and simplifies the troubleshooting process.
Patent Information
- Application Number
- CN202510872424.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Because CPU resources are occupied by a large number of RAS tasks, the CPU performance is degraded.
The RAS task is performed independently of the CPU through the Board Management Controller (BMC), including fault detection, result sending, receiving modification instructions and adjusting performance impact parameters, and replacing the CPU with RAS tasks.
Free up CPU resources, improve CPU performance, simplify troubleshooting processes, and reduce maintenance costs.
Smart Images

Figure CN120386662B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a parameter adjustment method, device, electronic device, and storage medium. Background Art
[0002] In current server management, the CPU (Central Processing Unit) typically handles RAS (Reliability, Availability, and Serviceability) tasks, such as hardware status monitoring and parameter adjustment. However, with growing data processing demands, especially in data-intensive applications, CPU resources are being overwhelmed by RAS tasks. Frequent RAS tasks can create a CPU resource bottleneck, leading to performance degradation.
[0003] Therefore, in the related technologies, the CPU performance degradation caused by the CPU resources being occupied by a large number of RAS tasks has become a problem that needs to be solved urgently. Summary of the Invention
[0004] The present application provides a parameter adjustment method, device, electronic device and storage medium to at least solve the problem in the related art that CPU performance is reduced due to CPU resources being occupied by a large number of RAS tasks.
[0005] The present application provides a parameter adjustment method, which is applied to a baseboard management controller set on a server, and the method includes: determining a fault detection result of the server based on a performance-affecting parameter of the server; sending the fault detection result to an external device of the server; receiving a modification instruction fed back by the external device; wherein the modification instruction is an instruction generated by the external device when it is determined that a first performance-affecting parameter indicated by the fault detection result meets a parameter modification condition; and adjusting the first performance-affecting parameter according to the modification instruction.
[0006] The present application also provides a parameter adjustment device, including: a determination module, used to determine the fault detection result of the server based on the performance impact parameter of the server; a sending module, used to send the fault detection result to an external device of the server; a receiving module, used to receive a modification instruction fed back by the external device; wherein the modification instruction is an instruction generated by the external device when it is determined that the first performance impact parameter indicated by the fault detection result meets the parameter modification condition; and an adjustment module, used to adjust the first performance impact parameter according to the modification instruction.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned parameter adjustment methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned parameter adjustment methods are implemented.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned parameter adjustment methods when executed by a processor.
[0010] Through this application, the fault detection result of the server can be determined based on the performance impact parameters of the server, that is, the baseboard management controller set on the server can perform the following RAS tasks independently of the CPU: obtain the fault detection result, and send the fault detection result to the external device of the server, and receive the modification instruction fed back by the external device. The modification instruction is generated by the external device when it determines that the target performance impact parameter indicated by the fault detection result meets the parameter modification condition. Then, the baseboard management controller can take over the CPU to adjust the first performance impact parameter, and replace the CPU to process the RAS task (referring to adjusting the first performance impact parameter), thereby releasing CPU resources and improving CPU performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 A structural block diagram of a server for a parameter adjustment method provided in an embodiment of the present application;
[0013] Figure 2 A flow chart of a parameter adjustment method provided in an embodiment of the present application;
[0014] Figure 3 One of the flowcharts of a method for determining a modification instruction provided in an embodiment of the present application;
[0015] Figure 4 This is a second flowchart of a method for determining a modification instruction provided in an embodiment of the present application;
[0016] Figure 5 Flowchart 3 of a method for determining a modification instruction provided in an embodiment of the present application;
[0017] Figure 6 This is a structural block diagram of a parameter adjustment device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0019] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0020] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0021] The parameter adjustment method embodiment provided in the embodiment of the present application can be executed in a server or similar computing device. Specifically, the parameter adjustment method embodiment provided in the embodiment of the present application can be run on a baseboard management controller on the server. The baseboard management controller can be a hardware manager set in the server, which can mainly monitor the hardware status of the device, perform remote management operations (such as parameter adjustment), and provide monitoring and control functions for the server's external devices. Figure 1 As shown, Figure 1 This is a hardware structure diagram of a server of a parameter adjustment method according to an embodiment of the present application. Figure 1 As shown, the server may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. The server may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server. Figure 1 More or fewer components than shown, or with Figure 1Different configurations shown.
[0022] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the parameter adjustment method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the server via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0023] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the server's communication provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0024] The embodiment of the present application provides a parameter adjustment method, which is applied to the baseboard management controller on the above server, such as Figure 2 As shown, the method includes the following steps S202-208:
[0025] S202, determining a fault detection result of the server based on performance influencing parameters of the server;
[0026] It's understood that performance-impacting parameters are indicators or data that reflect the health and operational efficiency of a server. These parameters can cover both hardware and software aspects, including but not limited to CPU utilization, memory usage, disk I / O rates, network latency, system error counts, CPU temperature, memory error rates, fan speeds, and power supply voltages. Fault detection results refer to the results of an analysis and evaluation based on performance-impacting parameters, determining whether a server has a fault or performance bottleneck.
[0027] Specifically, the baseboard management controller (BMC) can regularly collect performance-impacting parameters for servers using built-in hardware sensors or software monitoring tools. The BMC can analyze these parameters and identify abnormalities or data that exceeds preset thresholds. Examples include CPU utilization consistently exceeding 90%, memory usage reaching the system limit, increased disk I / O request latency, and decreased network throughput.
[0028] In an exemplary embodiment, the performance influencing parameters include at least hardware status parameters; the above-mentioned step S202 can be implemented by: determining the parameter value corresponding to the hardware status parameter; comparing the real-time status value of the hardware status parameter at the current moment with the parameter value corresponding to the hardware status parameter to obtain the current comparison result; and determining the fault detection result based on the current comparison result.
[0029] It should be noted that the parameter values corresponding to the hardware status parameters are pre-set parameter thresholds, including upper and lower limits, and are used to determine whether the server's hardware status is normal. The real-time status values of the hardware status parameters at the current moment are the collected values of the performance-impacting parameters of the server. By comparing the collected values with the parameter values, a comparison result is obtained, which can be used to instantly assess the health of the server. Specifically, when the real-time status value of the hardware status parameter exceeds the upper limit of the parameter value or falls below the lower limit of the parameter value, it may indicate that the server is experiencing performance issues or hardware failure.
[0030] It is understood that parameter values can be automatically set by the administrator or BMC based on the server model, environmental conditions, business requirements, management manuals, and historical operating data. For example, the CPU temperature threshold can be set at 25°C, and the memory error rate threshold can be set at no more than 1 in every 10 million accesses.
[0031] Specifically, the BMC continuously collects the server's hardware status parameters, such as real-time status values such as CPU temperature, memory error rate, and power supply voltage, through integrated sensors or hardware interfaces. The real-time status values of the currently collected hardware status parameters are compared with the parameter values corresponding to the hardware status parameters to generate the current comparison result. For example, if the current CPU temperature is 45°C, which exceeds the set 35°C, the current comparison result is abnormal. Based on the current comparison result, the BMC analyzes whether the hardware status parameters are beyond the normal operating range. If the comparison result shows that there are abnormal hardware status parameters, further analysis is performed on possible causes. For example, excessive CPU temperature may be due to a radiator failure or excessive computing load. If multiple hardware status parameters exceed the threshold at the same time, the BMC can initiate more complex fault detection processes, such as checking system logs, performing hardware diagnostic tests, etc., to determine the root cause of the problem.
[0032] In the above embodiment, the BMC can work independently of the server's operating system. By comparing the real-time status value with the preset parameter value, it can quickly and accurately detect the server's fault status, provide key information for server maintenance and optimization, and thus ensure the high reliability, availability and serviceability of the server.
[0033] In an exemplary embodiment, based on the current comparison result, a fault detection result is determined, including: when the current comparison result is that the real-time state value is greater than the parameter value corresponding to the hardware state parameter, obtaining the recorded historical comparison result of the server; wherein the historical comparison result is determined by comparing the previous state value of the hardware state parameter at the previous moment before the current moment with the parameter value corresponding to the hardware state parameter; when the historical comparison result is that the previous state value is greater than or equal to the parameter value corresponding to the hardware state parameter, it is determined that there is a fault in the server, and a fault detection result is generated.
[0034] It is understandable that the BMC determines whether the server has a persistent fault state based on the comparison result of the real-time status value and the parameter value corresponding to the hardware status parameter in combination with historical data, and generates a fault detection result accordingly.
[0035] When the current comparison result is abnormal, the stored historical comparison results are further retrieved. These historical comparison results record the comparison between the previous state value of the hardware status parameter and the parameter value at a certain moment before the current moment. If both the current comparison result and the historical comparison result show that the real-time state value and the previous state value of the hardware status parameter are greater than or equal to the parameter value, this is considered to be a persistent fault state of the server. After confirming the above situation, a fault detection result will be generated, indicating the specific fault type, such as "CPU temperature continues to be abnormally high", and may be accompanied by corresponding warning information or alarm signals. The fault detection result can also include the duration, severity, possible scope of impact of the fault, and recommended solution steps or emergency operation guidelines.
[0036] In some embodiments, the BMC can continuously record the status values of hardware status parameters and the comparison results of the status values with parameter thresholds. The comparison results can be stored in a local database or log file to facilitate subsequent analysis and fault tracing. The recorded comparison results can include timestamps to facilitate chronological tracking of server status changes and determine whether abnormal status is a rare event or a persistent problem.
[0037] In the above embodiment, by combining real-time monitoring and historical data analysis, it is possible to effectively identify and confirm persistent server failures, avoid false alarms caused by anomalies at a single point in time, and ensure the timely discovery and handling of real problems, which is of great significance for ensuring the stable operation of the server and the security of data.
[0038] S204, sending the fault detection result to an external device of the server;
[0039] The BMC selects an appropriate communication protocol to send fault detection results based on the type of external device and the communication environment. Common protocols include SNMP (Simple Network Management Protocol), SMTP (Simple Mail Transfer Protocol), and HTTPS (Hypertext Transfer Protocol Secure).
[0040] Specifically, the BMC can convert fault detection results into a format that external devices can understand. For example, it can encode, compress, or encrypt the fault detection results to determine the data packet to be sent. The data packet can contain key details such as the server ID (identification), fault type, occurrence time, fault severity, and recommended treatment steps.
[0041] S206, receiving a modification instruction fed back by the external device; wherein the modification instruction is an instruction generated by the external device when it determines that the first performance-affecting parameter indicated by the fault detection result meets the parameter modification condition;
[0042] It should be noted that the first performance-affecting parameter refers to the parameter directly related to the fault that is highlighted in the fault detection result. The parameter modification condition is a parameter condition set to determine whether the performance-affecting parameter needs to be adjusted.
[0043] When the BMC detects a fault, it generates detailed fault detection results, which include the specific performance-impacting parameters that caused the fault (such as CPU overload, memory leaks, and disk read / write errors). These performance-impacting parameters are analyzed to determine whether they meet the preset parameter modification conditions.
[0044] It is understandable that when determining whether the preset parameter modification conditions are met, the fault detection results can be analyzed to determine the nature of the fault (such as persistence, frequency), the severity of the fault, and the specific impact of the fault on the server performance. Specifically, the external device can analyze whether the parameter value of the performance-affecting parameter (i.e., the parameter threshold) is reasonable, that is, whether the parameter threshold is based on the actual working status of the server and the recommended value in the management manual. The external device can also determine whether the currently set parameter threshold is still applicable based on changes in the server operating environment (such as the temperature and humidity of the computer room) and business load. The external device can also analyze the frequency with which the status value exceeds the parameter threshold. If the status value of a performance-affecting parameter frequently reaches the parameter threshold and causes an alarm, but no serious fault has actually occurred, it may indicate that the parameter threshold setting is too sensitive.
[0045] Specifically, when the external device determines, based on the fault detection result, that among the performance-influencing parameters, a first performance-influencing parameter has a parameter value that meets the parameter modification condition, the external device feeds back a modification instruction to the BMC.
[0046] S208: Adjust the first performance impact parameter according to the modification instruction.
[0047] It is understood that after receiving a modification instruction from an external device, the BMC can actually adjust the first performance-impacting parameter that caused the fault or performance issue based on the instruction content. Specifically, the BMC can parse the received modification instruction to determine the first performance-impacting parameter to which the modification instruction refers and the parameter value contained in the modification instruction. The BMC can then modify the parameter value stored in the register or configuration file based on the parameter value contained in the modification instruction.
[0048] There are multiple implementation schemes for the above-mentioned step S208. In one optional embodiment: determining a second performance-affecting parameter from the modification instruction; wherein the second performance-affecting parameter is obtained by one of the following methods: accessing a parameter configuration unit of the baseboard management controller through the external device and modifying the parameter value of the first performance-affecting parameter in the parameter configuration unit; logging into a command line of the baseboard management controller through the external device and modifying the parameter value of the first performance-affecting parameter based on the command line instruction in the command line; logging into a server where the baseboard management controller is located through the external device and modifying the parameter value of the first performance-affecting parameter based on the parameter management unit in the server; or adjusting the first performance-affecting parameter based on the second performance-affecting parameter.
[0049] The BMC determines the second performance-affecting parameter from the modification instruction, where the second performance-affecting parameter is obtained by modifying the parameter value of the first performance-affecting parameter by an external device.
[0050] Specifically, an administrator or automated tool directly accesses the BMC's parameter configuration unit through a network connection, finds the parameter value of the first performance-affecting parameter, and modifies the parameter value to obtain the second performance-affecting parameter. For example, the first performance-affecting parameter and the second performance-affecting parameter may both refer to memory temperature, but the memory temperature threshold of the first performance-affecting parameter is different from the memory temperature threshold of the second performance-affecting parameter.
[0051] Administrators can also use command line access tools based on external devices to log in to the BMC command line interface and enter specific commands (such as set_parameter_value<param_id><new_value> ), modify the parameter value of the first performance-affecting parameter to obtain the second performance-affecting parameter. The administrator can also log in to the server operating system through an external device and use the server's built-in parameter management tool to modify the parameter value of the first performance-affecting parameter to obtain the second performance-affecting parameter.
[0052] In a specific application, after analyzing fault detection results, an external device determines that lowering the CPU temperature threshold can reduce frequent alarms. An administrator can use a server management tool to log in to the BMC web interface, locate the threshold corresponding to the CPU temperature, modify the threshold, and save the setting. After receiving the modification instruction, the BMC updates the temperature threshold corresponding to the CPU temperature.
[0053] In the above embodiment, the administrator can adjust the first performance-affecting parameter remotely and in batches without having to access the server operating system, thereby significantly improving management efficiency, simplifying the troubleshooting process, and reducing maintenance costs.
[0054] Through the above S202-S208, the fault detection result of the server can be determined according to the performance impact parameter of the server, that is, the baseboard management controller set on the server can perform the following RAS tasks independently of the CPU: obtain the fault detection result, and send the fault detection result to the external device of the server, and receive the modification instruction fed back by the external device. The modification instruction is generated by the external device when it determines that the target performance impact parameter indicated by the fault detection result meets the parameter modification condition. Then, the baseboard management controller can take over the CPU to modify the first performance impact parameter, and replace the CPU to process the RAS task (referring to adjusting the first performance impact parameter), thereby releasing CPU resources and improving CPU performance.
[0055] In an exemplary embodiment, adjusting a first performance impact parameter based on a second performance impact parameter includes: obtaining a pre-set register mapping table; wherein the register mapping table includes: multiple performance impact parameters, and register identifiers corresponding to the multiple performance impact parameters respectively; matching the first performance impact parameter with the register mapping table to determine a target register identifier corresponding to the first performance impact parameter; and writing the parameter value of the second performance impact parameter to a target register corresponding to the target register identifier to adjust the first performance impact parameter in the target register to the second performance impact parameter.
[0056] The register mapping table is a pre-compiled data table that lists all performance management-related registers in the server hardware (for example, registers used to control CPU frequency, voltage, or memory error thresholds), as well as the corresponding relationships between these registers and performance-impacting parameters. This mapping table enables the BMC to quickly locate the hardware control points associated with performance-impacting parameters, enabling direct and efficient hardware management.
[0057] Specifically, the BMC can match the first performance-affecting parameter with the register mapping table. The first performance-affecting parameter is usually a performance indicator that needs to be optimized or intervened, such as CPU temperature, memory error rate, etc., and determine the register directly related to the first performance-affecting parameter. Once a match is found, the BMC can determine the target register identifier, that is, the register in the hardware that specifically controls the first performance-affecting parameter. After determining the target register identifier, the BMC writes the new parameter value of the second performance-affecting parameter obtained from the modification instruction to the corresponding target register. Modifying the value of the target register can cause the hardware status to change immediately.
[0058] In some embodiments, when the temperature threshold set for memory temperature, a performance-impacting parameter, does not meet the recommended value in the management manual, the BMC may receive a modification instruction requesting a lowering of the temperature threshold. The BMC first locates the register identifier for controlling CPU temperature from the register mapping table and then writes the newly set temperature threshold to that register. This adjustment can alleviate the problem of frequent alarms caused by CPU temperature. The BMC also records this operation, including information such as the voltage values before and after the modification and the operation time, for subsequent analysis and troubleshooting.
[0059] In the above embodiment, the BMC can accurately adjust the first performance-affecting parameter by operating the hardware register based on the second performance-affecting parameter.
[0060] For another optional implementation of step S208, adjusting the first performance impact parameter based on the second performance impact parameter includes: calling the file system interface of the baseboard management controller; wherein the file system interface is an interface for managing a configuration file stored in a non-volatile memory of the baseboard management controller; based on the file system interface, reading the configuration file from the non-volatile memory; and writing the parameter value of the second performance impact parameter to the configuration file to adjust the first performance impact parameter in the configuration file to the second performance impact parameter.
[0061] The BMC integrates a file system interface. By calling this interface, configuration files stored in the BMC's non-volatile memory can be directly accessed and modified without going through the server's main operating system. Non-volatile memory typically refers to the BMC's FLASH memory, which stores various configuration information, including performance management parameter settings.
[0062] Specifically, the BMC reads the configuration file through the file system interface and, based on the modification instructions, writes the new parameter value of the second performance-impacting parameter into the previously read configuration file, overwriting the original parameter value. For example, if the second performance-impacting parameter is a voltage threshold, the field related to the voltage setting can be found in the configuration file and its value updated. Once the configuration file is updated, the BMC immediately applies the new parameter value to the hardware control logic, allowing the effect of the parameter adjustment to be immediately visible even without restarting the server during runtime.
[0063] In the above embodiment, by calling the BMC's file system interface to directly read and update parameter values in the configuration file, performance management parameters can be dynamically adjusted during server operation without restarting the server or affecting business continuity. This mechanism is a key embodiment of the flexibility and efficiency of modern server management technology and a key strategy for ensuring stable server operation in complex business scenarios.
[0064] In an exemplary embodiment, the method of writing the parameter value of the second performance affecting parameter to the configuration file to adjust the first performance affecting parameter in the configuration file to the second performance affecting parameter includes: identifying the file type of the configuration file; performing format conversion on the configuration file based on the file type, and converting the configuration file into a structured data object; wherein the structured data object includes: multiple keys, and key values corresponding to the multiple keys; traversing the multiple keys based on the parameter name of the first performance affecting parameter to determine the key to be modified; writing the parameter value of the second performance affecting parameter to the key value corresponding to the key to be modified, so as to adjust the first performance affecting parameter in the configuration file to the second performance affecting parameter.
[0065] It should be noted that configuration files can be stored in a variety of formats, the most common of which are text formats (such as INI, JSON, XML), binary or database formats. Identifying the file type is a prerequisite for parsing and modifying the configuration. The type of the configuration file is identified by the file extension, file header tag or predefined metadata field. Furthermore, the BMC performs format conversion based on the file type, converting the original configuration file into a structured data object that is easy for computer programs to operate, such as a dictionary, array or object, to facilitate the search and modification of specific parameters. For example, for JSON or XML, it can be directly parsed into a dictionary or object tree structure; the converted structured data object includes multiple keys (Key) and key values (Value), the key represents the parameter name in the configuration, and the key value represents the specific value or status of the parameter.
[0066] Specifically, based on the parameter name of the first performance-affecting parameter, the BMC traverses all keys in the converted structured data object to find a key that matches the first performance-affecting parameter. Once the key corresponding to the first performance-affecting parameter is found, it is used as the key to be modified and the parameter value is prepared to be replaced. The parameter value of the second performance-affecting parameter is written to the key value corresponding to the key to be modified to achieve the conversion from the second performance-affecting parameter to the first performance-affecting parameter. The modified structured data object is converted back to the original file format and written back to the BMC's non-volatile memory to update the configuration file.
[0067] In the above embodiment, by identifying the file type, format conversion, key value search and modification, and security verification of the configuration file, dynamic optimization of the first performance influencing parameter can be achieved during the operation of the server, and the system performance can be improved or potential hardware problems can be addressed by adjusting the second performance influencing parameter.
[0068] The embodiments described above are only part of the embodiments of the present application, not all of the embodiments. In order to better understand the above method, the above process is described below in conjunction with the embodiments, but it is not intended to limit the technical solutions of the embodiments of the present application. Specifically:
[0069] The BMC analyzes the collected performance-impacting parameters to identify abnormalities or data exceeding preset thresholds. Examples include CPU utilization consistently exceeding 90%, memory usage reaching the system limit, increased disk I / O request latency, and decreased network throughput. Furthermore, the BMC can convert fault detection results into a format understandable to external devices and send them to the device.
[0070] External devices can analyze fault detection results to determine the nature of the fault (such as persistence and frequency), the severity of the fault, and the specific impact of the fault on server performance. They can then determine whether the parameter values set for performance-affecting parameters (i.e., parameter thresholds) are reasonable. If the parameter values are unreasonable, modification instructions are generated and sent to the BMC, so that the BMC can modify the parameters based on the modification instructions.
[0071] Among them, reference Figure 3 FIG. 1 is a flow chart showing how an external device modifies a parameter threshold, including the following steps:
[0072] S301: The administrator uses an administrator terminal (i.e., an external device) with network connectivity to enter the BMC's IP address (Internet Protocol Address) through a browser and access the IPMI (Intelligent Platform Management Interface) web login page.
[0073] S302: After successful login, find the "RAS Settings" related menu on the IPMI interface, that is, the parameter configuration unit.
[0074] S303: After entering, you can see various RAS threshold settings such as memory error threshold, CPU temperature threshold, etc. The administrator can modify the threshold according to actual needs and click "Save" or "Apply" to generate the modification instructions;
[0075] S304: Sending a modification instruction to a baseboard management controller (BMC);
[0076] S305: The BMC receives the modification instruction and updates the system registers and its own configuration files. If the modified threshold value does not take effect, check the BMC log to see if there are any related error prompts. This may be caused by the threshold value exceeding the reasonable range or a BMC configuration conflict. At the same time, the IPMI interface can also display in real time the hardware status information collected and organized by the BMC from the hardware sensors, such as memory usage, hardware error counts, etc. Administrators can regularly view data such as memory usage and hardware error counts on the IPMI interface. If an abnormal hardware status is found, such as a sudden increase in the memory error count, the cause can be analyzed in a timely manner, and combined with the server business load, it can be determined whether further adjustment of the threshold value or hardware maintenance is required.
[0077] In addition, reference Figure 4 FIG. 1 is another flow chart showing how an external device modifies a parameter threshold, including the following steps:
[0078] Download the server management tool for the corresponding model from the server manufacturer's official website and complete the installation on the external device according to the installation wizard.
[0079] S401: Open the server management tool, enter the server IP, administrator account and password to establish a connection with the server. If the connection fails, check whether the server network configuration is correct and whether the firewall is blocking the connection request of the management tool. You can use the ping command to test the server network connectivity;
[0080] S402: Find the "Hardware Configuration" or "RAS Management" module, i.e., the parameter management unit, in which various RAS thresholds and hardware status can be graphically displayed;
[0081] S403: Parameter value modification. The administrator modifies the threshold value through mouse operation and clicks "Confirm" or "Submit" to generate the modification instruction. For example, when adjusting the temperature threshold, you can refer to the recommended temperature range in the server hardware manual and the current server heat dissipation conditions for setting. After the modification is completed, click the "Confirm" or "Submit" button. The server management tool can also obtain hardware status data from the BMC in real time and display it. If the data display is abnormal, such as the fan speed does not match the actual hardware status, you can check whether the communication between the server management tool and the BMC is normal, such as trying to reconnect or update the server management tool version;
[0082] S404: Sending a modification instruction to the BMC;
[0083] S405: The BMC receives the modification instruction and updates the system registers and its own configuration files. At the same time, the management tool can also obtain hardware status data from the BMC in real time and display it, such as charts of server power consumption, fan speed, etc.
[0084] In addition, reference Figure 5 FIG. 1 is another flow chart showing how an external device modifies a parameter threshold, including the following steps:
[0085] S501: Connect to the BMC command line using an external device. Use a serial cable to connect the management terminal to the server serial port. Open the serial communication software on the terminal and set the correct serial port parameters to establish a connection. Alternatively, use the SSH client tool to connect to the BMC IP address as an administrator over the network. Once connected, enter the BMC command line interface. If the serial port connection fails, check whether the serial cable is damaged and whether the serial port parameter settings are consistent with those on the server. If the network connection fails, check whether the SSH (Secure Shell Service) is running properly on the BMC, whether the IP address is correct, and whether there are any network security policy restrictions.
[0086] S502: Execute command operations, connect successfully and enter the BMC command line interface. Based on the OpenBMC system, use the command cd / xyz / openbmc_project / ras / to enter the RAS related namespace, that is, the parameter setting unit;
[0087] S503: Modify parameter values. Use specific commands to modify specific thresholds, such as the memory error threshold, and generate modification instructions. Then use query commands to confirm whether the modification has taken effect. You can also use related commands to view the current status of hardware such as memory and hard disks. For example, if the return data for the show_memory_status command is abnormal, check whether the server memory hardware is faulty. Further troubleshooting can be performed using hardware diagnostic tools. If an error is reported during the execution of any command, analyze the error message to determine whether it is caused by a command syntax error, insufficient permissions, or a server system failure.
[0088] S504: Sending a modification instruction to the BMC;
[0089] S505: The BMC receives the modification instruction and updates the system register and its own configuration file.
[0090] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0091] The embodiments of the present application also provide a parameter adjustment device for realizing the above-mentioned embodiments and preferred embodiments, which have been described and will not be repeated here. As used below, the term "module" can realize a combination of software and / or hardware of a predetermined function. Although the modules described in the following embodiments are preferably realized with software, the realization of hardware, or a combination of software and hardware is also possible and contemplated.
[0092] Figure 6 : is a structural block diagram of a parameter adjustment device according to an embodiment of the present application, the device comprising:
[0093] A determination module 602 is configured to determine a fault detection result of the server according to a performance influencing parameter of the server;
[0094] A sending module 604 is configured to send the fault detection result to an external device of the server;
[0095] A receiving module 606 is configured to receive a modification instruction fed back by the external device; wherein the modification instruction is an instruction generated by the external device when it is determined that the first performance-affecting parameter indicated by the fault detection result meets a parameter modification condition;
[0096] The adjustment module 608 is configured to adjust the first performance influencing parameter according to the modification instruction.
[0097] Through the above-mentioned device, the fault detection result of the server is determined according to the performance impact parameter of the server, that is, the baseboard management controller set on the server can perform the following RAS tasks independently of the CPU: obtain the fault detection result, and send the fault detection result to the external device of the server, and receive the modification instruction fed back by the external device. The modification instruction is generated by the external device when it is determined that the target performance impact parameter indicated by the fault detection result meets the parameter modification condition. Then, the baseboard management controller can take over the CPU to modify the first performance impact parameter, and replace the CPU to process the RAS task (referring to adjusting the first performance impact parameter), thereby releasing CPU resources and improving CPU performance.
[0098] In an exemplary embodiment, the adjustment module 608 is also used to determine a second performance impact parameter from the modification instruction; wherein, the second performance impact parameter is obtained in one of the following ways: accessing the parameter configuration unit of the baseboard management controller through the external device, and modifying the parameter value of the first performance impact parameter in the parameter configuration unit; logging into the command line of the baseboard management controller through the external device, and modifying the parameter value of the first performance impact parameter based on the command line instruction in the command line; logging into the server where the baseboard management controller is located through the external device, and modifying the parameter value of the first performance impact parameter based on the parameter management unit in the server; adjusting the first performance impact parameter based on the second performance impact parameter.
[0099] In an exemplary embodiment, the adjustment module 608 is further used to obtain a pre-set register mapping table; wherein the register mapping table includes: multiple performance impact parameters, and register identifiers corresponding to the multiple performance impact parameters respectively; matching the first performance impact parameter with the register mapping table to determine the target register identifier corresponding to the first performance impact parameter; writing the parameter value of the second performance impact parameter to the target register corresponding to the target register identifier to adjust the first performance impact parameter in the target register to the second performance impact parameter.
[0100] In an exemplary embodiment, the adjustment module 608 is also used to call the file system interface of the baseboard management controller; wherein the file system interface is an interface for managing a configuration file stored in a non-volatile memory of the baseboard management controller; based on the file system interface, the configuration file is read from the non-volatile memory; and the parameter value of the second performance impact parameter is written to the configuration file to adjust the first performance impact parameter in the configuration file to the second performance impact parameter.
[0101] In an exemplary embodiment, the adjustment module 608 is also used to identify the file type of the configuration file; perform format conversion on the configuration file based on the file type, and convert the configuration file into a structured data object; wherein the structured data object includes: multiple keys, and key values corresponding to the multiple keys; traverse the multiple keys based on the parameter name of the first performance affecting parameter to determine the key to be modified; write the parameter value of the second performance affecting parameter to the key value corresponding to the key to be modified, so as to adjust the first performance affecting parameter in the configuration file to the second performance affecting parameter.
[0102] In an exemplary embodiment, the performance influencing parameters include at least hardware status parameters; the determination module 602 is further used to determine the parameter value corresponding to the hardware status parameter; compare the real-time status value of the hardware status parameter at the current moment with the parameter value corresponding to the hardware status parameter to obtain a current comparison result; and determine the fault detection result based on the current comparison result.
[0103] In an exemplary embodiment, the determination module 602 is also used to obtain the recorded historical comparison result of the server when the current comparison result is that the real-time status value is greater than the parameter value corresponding to the hardware status parameter; wherein the historical comparison result is determined by comparing the previous status value of the hardware status parameter at the previous moment before the current moment with the parameter value corresponding to the hardware status parameter; when the historical comparison result is that the previous status value is greater than or equal to the parameter value corresponding to the hardware status parameter, it is determined that there is a fault in the server and the fault detection result is generated.
[0104] For the description of the features in the embodiment corresponding to the parameter adjustment device, reference can be made to the relevant description of the embodiment corresponding to the parameter adjustment method, and no further details will be given here.
[0105] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above parameter adjustment method embodiments.
[0106] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned parameter adjustment method embodiments when running.
[0107] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0108] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned parameter adjustment method embodiments are implemented.
[0109] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned parameter adjustment method embodiments are implemented.
[0110] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0111] The above is a detailed introduction to a parameter adjustment method provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A parameter adjustment method, characterized in that: Applied to a baseboard management controller provided on a server, the method includes: Determining a fault detection result of the server according to a performance influencing parameter of the server; sending the fault detection result to an external device of the server; receiving a modification instruction fed back by the external device; wherein the modification instruction is an instruction generated by the external device when it is determined that the first performance-affecting parameter indicated by the fault detection result meets a parameter modification condition; adjusting the first performance influencing parameter according to the modification instruction; The adjusting the first performance-affecting parameter according to the modification instruction includes: Determine a second performance-affecting parameter from the modification instruction; wherein the second performance-affecting parameter is obtained by one of the following ways: accessing a parameter configuration unit of the baseboard management controller through the external device and modifying a parameter value of the first performance-affecting parameter in the parameter configuration unit; logging into a command line of the baseboard management controller through the external device and modifying the parameter value of the first performance-affecting parameter based on a command line instruction in the command line; logging into a server where the baseboard management controller is located through the external device and modifying the parameter value of the first performance-affecting parameter based on a parameter management unit in the server; The first performance impact parameter is adjusted based on the second performance impact parameter.
2. The method according to claim 1, characterized in that The adjusting the first performance influencing parameter based on the second performance influencing parameter includes: Obtaining a preset register mapping table; wherein the register mapping table includes: a plurality of performance-affecting parameters, and register identifiers corresponding to the plurality of performance-affecting parameters; Matching the first performance impact parameter with the register mapping table to determine a target register identifier corresponding to the first performance impact parameter; The parameter value of the second performance affecting parameter is written into the target register corresponding to the target register identifier, so as to adjust the first performance affecting parameter in the target register to the second performance affecting parameter.
3. The method according to claim 1, characterized in that The adjusting the first performance influencing parameter based on the second performance influencing parameter includes: Calling a file system interface of the baseboard management controller; wherein the file system interface is an interface for managing configuration files stored in a non-volatile memory of the baseboard management controller; Reading the configuration file from the non-volatile memory based on the file system interface; The parameter value of the second performance influencing parameter is written into the configuration file, so as to adjust the first performance influencing parameter in the configuration file to the second performance influencing parameter.
4. The method according to claim 3, characterized in that Writing the parameter value of the second performance influencing parameter into the configuration file to adjust the first performance influencing parameter in the configuration file to the second performance influencing parameter includes: Identifying a file type of the configuration file; Performing format conversion on the configuration file based on the file type, converting the configuration file into a structured data object; wherein the structured data object includes: a plurality of keys, and key values corresponding to the plurality of keys; Traversing the plurality of keys based on the parameter name of the first performance influencing parameter to determine a key to be modified; The parameter value of the second performance impact parameter is written into the key value corresponding to the key to be modified, so as to adjust the first performance impact parameter in the configuration file to the second performance impact parameter.
5. The method according to claim 1, characterized in that The performance influencing parameters include at least hardware status parameters; Determining the fault detection result of the server according to the performance influencing parameter of the server includes: Determining a parameter value corresponding to the hardware status parameter; Compare the real-time state value of the hardware state parameter at the current moment with the parameter value corresponding to the hardware state parameter to obtain a current comparison result; Based on the current comparison result, a fault detection result is determined.
6. The method according to claim 5, characterized in that The determining of a fault detection result based on the current comparison result includes: When the current comparison result is that the real-time state value is greater than the parameter value corresponding to the hardware state parameter, obtaining a recorded historical comparison result of the server; wherein the historical comparison result is determined by comparing a previous state value of the hardware state parameter at a previous moment before the current moment with the parameter value corresponding to the hardware state parameter; If the historical comparison result is that the previous state value is greater than or equal to the parameter value corresponding to the hardware state parameter, it is determined that a fault exists in the server, and the fault detection result is generated.
7. A parameter adjustment device, characterized in that: Applicable to a baseboard management controller provided on a server, the device comprises: A determination module, configured to determine a fault detection result of the server according to a performance influencing parameter of the server; A sending module, configured to send the fault detection result to an external device of the server; a receiving module, configured to receive a modification instruction fed back by the external device; wherein the modification instruction is an instruction generated by the external device when it is determined that the first performance-affecting parameter indicated by the fault detection result meets a parameter modification condition; An adjustment module is used to determine a second performance impact parameter from the modification instruction; wherein the second performance impact parameter is obtained in one of the following ways: accessing the parameter configuration unit of the baseboard management controller through the external device and modifying the parameter value of the first performance impact parameter in the parameter configuration unit; logging into the command line of the baseboard management controller through the external device and modifying the parameter value of the first performance impact parameter based on the command line instruction in the command line; logging into the server where the baseboard management controller is located through the external device and modifying the parameter value of the first performance impact parameter based on the parameter management unit in the server; adjusting the first performance impact parameter based on the second performance impact parameter.
8. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Equipment fault detection method, electronic equipment and storage medium
CN118011127A
Performance tuning method and device
CN119690341A