Server management system, method, electronic device, and storage medium
By introducing a dedicated neural network into the BMC system, the problems of difficulty in identifying hardware failures caused by complex factors and difficulties in system expansion are solved, enabling efficient fault handling and hardware adjustment, and improving the intelligence and adaptability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-06-02
AI Technical Summary
Existing BMC systems struggle to identify hardware failures caused by complex factors and cannot adjust server hardware in a timely manner, resulting in significant manpower expenditure on fault analysis and difficulties in system expansion.
A dedicated neural network is introduced to create a management information generation module, enhancing the BMC function. The network control module receives weights, the data processing module acquires hardware information, the management information generation module generates server management information, and the output control module adjusts the hardware and generates alarm information.
It improves the accuracy of fault identification, reduces missed reports, saves manpower, improves fault handling efficiency, and enhances the system's scalability.
Smart Images

Figure CN122132205A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of server management technology, specifically to server management systems, methods, electronic devices, and storage media. Background Technology
[0002] The Baseboard Management Controller (BMC) can manage the server independently of the server's main processor. The BMC is responsible for monitoring the server's hardware status, such as temperature, voltage, and fan speed, and performing corresponding management and control.
[0003] When diagnosing server hardware faults, the BMC (Browser Control Center) can only identify the presence of hardware faults and issue alerts based on simple thresholds and limited rules. Subsequent analysis and location of the hardware faults require diagnostic tools to help administrators determine the cause and location of the fault, allowing them to decide on appropriate actions to repair or prevent it. However, relying solely on simple thresholds and limited rules makes it difficult for the BMC to identify hardware faults caused by complex, multi-factor interactions. Furthermore, the reliance on diagnostic tools and administrator analysis results in a significant manpower requirement for fault analysis, hindering timely adjustments to the server hardware and making it difficult to quickly resolve or prevent hardware faults. Summary of the Invention
[0004] This invention provides a server management system, method, electronic device, and storage medium to solve the problem of difficulty in identifying hardware failures caused by complex factors and the inability to adjust server hardware in a timely manner in response to hardware failures.
[0005] In a first aspect, this application provides a server management system, which includes: a network control module, a data processing module, a management information generation module, and an output control module;
[0006] The network control module is used to receive weights sent by the remote management platform and transmit the weights to the management information generation module; The data processing module is used to acquire server hardware information and transmit it to the management information generation module. The management information generation module is used to generate server management information based on weights and server hardware information, and then transmit the server management information to the output control module. The output control module is used to generate hardware control information based on server management information, adjust server hardware using hardware control information, and transmit server management information and hardware control information to the network control module. The network control module is also used to generate alarm information based on server management information and hardware control information.
[0007] Secondly, this application provides a server management method, the method comprising: Receive weights sent by the remote management platform and obtain server hardware information; Generate server management information based on weights and server hardware information; Hardware control information is generated based on server management information, and server hardware is adjusted using the hardware control information. Alarm information is generated based on server management information and hardware control information.
[0008] Thirdly, this application provides a server management device, which includes: The information acquisition module is used to receive weights sent by the remote management platform and acquire server hardware information; The management information generation module is used to generate server management information based on weights and server hardware information. The hardware control module is used to generate hardware control information based on server management information and to adjust the server hardware using the hardware control information. The alarm information generation module is used to generate alarm information based on server management information and hardware control information.
[0009] Fourthly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the server management method of the second aspect or any corresponding embodiment described above.
[0010] Fifthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to perform the server management method of the second aspect or any corresponding embodiment described above.
[0011] Sixthly, this application provides a computer program product, including computer instructions for causing a computer to execute the server management method described in the second aspect or any corresponding embodiment thereof.
[0012] This application utilizes a network control module to transmit received weights to a management information generation module; a data processing module to transmit acquired server hardware information to the management information generation module; the management information generation module generates server management information based on the weights and server hardware information; an output control module generates hardware adjustment information based on the server management information, adjusts the server hardware using the hardware adjustment information, and transmits the server management information and hardware adjustment information to the network control module; the network control module also generates alarm information based on the server management information and hardware adjustment information. This addresses the problem of difficulty in identifying hardware faults caused by complex factors and the inability to adjust server hardware in a timely manner for hardware faults. By introducing a dedicated neural network to create the management information generation module, this system enhances the BMC (Browser Control Center) function, enabling the identification of subtle fault characteristics and preventing missed fault reports; it eliminates the need for manual analysis of data or faults, improving fault handling efficiency and saving manpower. Furthermore, the dedicated neural network can be adjusted according to different needs, improving the system's scalability. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this application, the drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of the server management system according to an embodiment of this application; Figure 2 This is a schematic diagram of the application environment of the server management system according to an embodiment of this application; Figure 3 This is a schematic diagram of the inference phase workflow of the server management system according to an embodiment of this application; Figure 4 This is a schematic diagram of another server management system according to an embodiment of this application; Figure 5 This is a schematic diagram of the neural network training process according to an embodiment of this application; Figure 6 This is a flowchart illustrating a server management method according to an embodiment of this application; Figure 7 This is a structural block diagram of a server management device according to an embodiment of this application; Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0016] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0017] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] The Baseboard Management Controller (BMC) can operate independently of the server's main processor. The BMC monitors the server's hardware status, such as temperature, voltage, and fan speed, and manages and controls these parameters accordingly. However, in hardware fault diagnosis, the BMC often relies on simple thresholds and limited rules to identify faults. For hardware faults caused by multiple factors, the BMC struggles to accurately pinpoint the root cause, resulting in poor accuracy and timeliness of diagnosis. Regarding predictive maintenance, the BMC lacks in-depth analysis capabilities of hardware performance trends and cannot anticipate potential hardware failures, leading to unexpected server malfunctions and impacting business continuity. In energy management, the BMC struggles to dynamically and accurately adjust energy supply strategies based on real-time server workload and hardware status, resulting in energy waste. Furthermore, the BMC collects real-time equipment operating status information through hardware status monitoring technology. When abnormalities are detected, such as overheating, abnormal voltage, or hardware errors, it issues alarm signals to notify administrators. Administrators then use various fault diagnosis tools, such as hardware diagnostic programs and system log analysis tools, to analyze and locate the fault, determining its cause and location. However, some complex hardware faults do not show obvious signs in the early stages, or the fault characteristics are similar to normal operating fluctuations, making it difficult for the BMC to accurately identify them and easily leading to missed detections.
[0019] Insufficient BMC fault handling capabilities mean that after collecting the aforementioned hardware status data, the BMC needs to analyze the hardware status and determine whether alarms have been triggered. If the BMC's data processing capabilities are insufficient, or if a large amount of monitoring data needs to be processed simultaneously, it may lead to delays in hardware fault handling, affecting the real-time performance of alarms. For example, when monitoring large-scale server clusters, alarm delays may occur due to the need to process information from numerous nodes simultaneously. Furthermore, processing a large volume of alarm information requires specialized personnel for monitoring and analysis. A high false alarm rate will cause maintenance personnel to spend a significant amount of time on ineffective alarm handling, increasing labor costs.
[0020] Fault diagnosis tools typically require collecting and analyzing large amounts of hardware status data. This process can be time-consuming, especially for large-scale hardware systems or large datasets. For example, when diagnosing a server cluster in a large data center, collecting hardware status data from all servers may take several minutes or even longer, and analyzing this data also requires time. This leads to low efficiency in fault diagnosis and impacts system recovery time. As hardware systems grow in scale and functionality, fault diagnosis tools need to be highly scalable to adapt to new hardware devices and system architectures. However, in practice, expanding diagnostic tools can face technical challenges and cost issues. For instance, when a data center adds a batch of different types of servers, it may require extensive modifications and upgrades to existing fault diagnosis tools to support the new equipment. This not only requires significant R&D resources but may also affect the normal operation of the system.
[0021] Based on the above, this application provides a server management system, which is a management-enhanced BMC system based on neural network algorithms. By introducing a dedicated neural network, the functionality of the BMC chip is enhanced. This system can solve the following problems existing in server management technologies: missed reports due to unclear fault characteristics, real-time issues due to insufficient system processing capacity and time, the time-consuming data acquisition and analysis requiring a large amount of manpower, and difficulties in system expansion.
[0022] The specific application environment architecture or specific hardware architecture on which the execution of the server management method depends is described here.
[0023] According to an embodiment of this application, a server management system is provided. It should be noted that the server management system can be created based on integrated circuit components with executable instructions, such as BMC.
[0024] This embodiment provides a server management system. Figure 1This is a structural diagram of the server management system according to an embodiment of this application, such as... Figure 1 As shown, the server management system includes: a network control module, a data processing module, a management information generation module, and an output control module; The network control module is used to receive weights sent by the remote management platform and transmit the weights to the management information generation module; The data processing module is used to acquire server hardware information and transmit it to the management information generation module. The management information generation module is used to generate server management information based on weights and server hardware information, and then transmit the server management information to the output control module. The output control module is used to generate hardware control information based on server management information, adjust server hardware using hardware control information, and transmit server management information and hardware control information to the network control module. The network control module is also used to generate alarm information based on server management information and hardware control information.
[0025] Specifically, the server management system is built on BMC. Figure 2 It is the application environment of the server management system. The remote management platform communicates with the server management system in BMC through the network, and the server management system can exchange data with the server.
[0026] The network control module can receive and save the neural network weights sent by the remote management platform, and then send the weights to the management information generation module.
[0027] The data acquisition and preprocessing module serves as the data processing module. This module acquires server hardware information, such as real-time and comprehensive data collection of various server hardware operating data through the BMC's sensors and hardware interfaces. This includes, but is not limited to, data on the temperature, voltage, current, usage rate, and operating frequency of components such as the CPU, GPU, memory, hard drive, and power supply. The raw data is cleaned to remove outliers, duplicates, and noise. Then, the data is normalized to map data of different ranges and magnitudes to a specific standard range, improving the efficiency and accuracy of the neural network algorithm. Finally, the server hardware information is obtained and transmitted to the management information generation module.
[0028] This embodiment creates a management information generation module based on a neural network, such as a cerebellar neural network. Depending on the actual needs, convolutional neural networks, recurrent neural networks, self-attention networks, etc., can also be used.
[0029] The management information generation module adjusts the included neural network based on the weights, inputs server hardware information into the adjusted neural network, generates server management information, and transmits the server management information to the output control module. Server management information includes, for example, predicted hardware information for a future period, fault types, fault risk levels, and server hardware adjustment methods. To generate a specific type of server management information, corresponding training samples need to be constructed to train the neural network.
[0030] The output control module generates hardware control information based on server management information. For example, the output control module converts the output of the network control module into a data format that each controller can receive, thus obtaining hardware control information. This hardware control information is then sent to the peripheral controller, which uses this information to operate peripherals or adjust the server hardware.
[0031] In addition, the output control module will also transmit server management information, hardware control information and other information to the network control module; The network control module receives information from the output control module and determines whether an alarm is needed based on server management information and hardware control information. If an alarm is required, it generates an alarm message and sends it to the remote management platform or other channels. For example, if service management information determines that certain server hardware parameters exceed corresponding thresholds, and hardware control information determines that modifying server hardware parameters failed, an alarm message will be generated.
[0032] Combination Figure 3 The workflow of the inference phase of the server management system is described. The network control module receives the weights and starts working. The neural network in the management information generation module runs. Based on the output of the neural network, it is determined whether there is an abnormality in the server hardware. If an abnormality is found, an alarm is triggered. If no abnormality is found, the server can be shut down.
[0033] in addition, Figure 4 This is another server management system architecture, built upon the BMC (Block Controller Management System). The BMC contains the original processor, network interface, and system interface. New additions to the BMC include a cerebellum neural network module, a data acquisition and preprocessing module, an output control module, and a network control module. Four additional controllers and system interfaces are also added. These controllers can regulate connected peripherals, i.e., the server hardware. The network control module transmits data to the remote management platform via the processor and network interface within the BMC.
[0034] The server management system provided in this embodiment transmits received weights to the management information generation module via the network control module; the data processing module transmits acquired server hardware information to the management information generation module; the management information generation module generates server management information based on the weights and server hardware information; the output control module generates hardware adjustment information based on the server management information, adjusts the server hardware using the hardware adjustment information, and transmits the server management information and hardware adjustment information to the network control module; the network control module also generates alarm information based on the server management information and hardware adjustment information. This system enhances the BMC (Browser Control Center) function by introducing a dedicated neural network to create the management information generation module, enabling it to identify subtle fault characteristics and avoid missed fault reports; it eliminates the need for manual analysis of data or faults, improving fault handling efficiency and saving manpower. Furthermore, the dedicated neural network can be adjusted according to different needs, improving the system's scalability. It solves the problems of difficulty in identifying hardware faults caused by complex factors and the inability to adjust server hardware in a timely manner in response to hardware faults.
[0035] As an optional embodiment, the network control module is also used to receive training samples sent by the remote management platform and transmit the input information in the training samples to the data processing module; The data processing module is used to obtain target data from input information and transmit the target data to the management information generation module; The management information generation module is used to generate output results based on weights and target data, and transmit the output results to the network control module. The network control module is used to obtain updated weights based on the preset weight optimization algorithm, the output results, and the reference output information in the training samples, and then transmit the updated weights to the management information generation module.
[0036] Specifically, the network control module receives training samples sent by the remote management platform. The training samples include input information and reference output information. The network control module then transmits the input information to the data processing module.
[0037] The data processing module obtains target data from the input information. For example, it obtains various operational data of the server hardware, including but not limited to temperature, voltage, current, utilization, and operating frequency of components such as the CPU, GPU, memory, hard drive, and power supply. This data is cleaned to remove outliers, duplicates, and noisy data. Then, the data is normalized to map data of different ranges and magnitudes to a specific standard range, resulting in the target data. The target data is then transmitted to the management information generation module.
[0038] The management information generation module inputs the weights and target data into the cerebellar neural network, generates the output results, and transmits the output results to the network control module.
[0039] The network control module receives the output from the management information generation module, updates the weights using a preset weight optimization algorithm, the output results, and reference output information from the training samples, and then transmits the updated weights to the management information generation module. The preset weight optimization algorithm is a common neural network weight optimization algorithm, such as backpropagation (BP), gradient descent, or momentum method.
[0040] The above process is as follows Figure 5 As shown, the network control module works by receiving training samples and returning the weights after training the management information generation module.
[0041] In this embodiment, the quality of server hardware operation data is optimized. By relying on the cerebellum neural network and a general weight optimization algorithm, the accuracy of the output results is improved, the weights are dynamically updated, and the system's adaptability and intelligence to server management are enhanced.
[0042] As an optional embodiment, the system further includes: a peripheral controller; The output control module is used to convert server management information into hardware control information according to a preset data format, and transmit the hardware control information to the peripheral controller. The peripheral controller is used to generate operation instructions based on hardware control information and send the operation instructions to the server hardware so that the server hardware can execute the operation instructions.
[0043] Specifically, the server management system also includes peripheral controllers, such as: Figure 4 Controller 0, Controller 1, Controller 2, and Controller 3 in this document are all peripheral controllers. For simplicity, Figure 4 Only the operation of four peripheral controllers and system interfaces is listed.
[0044] The preset data format is one that each peripheral controller can receive, such as JSON. The output control module converts the received server management information into the preset data format to obtain hardware control information, which is then transmitted to the peripheral controller.
[0045] The peripheral controller generates operation commands based on hardware control information. These commands include, for example, increasing voltage, decreasing current, or increasing operating frequency. The peripheral controller then sends these commands to the server hardware, causing the server hardware to execute them and adjust the hardware to resolve or prevent hardware failures.
[0046] In this embodiment, the output control module converts server management information according to a preset format and adapts it to the peripheral controller. The latter generates targeted operation instructions to regulate the server hardware, which can promptly resolve or avoid hardware failures and improve the adaptability and accuracy of server hardware management.
[0047] As an optional embodiment, the management information generation module includes: a preset neural network; the preset neural network includes an input layer, an intermediate layer, and an output layer; The input layer is used to receive server hardware information and convert it into digital signals. The intermediate layer is used to obtain feature data based on the weights and digital signals; The output layer is used to obtain server management information based on feature data and transmit the server management information to the output control module.
[0048] Specifically, this embodiment creates a management information generation module based on a preset neural network, such as a cerebellar neural network, convolutional neural network, recurrent neural network, or self-attention network. This embodiment preferentially uses a cerebellar neural network to create the management information generation module.
[0049] The following explanation uses a cerebellar neural network as an example. The default neural network consists of an input layer, an intermediate layer, and an output layer.
[0050] The input layer is responsible for receiving input information from the data processing module, i.e., server hardware information. This information can be various types of data, such as physical signals collected by sensors, image data, text data, etc. The input layer converts it into digital signal form that the neural network can process, that is, converts server hardware information into digital signals.
[0051] The intermediate layer is responsible for processing and manipulating digital signals according to weights. Through a series of nonlinear transformations, it extracts feature data and patterns from the digital signals. The intermediate layer is also called the hidden layer. The neurons in the hidden layer adjust their connection weights through learning algorithms to achieve effective representation and transformation of input information, enabling the network to better adapt to different tasks and data distributions.
[0052] The output layer is responsible for generating the final output, namely server management information, based on the processing results (feature data) of the hidden layer. This server management information is then transmitted to the output control module.
[0053] According to an embodiment of this application, a server management method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be run in the above-mentioned server management system, or can be executed in a computer system such as a set of executable instructions, for example, a computer, a server, etc. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that shown here.
[0054] This embodiment provides a server management method. Figure 6 This is a flowchart of a server management method according to an embodiment of this application, such as... Figure 6 As shown, the process includes the following steps: Step S601: Receive the weights sent by the remote management platform and obtain the server hardware information.
[0055] Specifically, the network control module can receive and save the neural network weights sent by the remote management platform, and then send the weights to the management information generation module.
[0056] The processing module acquires server hardware information. For example, the data processing module collects various operational data of the server hardware in real time and comprehensively through the BMC's sensors and hardware interfaces, including but not limited to data on the temperature, voltage, current, usage rate, and operating frequency of components such as the CPU, GPU, memory, hard drive, and power supply. The collected raw data is cleaned to remove outliers, duplicates, and noise. Then, the data is normalized to map data of different ranges and magnitudes to a specific standard range, improving the efficiency and accuracy of the neural network algorithm in processing the data. Finally, the server hardware information is obtained and transmitted to the management information generation module.
[0057] Step S602: Generate server management information based on the weights and server hardware information.
[0058] Specifically, the management information generation module adjusts the included neural network according to the weights, inputs server hardware information into the adjusted neural network, generates server management information, and transmits the server management information to the output control module. Server management information includes, for example, predicted hardware information for a future period, fault types, fault risk levels, and server hardware adjustment methods. To generate a specific type of server management information, corresponding training samples need to be constructed to train the neural network.
[0059] Step S603: Generate hardware control information based on server management information, and adjust server hardware using the hardware control information.
[0060] Specifically, the output control module generates hardware control information based on server management information. For example, the output control module converts the output of the network control module into a data format that each controller can receive, thus obtaining hardware control information. This hardware control information is then sent to the peripheral controller, which uses this information to operate peripherals or adjust the server hardware.
[0061] Step S604: Generate alarm information based on server management information and hardware control information.
[0062] Specifically, the network control module receives information from the output control module and determines whether an alarm is needed based on server management information and hardware control information. If an alarm is required, an alarm message is generated and sent to the remote management platform or other channels. For example, if the service management information determines that certain server hardware parameters exceed corresponding thresholds, and the hardware control information determines that modifying the server hardware parameters failed, an alarm message will be generated.
[0063] The server management method provided in this embodiment enhances the BMC (Browser Management Center) function by introducing a dedicated neural network to create a management information generation module. This enables the identification of subtle fault characteristics, preventing missed fault reports. It eliminates the need for manual data or fault analysis, improving fault handling efficiency and saving manpower. Furthermore, the dedicated neural network can be adjusted according to different needs, improving the system's scalability. This solves the problems of difficulty in identifying hardware faults caused by complex factors and the inability to promptly adjust server hardware in response to such faults.
[0064] As an optional implementation, alarm information is generated based on server management information and hardware control information, including: Based on the server management information, determine whether there is server hardware information that exceeds the preset value range, and obtain the first judgment result; Obtain the risk level from the server management information and determine whether the risk level is the preset risk level to obtain the second judgment result; Based on the hardware control information, it is determined whether the operation command was successfully executed, and a third judgment result is obtained. The operation command is generated based on the hardware control information. An alarm message is generated based on the first, second, and third judgment results.
[0065] Specifically, based on server management information, the system determines whether any server hardware information exceeds preset value ranges, obtaining a first judgment result. For example, the server hardware information includes CPU temperature and power supply voltage, with corresponding preset value ranges of 0-80℃ and 1.1V-1.3V, respectively. If the CPU temperature is 85℃, exceeding the corresponding preset value range, and if the power supply voltage is 0.9V, also exceeding the corresponding preset value range, a basic alarm is triggered.
[0066] The risk level is obtained from the server management information, and it is determined whether the risk level is the preset risk level to obtain the second judgment result. For example, if the preset risk level is high, and the power supply voltage in the server hardware information is 0.9V and the cerebellum neural network predicts that the server will be shut down within 2 hours, then the risk level is high and an emergency alarm is triggered directly.
[0067] The system generates operation instructions based on hardware control information and sends these instructions to the server hardware, which then executes them.
[0068] Based on hardware control information, the system determines whether the operation command was executed successfully, obtaining a third judgment result. For example, if the execution result sent by the output control module is a failure, and the failure reason is related to a hardware fault, an alarm should also be issued. If the output control module reports that the hard drive frequency reduction command failed to execute or that there is a suspected interface fault, an alarm should also be issued even if the current hard drive usage is not exceeded.
[0069] Integrate the first, second, and third judgment results to generate alarm information.
[0070] In this embodiment, multiple dimensions are used to comprehensively judge hardware out-of-range, high-risk level, and instruction execution failure, accurately generate alarm information, reduce false alarms and missed alarms, push early warnings in a timely manner, and ensure server operation stability and business continuity.
[0071] As an optional embodiment, hardware control information is generated based on server management information, including: Obtain current and predicted hardware parameters from the server management information; Determine the trend of parameter changes based on current and predicted hardware parameters; Based on current hardware parameters, predicted hardware parameters, and parameter change trends, hardware control information is generated.
[0072] Specifically, in the field of predictive maintenance, BMC lacks the ability to deeply analyze hardware performance trends and cannot predict potential hardware failure risks in advance, which may cause servers to fail without warning, affecting business continuity.
[0073] This embodiment utilizes a neural network in the management information generation module to predict server hardware parameters, obtaining predicted hardware parameters for a future period, such as the temperature, voltage, current, utilization rate, and operating frequency of components like the CPU, GPU, memory, hard drive, and power supply. These predicted hardware parameters can then be used to adjust the server hardware.
[0074] Obtain the current and predicted hardware parameters from the server management information. By comparing the corresponding current and predicted hardware parameters, determine the trend of hardware parameter changes. For example, the hard disk latency is on an increasing trend, and the hard disk latency increases by 1-3ms per day.
[0075] Based on current hardware parameters, predicted hardware parameters, and parameter change trends, hardware control information is generated. For example, key parameters (such as power supply voltage of 0.9V and predicted hard drive read / write latency of 10ms) are extracted from both current and predicted hardware parameters. These key parameters are compared with corresponding thresholds. If they exceed the threshold, server hardware adjustments are needed. For instance, if hard drive read / write latency exceeds the threshold, operations such as disk defragmentation, checking and repairing disk errors, and optimizing the disk address structure are performed. This information is then written into the hardware control information. Parameter change trends are compared with the normal operating trend baseline. If deviations from the baseline indicate a hidden fault in the corresponding server hardware, adjustments are required. For example, if the normal operating trend baseline for hard drive latency is an increase of 1-3ms per day, and the parameter change trend is an increase of 5ms per day, a problem exists, requiring operations such as disk defragmentation, checking and repairing disk errors, and optimizing the disk address structure. This information is also written into the hardware control information. Additionally, hardware control information can be generated by combining various current and predicted hardware parameters. For example, based on current hardware parameters, it can be determined that memory bandwidth is insufficient while CPU utilization is normal. Based on predicted hardware parameters, it can be determined that the hard drive latency is increasing. After correlation verification, it can be determined that the root cause of the fault is the memory bandwidth bottleneck, rather than the hard drive itself, and that memory bandwidth needs to be increased in the future. The above information is then written into the hardware control information.
[0076] In this embodiment, by combining current and predicted hardware parameters, analyzing the trend of parameter changes to generate hardware control information, it is possible to accurately identify explicit and implicit faults, avoid hardware failure risks in advance, make up for the shortcomings of traditional BMC prediction, and ensure the continuity of server services.
[0077] As an optional embodiment, the method further includes: Receive training samples sent by the remote management platform and extract target data from the input information contained in the training samples; Input the target data and weights into a preset neural network to obtain the output of the preset neural network; Based on the preset weight optimization algorithm, the output results, and the reference output information in the training samples, the updated weights are obtained, so that the preset neural network can generate new server management information based on the updated weights and server hardware information.
[0078] Specifically, the network control module receives training samples sent by the remote management platform. The training samples include input information and reference output information. The network control module then transmits the input information to the data processing module.
[0079] The data processing module obtains target data from the input information. For example, it obtains various operational data of the server hardware, including but not limited to temperature, voltage, current, utilization, and operating frequency of components such as the CPU, GPU, memory, hard drive, and power supply. This data is cleaned to remove outliers, duplicates, and noisy data. Then, the data is normalized to map data of different ranges and magnitudes to a specific standard range, resulting in the target data. The target data is then transmitted to the management information generation module.
[0080] The management information generation module inputs the weights and target data into the cerebellar neural network, generates the output results, and transmits the output results to the network control module.
[0081] The network control module receives the output from the management information generation module, updates the weights using a preset weight optimization algorithm, the output results, and reference output information from the training samples, and then transmits the updated weights to the management information generation module. The preset weight optimization algorithm is a common neural network weight optimization algorithm, such as backpropagation (BP), gradient descent, or momentum method.
[0082] As an optional embodiment, the above step "obtaining the updated weights based on the preset weight optimization algorithm, the output results, and the reference output information in the training samples" may include steps A1 to A4.
[0083] Step A1: Collect key data.
[0084] Specifically, the system obtains reference output information from the training samples, such as standard answers for fault diagnosis and status assessment, from the remote management system; at the same time, it receives the actual output results after the cerebellar neural network operation, and summarizes both types of data into the network control module to provide the core basis for weight adjustment.
[0085] Step A2: Calculate the error.
[0086] Specifically, the network control module compares the actual output with the reference output information and quantifies the degree of deviation between the two through preset methods such as mean square error and cross-entropy, thereby clarifying the prediction / diagnosis error of the neural network under the current weights.
[0087] Step A3: Algorithm-driven adjustment.
[0088] Specifically, the network control module calls general optimization algorithms such as the backpropagation algorithm and gradient descent method. First, the error is propagated backward from the output layer of the cerebellar neural network to the intermediate layer and the input layer. The contribution (gradient) of the connection weights of each neuron to the error is calculated layer by layer. Then, according to the preset learning rate, the weights are corrected along the gradient descent direction, that is, the weights that have a positive effect on reducing the error are increased and the weights that have a negative effect are decreased, so as to minimize the error and improve the accuracy of model diagnosis and prediction.
[0089] Step A4, Output and Feedback.
[0090] Specifically, the network control module saves the updated weights and feeds them back to the cerebellum neural network module to replace the old weights for use in the next round of training sample calculations or real-time data inference.
[0091] In this embodiment, core data is accurately collected, output deviations are quantified, weights are optimized in reverse using a general algorithm and fed back for reuse, and the accuracy of neural network diagnosis and prediction is iteratively improved, effectively making up for the shortcomings of traditional BMC such as missed reports and insufficient real-time performance.
[0092] This embodiment also provides a server management device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0093] This embodiment provides a server management device, such as... Figure 7 As shown, it includes: The information acquisition module 701 is used to receive weights sent by the remote management platform and acquire server hardware information; The management information generation module 702 is used to generate server management information based on the weights and server hardware information. The hardware control module 703 is used to generate hardware control information based on server management information and to adjust the server hardware using the hardware control information. The alarm information generation module 704 is used to generate alarm information based on server management information and hardware control information.
[0094] In some optional implementations, the alarm information generation module 704 includes: The first judgment unit is used to determine whether there is server hardware information that exceeds the preset value range based on server management information, and to obtain the first judgment result. The second judgment unit is used to obtain the risk level from the server management information, determine whether the risk level is the preset risk level, and obtain the second judgment result. The third judgment unit is used to determine whether the operation instruction has been successfully executed based on the hardware control information and to obtain the third judgment result. The operation instruction is generated based on the hardware control information. The first generation unit is used to generate alarm information based on the first judgment result, the second judgment result, and the third judgment result.
[0095] In some alternative implementations, the hardware control module 703 includes: The acquisition unit is used to obtain current hardware parameters and predicted hardware parameters from the server management information; The determination unit is used to determine the trend of parameter changes based on the current hardware parameters and the predicted hardware parameters; The second generation unit is used to generate hardware control information based on the current hardware parameters, predicted hardware parameters, and parameter change trends.
[0096] In some alternative embodiments, the device further includes: The acquisition module is used to receive training samples sent by the remote management platform and extract target data from the input information contained in the training samples. The module is used to input the target data and weights into a preset neural network and obtain the output of the preset neural network. The generation module is used to obtain updated weights based on the preset weight optimization algorithm, the output results, and the reference output information in the training samples, so that the preset neural network can generate new server management information based on the updated weights and server hardware information.
[0097] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0098] In this embodiment, the server management device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0099] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0100] The following is a detailed reference. Figure 8This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 801, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 802 or a program loaded from memory 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device. The processor 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0101] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0102] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a memory 808, or installed from a ROM 802. When the computer program is executed by the processor 801, it performs the functions defined in the server management method of the embodiments of the present invention.
[0103] Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0104] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the server management method shown in the above embodiments is implemented.
[0105] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0106] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A server management system, characterized in that, The system includes: a network control module, a data processing module, a management information generation module, and an output control module; The network control module is used to receive weights sent by the remote management platform and transmit the weights to the management information generation module; The data processing module is used to acquire server hardware information and transmit the server hardware information to the management information generation module; The management information generation module is used to generate server management information based on the weights and the server hardware information, and transmit the server management information to the output control module. The output control module is used to generate hardware control information based on the server management information, adjust the server hardware using the hardware control information, and transmit the server management information and the hardware control information to the network control module. The network control module is also used to generate alarm information based on the server management information and the hardware control information.
2. The system according to claim 1, characterized in that, The network control module is also used to receive training samples sent by the remote management platform and transmit the input information in the training samples to the data processing module. The data processing module is used to obtain target data from the input information and transmit the target data to the management information generation module. The management information generation module is used to generate output results based on the weights and the target data, and transmit the output results to the network control module. The network control module is used to obtain updated weights based on a preset weight optimization algorithm, the output results, and reference output information in the training samples, and then transmit the updated weights to the management information generation module.
3. The system according to claim 1, characterized in that, The system also includes: a peripheral controller; The output control module is used to convert the server management information into hardware control information according to a preset data format, and transmit the hardware control information to the peripheral controller. The peripheral controller is used to generate operation instructions based on the hardware control information and send the operation instructions to the server hardware so that the server hardware executes the operation instructions.
4. The system according to claim 1, characterized in that, The management information generation module includes: a preset neural network; the preset neural network includes an input layer, an intermediate layer, and an output layer; The input layer is used to receive the server hardware information and convert the server hardware information into digital signals; The intermediate layer is used to obtain feature data based on the weights and the digital signal; The output layer is used to obtain the server management information based on the feature data and transmit the server management information to the output control module.
5. A server management method, characterized in that, The method includes: Receive weights sent by the remote management platform and obtain server hardware information; Based on the weights and the server hardware information, server management information is generated; Hardware control information is generated based on the server management information, and the server hardware is adjusted using the hardware control information. An alarm message is generated based on the server management information and the hardware control information.
6. The method according to claim 5, characterized in that, The step of generating alarm information based on the server management information and the hardware control information includes: Based on the server management information, determine whether there is server hardware information that exceeds the preset value range, and obtain a first determination result; The risk level is obtained from the server management information, and it is determined whether the risk level is a preset risk level to obtain a second judgment result; Based on the hardware control information, it is determined whether the operation instruction was successfully executed, and a third determination result is obtained, wherein the operation instruction is generated based on the hardware control information; The alarm information is generated based on the first judgment result, the second judgment result, and the third judgment result.
7. The method according to claim 5, characterized in that, The step of generating hardware control information based on the server management information includes: Obtain the current hardware parameters and predicted hardware parameters from the server management information; Based on the current hardware parameters and the predicted hardware parameters, determine the parameter change trend; The hardware control information is generated based on the current hardware parameters, the predicted hardware parameters, and the trend of parameter changes.
8. The method according to claim 5, characterized in that, The method further includes: Receive training samples sent by the remote management platform, and obtain target data from the input information contained in the training samples; The target data and the weights are input into a preset neural network to obtain the output of the preset neural network; Based on the preset weight optimization algorithm, the output result, and the reference output information in the training samples, the updated weights are obtained, so that the preset neural network can generate new server management information based on the updated weights and the server hardware information.
9. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the server management method of any one of claims 5 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the server management method according to any one of claims 5 to 8.