Server configuration management system based on Redfish protocol

Through the server configuration management system based on the Redfish protocol, the existing remote server management solutions are solved, and efficient, intelligent and secure remote server management is achieved.

CN120223498AInactive Publication Date: 2025-06-27POWERLEADER COMPUTER SYST CO LTD

Patent Information

Application Number
CN202510664026.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing remote server management solutions have problems such as low efficiency, poor compatibility, insufficient security, insufficient fault prediction capabilities, and complex firmware upgrades.

Method used

The server configuration management system based on the Redfish protocol is adopted, including the API interface layer, the Redfish device layer and the management control layer. Through technical means such as Redfish API, AI intelligent analysis, automated management and security authentication, efficient, intelligent and secure remote management of servers is achieved.

Benefits of technology

Improve operation and maintenance efficiency, reduce management costs, enhance security, improve fault prediction capabilities, and simplify the firmware upgrade process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223498A_ABST
    Figure CN120223498A_ABST
Patent Text Reader

Abstract

The invention relates to a server configuration management system based on a Redfish protocol, and the system comprises an API interface layer which is used for providing an access interface for the outside; the Redfish equipment layer is used for realizing remote management of a server through a Redfish protocol, and the Redfish equipment layer is connected with the API interface layer; and the management control layer is used for carrying out comprehensive management on configuration, faults, compatibility and firmware conditions of the multi-server cluster, and the management control layer is connected with the Redfish equipment layer. Through Redfish API, AI intelligent analysis, automatic management, security authentication and other technical means, a set of efficient, intelligent and safe server remote management scheme is constructed, the operation and maintenance efficiency is effectively improved, and the management cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of server operation and maintenance, and particularly to a server configuration management system, device and storage medium based on the Redfish protocol. Background Art

[0002] With the development of technologies such as cloud computing, artificial intelligence, and big data, the scale of data centers has been expanding day by day, and the demand for remote management of servers has become increasingly strong. Modern enterprises and cloud service providers (such as AWS, Google Cloud, Microsoft Azure) need to manage thousands of servers, but traditional management methods rely too much on on-site operation and maintenance, making it difficult to meet the management requirements of high efficiency, low cost, and intelligence. Remote server management is a core link in data center infrastructure management (DCIM), involving functions such as hardware monitoring, power management, BIOS configuration, firmware upgrade, and fault prediction of servers. Existing remote management protocols (such as IPMI, SNMP) have problems such as low security, limited functions, and insufficient standardization.

[0003] Currently, the main methods of server remote management include:

[0004] (1) IPMI (Intelligent Platform Management Interface), a traditional server remote management protocol, widely used in enterprise-level servers. It provides functions such as power management, hardware monitoring, and remote control. The disadvantages are low security, an old protocol, lack of support for modern RESTful APIs, and complex configuration.

[0005] (2) SNMP (Simple Network Management Protocol), mainly used for network device management and also applicable to server monitoring. The disadvantages are complex data structures, poor security, lack of support for REST APIs, and unsuitability for large-scale server management.

[0006] (3) Redfish (a remote management standard developed by DMTF), which adopts RESTful APIs and can interact through JSON, suitable for Web and cloud environments. It provides functions such as server power management, BIOS configuration, RAID management, and firmware update. It supports HTTPS encryption, has higher security, and can replace IPMI.

[0007] Most existing remote management solutions rely on vendor - specific protocols, with poor compatibility and low efficiency in server remote management; traditional management methods require manual configuration of each server, which is inefficient and error - prone, and difficult for batch management; existing solutions usually only monitor basic hardware status, unable to achieve in - depth analysis, and the hardware status monitoring is not comprehensive; traditional operation and maintenance rely on manual monitoring, making it difficult to detect potential faults in a timely manner, and the fault prediction ability is insufficient; server firmware upgrades or hardware replacements may cause compatibility problems, and existing methods mainly rely on manual testing, which is time - consuming and laborious, and difficult for compatibility analysis; existing remote management solutions are relatively weak in authentication and authorization control, vulnerable to attacks, and lack security; traditional firmware upgrades require manual operations, prone to problems such as version mismatch and upgrade failure. Summary of the Invention

[0008] The present invention provides a server configuration management system, device, and storage medium based on the Redfish protocol, aiming to solve at least one of the technical problems existing in the prior art.

[0009] The technical solution of the present invention is a server configuration management system based on the Redfish protocol, including:

[0010] An API interface layer for providing access interfaces externally;

[0011] A Redfish device layer for realizing remote management of the server through the Redfish protocol, and the Redfish device layer is connected to the API interface layer;

[0012] A management control layer for comprehensively managing the configuration, faults, compatibility, and firmware conditions of a multi - server cluster, and the management control layer is connected to the Redfish device layer.

[0013] Further, the API interface layer includes a unified API proxy module, an authentication and authorization management module, and a log record and access audit module connected in sequence.

[0014] The unified API proxy module is built based on the Python and Flask frameworks and encapsulates multiple source RESTful APIs to provide a unified interface externally;

[0015] The authentication and authorization management module performs identity authentication through JWT (JSON Web Token) and API access control based on role - based access control (RBAC);

[0016] The log record and access audit module records behaviors such as API calls, user identities, and access times through the Elasticsearch full - text search engine and the Kibana analysis and visualization platform, and generates visual logs and security audit reports.

[0017] Furthermore, the Redfish device layer includes: a server power management module, a hardware status monitoring module, a RAID configuration management module, a BIOS remote configuration module, and a firmware update management module;

[0018] The server power management module sends RESTful requests for operating the server power through the Redfish API, and the RESTful requests for operating the server power can remotely batch-operate the power-on, power-off, and restart of the server;

[0019] The hardware status monitoring module collects the health status data of the server through the Redfish API and stores it in the Prometheus time series database, and performs visual display through the Grafana analysis monitoring platform. The health status data of the server at least includes CPU, memory, disk, temperature, and fan status data;

[0020] The RAID configuration management module configures the storage controller through the Redfish API, including remotely creating RAID0, RAID 1, and RAID 5 to improve storage redundancy and data security;

[0021] The BIOS remote configuration module improves the compatibility of the server by remotely adjusting the BIOS startup parameters;

[0022] The firmware update management module is used to remotely batch-update the BIOS firmware and BMC firmware.

[0023] Furthermore, the management control layer includes a server batch configuration module, a server fault prediction module, a compatibility analysis module, and a remote firmware management module;

[0024] The server batch configuration module realizes batch initialization of the server BIOS and RAID configuration through the Ansible operation and maintenance tool and the Redfish API;

[0025] The server fault prediction module collects the server health data through the Redfish API, stores the data using the Prometheus time series database, performs visual analysis through the Grafana analysis monitoring platform, and predicts the health status of the server through the first prediction model to evaluate the probability of the risk of failure;

[0026] The first prediction model includes an XGBoost model or an LSTM model;

[0027] The compatibility analysis module predicts whether system compatibility issues will occur after firmware upgrade or hardware replacement through the second prediction model, avoiding crashes or anomalies after the upgrade;

[0028] The second prediction model includes a logistic regression model, a support vector machine (SVM) model, or a random forest model;

[0029] The remote firmware management module remotely updates the BIOS and BMC firmware in batches through the Redfish API.

[0030] Furthermore, the XGBoost model includes a data input sub-module, at least one decision tree sub-module, a weighted integration sub-module, and a prediction sub-module connected in sequence.

[0031] The data input sub-module is used to collect operation data such as the server's CPU temperature, fan speed, and disk I / O; at least one decision tree sub-module is iteratively optimized through the Boosting method to gradually reduce errors; the prediction sub-module outputs a predicted value of the future failure risk percentage.

[0032] Furthermore, the LSTM model includes a historical data input sub-module, a long short-term memory (LSTM) network sub-module, and a prediction sub-module connected in sequence;

[0033] The historical data input sub-module is used to input the CPU load and temperature data of the server within a preset time range in the past; the long short-term memory (LSTM) network sub-module processes long-term dependent data through a memory cell; the prediction sub-module predicts the health status of the server within a preset time range in the future.

[0034] Furthermore, the logistic regression model includes a configuration data input sub-module, a linear transformation sub-module, and a Sigmoid function processing sub-module connected in sequence;

[0035] The configuration data input sub-module is used to input the hardware information of the server, and the hardware information at least includes the CPU model and BIOS version; the Sigmoid function processing sub-module is used to output the upgrade compatibility probability.

[0036] Furthermore, the support vector machine (SVM) model is used to predict whether a new firmware is suitable for specific server hardware, and the support vector machine (SVM) model includes a data input sub-module, a feature mapping sub-module, a hyperplane classification sub-module, and a result output sub-module connected in sequence;

[0037] The data input sub-module is used to input the hardware parameters of the server, and the hardware parameters of the server at least include BIOS version, CPU model, and storage controller model; the feature mapping sub-module is used to extract the server hardware parameters and map them to a high-dimensional space; the hyperplane classification sub-module is used to calculate the optimal decision boundary through SVM; the result output sub-module is used to output compatible or incompatible results.

[0038] Furthermore, the random forest model is used for multi-factor health assessment of the server. The random forest model includes a data input sub-module, a decision tree sub-module, and a voting decision sub-module connected in sequence; the data input sub-module is used to input the server status data; the decision tree sub-module includes multiple decision trees, and each decision tree is set in parallel; the voting decision sub-module makes a voting decision based on the output of the decision tree sub-module and outputs a health score.

[0039] The beneficial effects of the present invention are:

[0040] The server configuration management system based on the Redfish protocol constructs an efficient, intelligent, and secure server remote management solution through technical means such as Redfish API, AI intelligent analysis, automated management, and security authentication, effectively improving the operation and maintenance efficiency and reducing the management cost. Brief Description of the Drawings

[0041] Figure 1 It is a schematic diagram of the network topology of the server configuration management system based on the Redfish protocol.

[0042] Figure 2 It is a schematic diagram of an instance of the network topology of the server configuration management system based on the Redfish protocol. Detailed Embodiments

[0043] The following will clearly and completely describe the concept, specific structure, and technical effects generated by the present invention in combination with the embodiments and the drawings, so as to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0044] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to another feature, or indirectly fixed or connected to another feature. In addition, the up, down, left, right, top, bottom, etc. used in the present invention are only relative to the mutual positional relationship of the components of the present invention in the drawings.

[0045] In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this technology belongs. The terms used in the description of this specification are only for describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any combination of one or more of the related listed items.

[0046] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element.

[0047] Referring to Figure 1 , in some embodiments, the technical solution of the present invention is a server configuration management system based on the Redfish protocol, including:

[0048] The API interface layer is used to provide access interfaces externally;

[0049] The Redfish device layer is used to achieve remote management of the server through the Redfish protocol, and the Redfish device layer is connected to the API interface layer;

[0050] The management control layer is used to comprehensively manage the configuration, faults, compatibility, and firmware status of the multi-server cluster, and the management control layer is connected to the Redfish device layer.

[0051] The beneficial effects of the present invention are:

[0052] The described server configuration management system based on the Redfish protocol constructs a set of efficient, intelligent, and secure server remote management solutions through technical means such as Redfish API, AI intelligent analysis, automated management, and security authentication, effectively improving the operation and maintenance efficiency and reducing the management cost.

[0053] Specifically, the present invention mainly solves the following technical problems:

[0054] (1) Low efficiency of server remote management

[0055] Most existing remote management solutions rely on vendor-proprietary protocols, have poor compatibility, and complex operations. The present invention provides a unified remote management interface based on the Redfish standard, simplifies the management process, and improves the operation efficiency.

[0056] (2) Difficulty in batch management

[0057] Traditional management methods require manual configuration for each server, which is inefficient and error-prone. The present invention supports Ansible automated configuration, enabling one-click deployment of large-scale servers, batch BIOS configuration, batch RAID configuration, etc., improving the operation and maintenance efficiency.

[0058] (3) Incomplete monitoring of hardware status

[0059] Existing solutions usually only monitor the basic hardware status and cannot achieve in-depth analysis. The present invention combines Prometheus and Grafana, supports multi-dimensional data collection of CPU, memory, disk, temperature, fan, etc., and provides visual analysis to enhance the server health monitoring ability.

[0060] (4) Insufficient fault prediction ability

[0061] Traditional operation and maintenance rely on manual monitoring, making it difficult to detect potential faults in a timely manner. The present invention introduces AI (XGBoost / LSTM) intelligent analysis, predicts server faults based on historical data, provides early warnings, and reduces the risk of downtime.

[0062] (5) Difficult compatibility analysis

[0063] Server firmware upgrades or hardware replacements may cause compatibility issues. Existing methods mainly rely on manual testing, which is time-consuming and laborious. The present invention is based on an AI compatibility analysis model, which can evaluate compatibility before upgrades to avoid system crashes or anomalies.

[0064] (6) Insufficient security

[0065] Existing remote management solutions are relatively weak in authentication and permission control and are easily attacked. The present invention adopts JWT authentication + RBAC role-based permission management to ensure the security of API access and prevent unauthorized operations.

[0066] (7) Complicated firmware upgrade and error-prone

[0067] Traditional firmware upgrades require manual operations, which are prone to problems such as version mismatch and upgrade failure. The present invention supports Redfish API remote firmware management, which can automatically download and batch deploy firmware, improving the upgrade success rate.

[0068] Through technical means such as Redfish API, AI intelligent analysis, automated management, and security authentication, the present invention constructs a set of efficient, intelligent, and secure server remote management solutions, effectively improving operation and maintenance efficiency and reducing management costs. It adopts HTTPS, JWT authentication, and RBAC role management to enhance security; conducts server health prediction based on LSTM / XGBoost to early warn of hardware failures; uses Ansible / SaltStack for large-scale batch management; conducts API compatibility prediction based on SVM / random forest to support multi-brand management; and supports BIOS parameter adjustment, RAID configuration, and firmware update.

[0069] Referring to Figure 2 , the intelligent server remote management system based on the Redfish protocol of the present invention has the following main advantages compared with the existing most advanced management technologies (such as IPMI, SNMP, traditional Redfish API solutions):

[0070] Higher security. It adopts HTTPS + JWT authentication + RBAC (role-based access control) to prevent unauthorized access. Log auditing + anomaly detection can monitor behaviors such as hacker attacks and brute-force cracking. API access encryption avoids security vulnerabilities caused by plaintext transmission. AI-driven intelligent prediction combines LSTM (long short-term memory network) and XGBoost (gradient boosting tree) for server health prediction, which can early warn of hardware failures such as hard disks, power supplies, and fans 7 - 14 days in advance, avoiding sudden downtime. Existing technologies (such as traditional Redfish solutions) can only monitor the current status and cannot predict future failures.

[0071] Stronger batch management ability. The existing Redfish solutions are mainly applicable to single-server management, while the present invention supports Ansible / SaltStack batch management, which can control thousands of servers simultaneously, improving operation and maintenance efficiency. One-key batch adjustment of BIOS parameters, RAID configuration, and remote restart without the need for individual operations. Solve the Redfish API compatibility problem. In the existing technology, the Redfish API structures of different manufacturers are different, resulting in difficulties in cross-brand management. The present invention adopts the SVM / random forest algorithm to automatically adapt to different brand APIs, provides a standardized management interface, and realizes cross-brand compatibility.

[0072] Supports remote BIOS, RAID configuration, and firmware upgrade. The existing IPMI / Redfish solutions cannot remotely modify BIOS parameters, while the present invention supports remote BIOS adjustment, RAID configuration, and firmware upgrade without the need for on-site operations.

[0073] Furthermore, referring to Figure 1 , the API interface layer includes a unified API proxy module, an identity authentication and permission management module, and a log recording and access auditing module that are connected in sequence.

[0074] The unified API proxy module is built based on the Python and Flask frameworks and encapsulates multiple source RESTful APIs, providing a unified interface externally;

[0075] The identity authentication and permission management module performs identity authentication through JWT (JSON Web Token) and API access control based on role-based access control (RBAC);

[0076] The log recording and access auditing module records behaviors such as API calls, user identities, and access times through the Elasticsearch full-text search engine and the Kibana analysis and visualization platform, generating visual logs and security audit reports.

[0077] In a specific embodiment,

[0078] (1) The example code of the unified API proxy module is as follows:

[0079] from flask import Flask, jsonify, request

[0080] import requests

[0081] app = Flask(__name__)

[0082] BMC_IP = "(This is a BMC management IP address)"

[0083] USERNAME = "admin"

[0084] PASSWORD = "password"

[0085] @app.route(' / server / status', methods=['GET'])

[0086] def get_server_status():

[0087] url = f"(This is an http network address for querying the server status)"

[0088] response = requests.get(url, auth=(USERNAME, PASSWORD), verify=False)

[0089] return jsonify(response.json())

[0090] if __name__ == '__main__':

[0091] app.run(host='0.0.0.0', port=5000)

[0092] (2)The sample code for the identity authentication and permission management module is as follows:

[0093] import jwt

[0094] import datetime

[0095] SECRET_KEY = "your_secret_key"

[0096] def generate_token(username, role):

[0097] payload = {

[0098] "exp": datetime.datetime.utcnow() + datetime.timedelta(days=1),

[0099] "iat": datetime.datetime.utcnow(),

[0100] "sub": username,

[0101] "role": role}

[0102] return jwt.encode(payload, SECRET_KEY, algorithm="HS256")

[0103] (3)The sample code for the logging and access auditing module is as follows:

[0104] import logging

[0105] logging.basicConfig(filename="access.log", level=logging.INFO)

[0106] @app.before_request

[0107] def log_request():

[0108] logging.info(f"Method: {request.method}, URL: {request.url}")

[0109] Furthermore, referring to Figure 1 , the Redfish device layer includes: a server power management module, a hardware status monitoring module, a RAID configuration management module, a BIOS remote configuration module, and a firmware update management module;

[0110] The server power management module sends RESTful requests for operating the server power through the Redfish API, and the RESTful requests for operating the server power can remotely batch-operate the power-on, power-off, and restart of the server;

[0111] The hardware status monitoring module collects the health status data of the server through the Redfish API and stores it in the Prometheus time series database, and visualizes it through the Grafana analysis monitoring platform. The health status data of the server at least includes CPU, memory, disk, temperature, and fan status data;

[0112] The RAID configuration management module configures the storage controller through the Redfish API, including remotely creating RAID0, RAID 1, and RAID 5 to improve storage redundancy and data security;

[0113] The BIOS remote configuration module improves the compatibility of the server by remotely adjusting the BIOS startup parameters;

[0114] The firmware update management module is used to remotely batch-update the BIOS firmware and BMC firmware.

[0115] The Redfish device layer relies on the BMC (Baseboard Management Controller) and realizes the remote management of the server through the Redfish protocol, including power management, hardware monitoring, RAID configuration, BIOS configuration, firmware update, etc. In a specific embodiment,

[0116] (1)The server power management module is used for remote power-on, power-off, and reboot, which is suitable for batch operations of large-scale servers. It is implemented by sending RESTful requests through the Redfish API to operate the server power. The following is an example code for remote power-on:

[0117] curl -k -u admin:password -X POST -d '{"ResetType": "On"}' \

[0118] (Here is the http network address for the request to operate the server power)

[0119] (2)The hardware status monitoring module is used to collect data such as CPU, memory, disk, temperature, and fan status, and perform health analysis. It is implemented by collecting server health status data through the Redfish API, storing it in Prometheus for long-term storage, and using Grafana for visual display. The following is an example code for querying temperature information:

[0120] curl -k -u admin:password (Here is the http network address for the request to obtain server health status data)

[0121] (3)The RAID configuration management module is used to remotely create RAID 0, RAID 1, and RAID 5 to improve storage redundancy and data security. It is implemented by configuring the storage controller using the Redfish API. The following is an example code for creating RAID 1:

[0122] { "Name": "RAID 1",

[0123] "VolumeType": "Mirrored",

[0124] "Drives": [" / redfish / v1 / Systems / 1 / Storage / Drives / 0", " / redfish / v1 / Systems / 1 / Storage / Drives / 1"]}

[0125] (4)The BIOS remote configuration module is used to remotely adjust BIOS startup parameters to improve server compatibility. The following is an example code for modifying the startup mode to UEFI:

[0126] { "Attributes": {

[0127] "BootMode": "UEFI"}}

[0128] The (5) firmware update management module is used to remotely update the BIOS and BMC firmware in batches. The following is an example code for firmware upgrade:

[0129] curl -k -u admin:password -X POST -d '{"ImageURI": " (Here is the address of the firmware icon)"}' \

[0130] (Here is the http network address for remotely updating the BIOS and BMC firmware in batches)

[0131] Furthermore, referring to Figure 1 , the management control layer includes a server batch configuration module, a server fault prediction module, a compatibility analysis module, and a remote firmware management module;

[0132] The server batch configuration module realizes batch initialization of the server BIOS and RAID configurations through the Ansible operation and maintenance tool and the Redfish API;

[0133] The server fault prediction module collects server health data through the Redfish API, stores the data using the Prometheus time series database, performs visual analysis through the Grafana analysis and monitoring platform, and predicts the health status of the server through the first prediction model to evaluate the probability of the risk of failure;

[0134] The first prediction model includes an XGBoost model or an LSTM model;

[0135] The compatibility analysis module predicts whether system compatibility problems will occur after firmware upgrade or hardware replacement through the second prediction model, avoiding crashes or anomalies after upgrade;

[0136] The second prediction model includes a logistic regression model or a support vector machine SVM model or a random forest model;

[0137] The remote firmware management module remotely updates the BIOS and BMC firmware in batches through the Redfish API.

[0138] Specifically,

[0139] The (1) server batch configuration module is used to batch initialize the server, including configurations such as BIOS and RAID, and realizes batch management through the use of Ansible + Redfish API. The following is part of the implementation code:

[0140] - name: Configure BIOS boot mode

[0141] redfish_command:

[0142] category: Systems

[0143] command: SetBIOS

[0144] target: "{{ BMC_IP}}"

[0145] username: "admin"

[0146] password: "password"

[0147] bios_settings:

[0148] BootMode: "UEFI"

[0149] (2)The server fault prediction module is used to collect key data such as CPU, memory, disk, fan speed, temperature, power status, etc., and use the first prediction model to predict possible server faults and issue early warnings. Among them, the technical implementation methods include:

[0150] Data collection: Collect server health data (such as CPU temperature, fan speed, etc.) through the Redfish API.

[0151] Data storage: Use Prometheus for monitoring data storage and perform visual analysis through Grafana.

[0152] The first prediction model: Use the XGBoost or LSTM (Long Short-Term Memory Network) model to predict the server health status. The training data includes historical server fault records, sensor data, workload changes, etc.

[0153] Further, referring to Figure 1 ,the XGBoost model includes a data input sub-module, at least one decision tree sub-module, a weighted integration sub-module, and a prediction sub-module connected in sequence.

[0154] The data input sub-module is used to collect operation data such as server CPU temperature, fan speed, disk I / O, etc.; at least one decision tree sub-module is iteratively optimized by the Boosting method to gradually reduce errors; the prediction sub-module outputs a predicted value of the future fault risk percentage.

[0155] XGBoost (Extreme Gradient Boosting) is an optimized algorithm for gradient boosting decision trees (GBDT), which is commonly used for regression and classification problems of structured data. It improves the prediction ability of the model through multiple iterations (boosting), can efficiently process structured data, and is applicable to server failure prediction. Specifically, it includes:

[0156] Input data (features): Features are the basis for the model to learn. They are server status or sensor data, such as CPU temperature, memory temperature, fan speed, etc. The dimension and features of the input data are the key to determining the performance of the model.

[0157] Decision Tree: Each decision tree in the model is responsible for splitting different paths according to the input data and finally making predictions. Each tree has multiple nodes and leaf nodes. Each node splits according to a certain feature, and finally each leaf node gives a prediction value.

[0158] Boosting Iteration: In each iteration, by weighting and correcting the errors of the previous decision tree, the ability of the new tree to correct the previous errors is increased. Each new decision tree will correct the errors of the previous model and improve the accuracy of the model.

[0159] Weighted Integration: The prediction results of all decision trees are weighted and fused, and the final prediction result is obtained by weighted summation. The contribution of each tree is different, and different weights are assigned according to its accuracy.

[0160] Final Prediction Output: Through the prediction result after weighted integration, the model outputs the final prediction result, which is usually a regression value or a classification label.

[0161] Input data (server status data: CPU temperature, memory, disk usage, etc.), multiple decision trees (Boosting iterative training), weighted integration (improving model performance), predicting server failure risk (0 - 100%). The technical implementation method is as follows:

[0162] 1. Data input: Collect server operation data (such as CPU temperature, fan speed, disk I / O, etc.);

[0163] 2. Training of multiple decision trees: Iteratively optimize through the Boosting method to gradually reduce errors;

[0164] 3. Final prediction result: Output the percentage of the future failure risk of the server.

[0165] In a specific embodiment, the example code for XGBoost failure prediction is as follows:

[0166] import xgboost as xgb

[0167] import pandas as pd

[0168] # Read server operation data

[0169] df = pd.read_csv("server_health_data.csv")

[0170] # Select features

[0171] X = df[['CPU_Temp', 'Memory_Temp', 'Fan_Speed', 'Disk_Usage', 'Power_Consumption']]

[0172] y = df['Failure_Probability'] # Server failure probability

[0173] # Train the XGBoost model

[0174] model = xgb.XGBRegressor(n_estimators=100)

[0175] model.fit(X, y)

[0176] # Predict the health status of a new server

[0177] new_server_data = [[70, 60, 1200, 80, 200]] # 70°C CPU temperature, 60°C memory temperature, 1200 RPM fan speed, etc.

[0178] failure_risk = model.predict(new_server_data)

[0179] print(f"Server failure risk probability: {failure_risk[0]:.2%}")

[0180] The application effect is as follows: If the failure risk probability is higher than 80%, the system will give an early warning and provide optimization suggestions (such as reducing the CPU load, checking the cooling system, etc.).

[0181] (3)The compatibility analysis module is used to predict whether firmware upgrades and hardware replacements will cause system compatibility issues, avoiding crashes or anomalies after upgrades. By collecting server hardware configurations (CPU model, BIOS version, RAID controller, etc.) and training a compatibility prediction model, the compatibility of different components is analyzed, and algorithms such as logistic regression, SVM (Support Vector Machine), and random forest are used to predict the compatibility of different hardware combinations. The following is an example code for machine learning compatibility prediction:

[0182] import numpy as np

[0183] import joblib

[0184] # Load the trained compatibility analysis model

[0185] model = joblib.load("compatibility_model.pkl")

[0186] # Current server hardware configuration (example)

[0187] # [BIOS version, CPU generation, memory size, RAID version, power supply power]

[0188] new_config = np.array([[1.3, 10, 128, 5, 750]])

[0189] # Predict the upgrade risk

[0190] risk = model.predict(new_config)

[0191] if risk[0] == 1:

[0192] print("Danger: This upgrade may cause server compatibility issues!")

[0193] else:

[0194] print("Compatibility check passed, safe to upgrade.")

[0195] (4)The remote firmware management module remotely updates BIOS and BMC firmware through the Redfish API, supporting batch updates. The following is an example code:

[0196] curl -k -u admin:password -X POST -d '{"ImageURI": " (Firmware icon address here)"}' \

[0197] (Here is the request http network address for remotely updating the BIOS and BMC firmware)

[0198] Logistic regression is used for binary classification problems and is used to predict whether the firmware upgrade is compatible (0 = incompatible, 1 = compatible).

[0199] Specifically, the processing process of logistic regression sequentially includes steps of inputting data (current hardware configuration, BIOS version, RAID controller, etc.), linear transformation (weighted summation), Sigmoid activation function (outputting probability), and predicting upgrade compatibility (0 = incompatible, 1 = compatible). The technical implementation is as follows:

[0200] 1. Input hardware information (such as CPU model, BIOS version, etc.);

[0201] 2. Linear transformation calculation (weighted summation);

[0202] 3. Process by the Sigmoid function to output the upgrade compatibility probability (such as 90% compatible).

[0203] Furthermore, referring to Figure 1 , the LSTM model includes a historical data input sub-module, a long short-term memory network LSTM sub-module, and a prediction sub-module connected in sequence;

[0204] The historical data input sub-module is used to input the CPU load and temperature data of the server within a preset time range in the past; the long short-term memory network LSTM sub-module processes long-term dependent data through the memory gate Memory Cell; the prediction sub-module predicts the health status of the server within a preset time range in the future.

[0205] Specifically, LSTM (long short-term memory network) is a type of RNN (recurrent neural network), which is a special type of recursive neural network (RNN) specifically used to process and predict time series data. LSTM can remember long-term dependencies and is suitable for processing tasks with time correlations, such as predicting the health status of a server. It is applicable to time series prediction and can predict the future health status of a server based on historical data.

[0206] Specific component units and their characteristics:

[0207] (1) Input data (time series): The data is arranged in chronological order and usually contains multiple parameters for server health monitoring (such as CPU temperature, memory usage, fan speed, etc.). The input data is gradually passed into the LSTM unit, and the data at each time step affects the output of the next step.

[0208] (2) LSTM Unit: The LSTM unit is the core part of this network. It uses three gates (input gate, forget gate, output gate) to control the flow of information so that the model can remember important information and discard unimportant information. Each unit has an internal memory state (Cell State) and multiple gates. It can handle long-term dependencies.

[0209] (3) Input Gate: Controls whether the current input information (such as newly arrived data) should be added to the memory of the LSTM. Determines which parts of the current input data will be updated to the memory state of the network.

[0210] (4) Forget Gate: Controls how the model forgets previous information. Based on the current input and the previous hidden state, the forget gate determines which information should be discarded from the memory.

[0211] (5) Output Gate: Controls the output of the LSTM unit. It determines which information will ultimately affect the prediction of the model. Through the operation of the output gate, part of the information in the hidden state is passed to the next unit or affects the final prediction result.

[0212] (6) Internal State (Cell State): The Cell State is the long-term memory part in the LSTM unit. It is responsible for storing information on long-term dependencies. Through the control of the forget gate and the input gate, information is updated or passed on.

[0213] (7) Hidden State: The hidden state is the output of the network at each moment. It passes information to the LSTM unit at the next moment. As the time series progresses, the hidden state captures and stores the context information of the input sequence.

[0214] (8) Prediction Output: The final prediction output of the LSTM model. It can be a regression prediction (such as a health status score) or a classification prediction (such as whether a failure will occur). Usually, it is a value processed by the last layer of the neural network, representing the prediction result.

[0215] Through these structures and components, both XGBoost and LSTM can effectively handle complex data patterns and provide efficient prediction capabilities for aspects such as server health status and failure prediction.

[0216] Historical server operation data (time-series data: CPU load and temperature in the past 24 hours), LSTM cells (Long Short-Term Memory networks), hidden state updates (storing long-term dependency information), predicting the server health status for the next N hours. The technical implementation includes:

[0217] 1. Data input: various server health metrics within the past 24 hours (such as CPU load, fan speed, etc.);

[0218] 2. LSTM network: processing long-term dependency data through memory cells;

[0219] 3. Predicting the future server status: such as the temperature trend within the next 6 hours and whether it will exceed the threshold.

[0220] Furthermore, the logistic regression model includes a configuration data input sub-module, a linear transformation sub-module, and a Sigmoid function processing sub-module connected in sequence;

[0221] The configuration data input sub-module is used to input the hardware information of the server, and the hardware information at least includes the CPU model and BIOS version; the Sigmoid function processing sub-module is used to output the upgrade compatibility probability.

[0222] Specifically, the structure of the logistic regression model is: input features (data), weight calculation, linear combination (Σw*x), Sigmoid activation function, output prediction (probability), where,

[0223] Input features (data): The input data are the features used by the model for prediction, which can be numerical or encoded categorical features. For example:

[0224] CPU model: Different CPU models may have compatibility differences, especially in terms of supported functions and performance.

[0225] BIOS version: Different BIOS versions may affect hardware support or operating system compatibility.

[0226] RAID controller: Different RAID controller models may have different driver programs and supported storage configurations.

[0227] Memory size, version, and type: Different memory versions and sizes may affect system performance and stability.

[0228] Hard drive type and version: Solid State Drives (SSDs) and Hard Disk Drives (HDDs) may have different compatibility issues after firmware upgrades.

[0229] Feature: The input features are weighted according to weights.

[0230] Weight Calculation:

[0231] Function: Each input feature has a related weight, and the weight determines the importance of the feature in the model.

[0232] Feature: The weights are learned and optimized through the training process.

[0233] Linear combination (Σw*x):

[0234] Function: All input features form a linear model through weighted summation (Σw*x).

[0235] Feature: This step is the core of logistic regression, where the weights are multiplied by the input features and summed up.

[0236] Sigmoid activation function:

[0237] Function: The Sigmoid function converts the result of the linear combination into a probability value between 0 and 1.

[0238] Feature: The formula of the Sigmoid function is: , where z is the result of the linear combination.

[0239] Output prediction (probability):

[0240] Function: The output is a probability value for classification. Usually, a probability value greater than a certain threshold (such as 0.5) is classified as the positive class, otherwise as the negative class.

[0241] Feature: The output of logistic regression is a probability value, which can be further converted into a classification label.

[0242] Furthermore, a support vector machine (SVM) model is used to predict whether a new firmware is suitable for a specific server hardware. The SVM model includes a data input sub-module, a feature mapping sub-module, a hyperplane classification sub-module, and a result output sub-module connected in sequence;

[0243] The data input sub-module is used to input the hardware parameters of the server. The hardware parameters of the server at least include the BIOS version, CPU model, and storage controller model. The feature mapping sub-module is used to extract the server hardware parameters and map them to a high-dimensional space. The hyperplane classification sub-module is used to calculate the optimal decision boundary through SVM. The result output sub-module is used to output the compatible or incompatible result.

[0244] Specifically, SVM (Support Vector Machine) is suitable for small-sample classification problems and can be used to predict whether a new firmware is applicable to specific server hardware.

[0245] Specifically, the processing process of SVM (Support Vector Machine) sequentially includes steps of inputting data (server hardware parameters: BIOS version, CPU, storage controller), feature mapping (high-dimensional feature space), hyperplane classification (SVM calculates the optimal decision boundary), and outputting the upgrade result (compatible / incompatible). The technical implementation is as follows:

[0246] 1. Feature extraction (mapping server hardware parameters to a high-dimensional space);

[0247] 2. Training the SVM classifier to find the optimal decision boundary;

[0248] 3. Predicting whether the firmware upgrade is compatible (outputting compatible / incompatible labels).

[0249] Furthermore, the random forest model is used for multi-factor health assessment of the server. The random forest model includes a data input sub-module, a decision tree sub-module, and a voting decision sub-module connected in sequence. The data input sub-module is used to input server status data. The decision tree sub-module includes multiple decision trees, and each decision tree is set in parallel. The voting decision sub-module makes a voting decision based on the output of the decision tree sub-module and outputs a health score.

[0250] Specifically, Random Forest consists of multiple decision trees and is suitable for multi-factor health assessment of the server.

[0251] The processing process of Random Forest sequentially includes steps of inputting server status data (CPU, memory, disk, etc.), parallel training of multiple decision trees, a voting mechanism (taking the majority decision), and outputting the server health score (0 - 100). The technical implementation is as follows:

[0252] 1. Collecting server hardware status data;

[0253] 2. Conducting multi-factor health assessment using multiple decision trees (training different trees with different data subsets);

[0254] 3. Making a voting decision and outputting the health score (such as 80 / 100, indicating good status).

[0255] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. As long as it achieves the technical effects of the present invention by the same means, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the scope of protection of the present disclosure and shall fall within the scope of protection of the present invention. Within the scope of protection of the present invention, various different modifications and variations may be made to its technical solutions and / or implementation manners.

Claims

1. A server configuration management system based on the Redfish protocol, characterized in that Including: API interface layer, which is used to provide access interfaces externally; Redfish device layer, which is used to achieve remote management of the server through the Redfish protocol, and the Redfish device layer is connected to the API interface layer; Management control layer, which is used to comprehensively manage the configuration, faults, compatibility and firmware status of the multi-server cluster, and the management control layer is connected to the Redfish device layer.

2. The server configuration management system based on the Redfish protocol according to claim 1, wherein The API interface layer includes a unified API proxy module, an identity authentication and permission management module, and a log recording and access auditing module that are connected in sequence, The unified API proxy module is built based on the Python and Flask frameworks and encapsulates multiple source RESTful APIs to provide a unified interface externally; The identity authentication and permission management module performs identity authentication through JWT (JSON Web Token) and API access control based on role-based access control (RBAC); The log recording and access auditing module records behaviors such as API calls, user identities, and access times through the Elasticsearch full-text search engine and the Kibana analysis and visualization platform, and generates visual logs and security audit reports.

3. The server configuration management system based on the Redfish protocol according to claim 1, characterized in that, The Redfish device layer includes: a server power management module, a hardware status monitoring module, a RAID configuration management module, a BIOS remote configuration module, and a firmware update management module; The server power management module sends RESTful requests for operating the server power through the Redfish API, and the RESTful requests for operating the server power can remotely batch-operate the power-on, power-off, and restart of the server; The hardware status monitoring module collects the health status data of the server through the Redfish API and stores it in the Prometheus time series database, and performs visual display through the Grafana analysis and monitoring platform. The health status data of the server at least includes CPU, memory, disk, temperature, and fan status data; The RAID configuration management module configures the storage controller through the Redfish API, including remotely creating RAID 0, RAID 1, and RAID 5 to improve storage redundancy and data security; The BIOS remote configuration module improves the compatibility of the server by remotely adjusting the BIOS startup parameters; The firmware update management module is used to remotely batch-update the BIOS firmware and BMC firmware.

4. The server configuration management system based on the Redfish protocol according to claim 1, wherein The management control layer includes a server batch configuration module, a server fault prediction module, a compatibility analysis module, and a remote firmware management module; The server batch configuration module realizes batch initialization of the server BIOS and RAID configuration through the Ansible operation and maintenance tool and the Redfish API; The server failure prediction module collects server health data through the Redfish API, stores the data using the Prometheus time series database, performs visual analysis through the Grafana analysis and monitoring platform, and predicts the health status of the server through the first prediction model to evaluate the probability of failure risk; The first prediction model includes an XGBoost model or an LSTM model; The compatibility analysis module predicts whether system compatibility issues will occur after firmware upgrade or hardware replacement through the second prediction model, avoiding crashes or anomalies after upgrade; The second prediction model includes a logistic regression model, a support vector machine SVM model, or a random forest model; The remote firmware management module remotely updates the BIOS and BMC firmware in batches through the Redfish API.

5. The Redfish protocol-based server configuration management system according to claim 4, wherein The XGBoost model includes a data input sub-module, at least one decision tree sub-module, a weighted integration sub-module, and a prediction sub-module connected in sequence, The data input sub-module is used to collect operation data such as server CPU temperature, fan speed, and disk I / O; at least one decision tree sub-module is iteratively optimized by the Boosting method to gradually reduce errors; the prediction sub-module outputs a predicted value of the future failure risk percentage.

6. The Redfish protocol-based server configuration management system according to claim 4, wherein The LSTM model includes a historical data input sub-module, a long short-term memory network LSTM sub-module, and a prediction sub-module connected in sequence; The historical data input sub-module is used to input the CPU load and temperature data of the server within a preset time range in the past; the long short-term memory network LSTM sub-module processes long-term dependent data through the memory gate Memory Cell; the prediction sub-module predicts the health status of the server within a preset time range in the future.

7. The Redfish protocol-based server configuration management system according to claim 4, wherein The logistic regression model includes a configuration data input sub-module, a linear transformation sub-module, and a Sigmoid function processing sub-module connected in sequence; The configuration data input sub-module is used to input the hardware information of the server, and the hardware information includes at least the CPU model and the BIOS version; the Sigmoid function processing sub-module is used to output the upgrade compatibility probability.

8. The Redfish protocol-based server configuration management system according to claim 4, wherein The support vector machine SVM model is used to predict whether the new firmware is applicable to specific server hardware, and the support vector machine SVM model includes a data input sub-module, a feature mapping sub-module, a hyperplane classification sub-module, and a result output sub-module connected in sequence; The data input sub-module is used to input the hardware parameters of the server, and the hardware parameters of the server at least include the BIOS version, CPU model, and storage controller model; the feature mapping sub-module is used to extract the server hardware parameters and map them to a high-dimensional space; the hyperplane classification sub-module is used to calculate the optimal decision boundary through SVM; the result output sub-module is used to output the compatible or incompatible results.

9. The Redfish protocol-based server configuration management system according to claim 4, wherein The random forest model is used for multi-factor health assessment of the server. The random forest model includes a data input sub-module, a decision tree sub-module, and a voting decision sub-module connected in sequence; the data input sub-module is used to input the server status data; the decision tree sub-module includes multiple decision trees, and each decision tree is set in parallel; the voting decision sub-module makes a voting decision according to the output of the decision tree sub-module and outputs a health score.

Citation Information

Patent Citations

  • Cluster remote management method and system

    CN116827757A

  • Intelligent management method and system based on server cluster

    CN119473803A

  • Automatic hard disk testing method and device and readable storage medium

    CN119883851A

Cited By

  • Monitoring method and system based on Redfish protocol, electronic equipment and storage medium

    CN122152644A