Dynamic load-oriented intelligent identification heterogeneous server resource expansion method

By building a server resource expansion system with intelligent identification and dynamic adjustment, the problems of insufficient automation and limited scalability of traditional server management technologies in large-scale clusters are solved, efficient adaptation to dynamic loads and resource optimization are achieved, and the stability of the system and resource utilization efficiency are improved.

CN120849097APending Publication Date: 2025-10-28JIANGSU WANWEI AISI NETWORK INTELLIGENT IND INNOVATION CENT
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510879291.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional server management technologies lack automation in large-scale clusters, have limited scalability, low intelligence, and inefficient resource utilization. They are unable to adapt to dynamic load changes, resulting in increased O&M burdens, decreased service availability, and increased risk of downtime.

Method used

Build an intelligent heterogeneous server resource expansion system for dynamic loads, identify new hardware resource modules through the I2C interface, automatically configure resources, dynamically adjust the trigger threshold of resource management based on historical data and real-time load fluctuations, use time series prediction algorithms to predict future loads and adaptively adjust resource allocation to achieve load balancing.

Benefits of technology

It optimizes resource management efficiency in data center and cloud computing environments, improves the flexibility of hardware expansion and the stability of system operation, and is suitable for modern data center scenarios such as large-scale cloud computing service platforms and artificial intelligence computing clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849097A_ABST
    Figure CN120849097A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of server management and cloud computing, in particular to a dynamic load-oriented intelligent identification heterogeneous server resource expansion method. The method aims at optimizing the resource management efficiency, the hardware expansion flexibility and the overall operation stability of the system in a data center and a cloud computing environment by integrating an advanced intelligent identification technology, a real-time data analysis capability and a self-adaptive resource scheduling mechanism. The method is particularly suitable for modern data center scenes requiring high concurrent processing capability, dynamic load adaptability and high reliability, such as a large-scale cloud computing service platform, an enterprise-level distributed storage system and an artificial intelligence computing cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of server management and cloud computing technology, and more specifically, to a method for intelligently identifying and expanding heterogeneous server resources for dynamic load. Background Technology

[0002] With the rapid development of cloud computing, big data, and artificial intelligence technologies, data centers are placing increasingly stringent demands on server computing power, storage capacity, network transmission speed, and system reliability. Traditional server management technologies primarily rely on Baseboard Management Controllers (BMCs) to monitor hardware status and perform initial resource configuration, providing remote management functions through the IPMI (Intelligent Platform Management Interface) protocol. However, with the continuous expansion of server clusters, the diversification of hardware devices, and the dynamic changes in application workloads, traditional technologies have gradually revealed numerous shortcomings.

[0003] Insufficient automation: When traditional BMCs identify and configure new hardware modules, they usually rely on administrators to manually input parameters or execute scripts, making it difficult to achieve plug-and-play automated management. This inefficient operation significantly increases the maintenance burden, especially in large-scale clusters.

[0004] Limited scalability: Existing systems have limited support for hot-swappable hardware. The process of adding new hardware modules often requires pausing some services or restarting the system, resulting in decreased service availability and failing to meet the high availability requirements of cloud computing environments.

[0005] Low level of intelligence: Traditional monitoring methods often use static thresholds (such as alarms when CPU utilization exceeds 80%) for anomaly detection. These fixed rules cannot adapt to dynamic fluctuations in load, which may lead to frequent false alarms or missed alarms. In addition, the lack of predictive analysis of hardware operating status makes it impossible for the system to detect potential faults in advance, and fault response is often delayed, increasing the risk of downtime.

[0006] Low resource utilization efficiency: Traditional resource allocation strategies are usually adjusted based on the current state and lack the ability to predict future load trends, resulting in resource waste (such as idle computing modules) or insufficiency (such as performance bottlenecks during peak periods).

[0007] Against this backdrop, academia and industry have begun exploring the introduction of intelligent technologies into server management. For example, combining sensor data with machine learning algorithms can enable fault prediction, or automated scripts can simplify hardware configuration processes. However, existing solutions still have limitations: on the one hand, many methods remain at the passive monitoring level, lacking proactive optimization capabilities; on the other hand, some intelligent solutions rely on complex deep learning models, incurring high computational costs and demanding hardware requirements, making them difficult to deploy widely in resource-constrained server environments. Therefore, there is an urgent need for a server management and resource expansion system that is innovative, practical, and efficient to address the diverse challenges of modern data centers. Summary of the Invention

[0008] The purpose of this invention is to provide a method for intelligently identifying and expanding heterogeneous server resources in response to dynamic loads, in order to solve at least one technical problem existing in the prior art.

[0009] Technical solution: A method for intelligently identifying and expanding heterogeneous server resources for dynamic load analysis, comprising: S1: Construct an intelligent heterogeneous server resource expansion system for dynamic load analysis, the system including servers. Nodes, hardware resource expansion modules, intelligent management modules, and management devices; S2: The intelligent management module identifies newly added hardware resource modules through the I2C interface and automatically configures the resources; S3: The intelligent management module dynamically adjusts the trigger thresholds for resource management based on historical data and real-time load fluctuations; S4: The intelligent management module uses time series forecasting algorithms to predict future loads and adaptively adjusts resource allocation based on the forecast results to achieve load balancing.

[0010] According to a further improvement of the present invention, the dynamic threshold adjustment in step S3 includes the following steps: S3.1: Use a sliding window to analyze the indicator data over the past 30 minutes and calculate the mean and standard deviation; S3.2: Adjust the regulation coefficient according to load fluctuations to generate a dynamic threshold; S3.3: Compare real-time data with dynamic thresholds to detect abnormal states and record abnormal events to the event log.

[0011] According to a further improvement of the present invention, the predictive load balancing uses an exponential smoothing method to predict the load trend for the next 5 to 10 minutes, and adjusts the power distribution or enables the standby computing module based on the predicted value.

[0012] According to a further improvement of the present invention, the intelligent identification includes scanning the expansion slot through the I2C interface, reading the EEPROM information of the hardware module, extracting the manufacturer ID, product ID and specification parameters, and verifying compatibility with the server node.

[0013] According to a further improvement of the present invention, the automatic configuration includes allocating PCIe addresses, configuring interrupt request (IRQ) and direct memory access (DMA) channels, and loading the corresponding drivers and initializing hardware modules.

[0014] According to a further improvement of the present invention, the management device provides a visual interface through OpenBMC to display real-time monitoring data, dynamic threshold curves and predicted load trends, and supports remote configuration and alarm management.

[0015] According to a further improvement of the present invention, the dynamic threshold adjustment further includes an optimization step, which adjusts the adjustment coefficient by calculating the false alarm rate and the false negative rate to improve the accuracy of the threshold.

[0016] According to a further improvement of the present invention, the predictive load balancing includes verifying the prediction error, and if the error exceeds a preset threshold, switching to a conservative strategy to ensure system stability.

[0017] Beneficial Effects: By integrating advanced intelligent recognition technology, real-time data analysis capabilities, and adaptive resource scheduling mechanisms, this invention optimizes resource management efficiency, hardware expansion flexibility, and overall system stability in data center and cloud computing environments. This invention is particularly suitable for modern data center scenarios requiring high concurrency processing capabilities, dynamic load adaptability, and high reliability, such as large-scale cloud computing service platforms, enterprise-level distributed storage systems, and artificial intelligence computing clusters. Attached Figure Description

[0018] Figure 1 This is the system architecture diagram of the present invention.

[0019] Figure 2 This is a functional block diagram of the intelligent management module of the present invention.

[0020] Figure 3 This is a flowchart of the automatic configuration unit of the present invention. Detailed Implementation

[0021] Example 1 The server management and hardware resource expansion system of the present invention, based on intelligent identification, dynamic threshold adjustment, and predictive load balancing, is implemented through the following detailed steps, covering hardware initialization, resource identification, automated configuration, dynamic monitoring, predictive adjustment, and visual management. The specific implementation methods are described step-by-step below to ensure the operability of the technical solution.

[0022] 1. Hardware resource initialization and data acquisition Step 1.1: Deployment and Configuration of Monitoring Tools A Baseboard Management Controller (BMC) is deployed in the server system, using the ASPEED AST2600 chip, which supports the IPMI 2.0 protocol and provides a high-performance sensor interface and independent management network functionality. The management network is configured via a dedicated Network Interface Card (NIC), with the IP address set to 192.168.1.100 and port 623, and HTTPS enabled to ensure secure communication. During BMC initialization, the core components of the server node are automatically scanned, including the processor (e.g., Intel Xeon Platinum 8360Y, 36 cores, 2.4GHz base frequency), memory modules (e.g., 128GB DDR4-3200 ECC memory), storage devices (e.g., 1TB Samsung 970 EVO NVMe SSD and 10TB Western Digital Gold SATA HDD), and network interfaces (e.g., Mellanox ConnectX-5 dual-port 40GbE network card). The data acquisition frequency is set to every 10 seconds, acquiring data from sensors via the I2C bus and PCIe interface to ensure a balance between real-time performance and system performance. During configuration, verify the sensor status through the BMC's web interface or command-line tools (such as ipmitool). For example, execute "ipmitool sensor list" to check if all sensors are returning data normally. If a sensor is found to be unresponsive, log it to " / var / log / bmc / sensor_error.log" and trigger an alarm.

[0023] Step 1.2: Target Definition and Data Acquisition Define the data collection targets, covering multi-dimensional operational metrics to support subsequent analysis and prediction. Collected metrics include CPU utilization (expressed as a percentage, e.g., 50%, or in units of cores, e.g., 2.5 cores), memory usage (in GB, e.g., 16GB used / 32GB total capacity), storage I / O (read / write speeds, e.g., 500MB / s read, 300MB / s write), temperature (processor 45°C, hard drive 40°C, motherboard 38°C), network traffic (input 2Gbps, output 1.5Gbps), and power status (12V voltage, 250W power consumption). The data collection process is implemented via the BMC's I2C bus. For example, temperature sensor data is read using I2C address 0x50. The data capture frequency is set to 10 seconds via the configuration file ( / etc / bmc / config.yaml) to avoid excessive system resource consumption. The collected data is stored in the BMC's built-in time-series database (path: / var / lib / bmc / tsdb). Each record contains a timestamp (accurate to milliseconds, e.g., 2025-04-06 10:00:00.123) and a metric value. To ensure data integrity, the acquisition module performs validation after each capture, discarding invalid data (such as missing temperature fields).

[0024] Step 1.3: Data Validation and Persistent Storage The collected data undergoes integrity verification, detecting null values ​​(such as missing temperature fields) or outliers (such as CPU utilization > 100% or temperature > 100°C). Outlier data is logged to the BMC's System Event Log (SEL), for example, "2025-04-06 10:00:05 - Sensor Error: CPU Temp Invalid". Valid data is compressed using the BMC's built-in storage engine and stored at " / var / lib / bmc / tsdb" with a retention period of 30 days. To support long-term analysis, data is backed up daily to external storage (mount point: / mnt / backup) via the NFS protocol. Backup tasks are scheduled using cron (cron:0 0 * * *). If a backup fails, a retry mechanism is triggered (up to 3 times, 10-minute intervals), and a log is recorded to " / var / log / bmc / backup_error.log". Data verification also includes checking timestamp continuity; for example, if the timestamp interval is > 15 seconds, it is marked as data loss and the administrator is notified.

[0025] Step 1.4: Scalability Support for Data Acquisition It supports dynamically expanding the data acquisition metrics, such as monitoring fan speed (2000 RPM), power efficiency (80% efficiency, 250W input / 200W output), or ambient humidity (40%) via additional sensors. Users can add custom acquisition endpoints through the BMC's web interface, such as configuring " / sensors / fan_speed" to monitor fan status. The data format is consistent with standard metrics and stored in a time-series database. Extensibility relies on the BMC's plugin architecture; for example, a Python script (fan_sensor.py) can be used to load new sensor drivers and update acquisition rules ( / etc / bmc / sensors.conf). To verify the expansion effect, test acquisition is performed, such as "ipmitool sdr type Fan" to confirm normal fan data. If the acquisition of a new metric fails, an error log is recorded (e.g., "Fan Sensor Not Found") and an alarm is sent via the SNMP interface (trap: sensor_extension_failed).

[0026] 2. Intelligent identification of hardware resources Step 2.1: Hardware Detection The BMC periodically scans server expansion slots (such as PCIe x16 slots or SAS interfaces) via the I2C interface to detect hot-plug events, such as the insertion of a 4TB SAS hard drive (model: Samsung PM1643) or a 40GbE network module (model: MellanoxConnectX-5). The scan frequency is set to every 5 seconds. The BMC reads the module's EEPROM information via the I2C address (e.g., 0xA0), including the manufacturer ID ("Samsung"), product ID ("PM1643"), version number ("v1.2"), and serial number ("SN123456"). A hot-plug event triggers an interrupt signal (GPIO pin 10), which the BMC captures and logs ("2025-04-06 10:00:10 - NewDevice Detected: SAS Storage"). If multiple devices are detected inserted simultaneously, the BMC processes them one by one according to slot priority (e.g., PCIe slot 1 takes precedence over slot 2) to ensure system stability.

[0027] Step 2.2: Parameter Extraction and Compatibility Verification The EEPROM data is parsed to extract hardware specifications, such as storage module capacity (4TB), interface type (SAS 12Gb / s), maximum power consumption (15W), or network module bandwidth (40GbE, QSFP+ interface). Compatibility verification includes checking the PCIe version (Gen4 supported), power requirements (module power consumption < motherboard power supply limit of 500W), and firmware version (EEPROM version matches the support list in the BMC database). If incompatible, for example, a PCIe Gen3 module inserted into a Gen4 slot, an error log is generated (“PCIe Version Mismatch: Gen3 vs Gen4”) and an alarm is sent via the BMC's SNMP interface (trap: device_incompatible). Modules that pass verification are marked as “recognized” and stored in the BMC cache (path: / tmp / bmc / cache). To improve efficiency, the BMC maintains a device support database ( / etc / bmc / device_support.db), which contains compatibility information for common hardware (e.g., Samsung PM1643 supports Gen4).

[0028] Step 2.3: Storage and Sharing of Recognition Results The recognition results are formatted as JSON objects and stored in the BMC memory cache ( / tmp / bmc / cache / device.json). They are then transmitted to the server management device via the IPMI protocol (command: ipmitool raw 0x0a 0x01). The management device retrieves the recognition results via a REST API (GET / api / devices), supporting remote viewing and configuration. Data synchronization occurs every 30 seconds; if synchronization fails, a retry is triggered (up to 5 times, with 10-second intervals). To ensure data consistency, the BMC validates the JSON format before each synchronization. If an error is found, a log entry ("Invalid JSON Format") is recorded, and the JSON is regenerated.

[0029] 3. Automated resource allocation Step 3.1: Resource Allocation Based on the identification results, hardware resources are automatically allocated. Storage modules are assigned PCIe addresses (e.g., "0000:01:00.0"), and network modules are assigned "0000:02:00.0". Interrupt request (IRQ) allocation is based on the Linux kernel's interrupt table; for example, storage modules are assigned IRQ 10, and network modules are assigned IRQ 11, with conflicts checked in " / proc / interrupts". Direct memory access (DMA) channels are configured for high-efficiency data transfer; for example, storage modules are assigned DMA channel 0 with high bandwidth priority (qos: high). The allocation process is completed through the BMC firmware, triggering a PCIe bus rescan by executing the command "echo 1 > / sys / bus / pci / rescan". If allocation fails (e.g., IRQ conflict), the BMC automatically reassigns (e.g., switching to IRQ 12) and logs it ("IRQ Conflict Resolved").

[0030] Step 3.2: Driver Loading and Initialization The driver is loaded based on the module type. For example, the "mpt3sas" kernel module is loaded for the SAS storage module (modprobe mpt3sas), and "mlx5_core" is loaded for the 40GbE network module. The driver loading status is verified using "lsmod |grep mpt3sas". If it fails, it is logged ("Driver Load Failed: mpt3sas"). Initialization includes creating a logical volume for the storage module (lvcreate -L 4T -n data_vol vg0) or configuring an IP address for the network module (ipaddr add 10.0.0.1 / 24 dev eth1). After initialization, "dmesg | grep mpt3sas" is executed to check the driver logs to ensure that the module is running normally. To improve reliability, BMC maintains a driver version database ( / etc / bmc / driver_db). If an outdated driver version is detected, it is automatically downloaded from the server (URL: HYPERLINK "https: / / driverrepo.example.com" \h https: / / driverrepo.example.com).

[0031] Step 3.3: Configuration Verification and Error Handling Perform a hardware self-test (POST) and verify the module status using the BMC command "ipmitool selftest". For example, the storage module should return "online" and the network module should return "link up". If configuration fails (e.g., driver loading timeout, with a timeout threshold set to 10 seconds), automatically roll back the operation, unload the driver (modprobe -r mpt3sas), and release the PCIe address (echo 0 > / sys / bus / pci / devices / 0000:01:00.0 / remove). Error reports are generated by the BMC in the format "2025-04-06 10:00:15 - Configuration Error: Driver Timeout". If multiple failures occur (more than 3), the BMC will pause configuration and notify the administrator via email (smtp: HYPERLINK "mailto:alert@example.com" \h alert@example.com).

[0032] Step 3.4: Resource Status Update Update the system resource table, for example, add 4TB of storage to the storage pool ( / dev / vg0 / data_vol, status: active), and add the 40GbE network card to the network interface list (eth1, status: available). Notify the management device via the IPMI command (ipmitoolraw 0x0a 0x02) to update the global resource view, with the storage path " / var / lib / manager / resources.json". If the update fails, trigger a retry mechanism (maximum 3 times, 5-second intervals). To support large-scale clusters, BMC synchronizes resource status through a Redis cache (redis: / / 192.168.1.200:6379) to ensure consistency across multiple nodes.

[0033] 4. Dynamic threshold adjustment and operational status monitoring Step 4.1: Data Acquisition and Preprocessing Real-time data is acquired from Step 1, such as CPU utilization (50%), memory usage (16GB / 32GB), and temperature (45°C), to generate time-series data. The data is smoothed using a 5-second moving average; for example, CPU utilization over the past 5 seconds [40%, 45%, 50%, 55%, 60%] is smoothed to 52%. The smoothed data is stored in the BMC cache ( / tmp / bmc / data) for subsequent analysis. To reduce noise, BMC filters out transient fluctuations (e.g., single usage >90% but lasting <3 seconds), retaining only persistent data.

[0034] Step 4.2: Dynamic threshold calculation A sliding window (30 minutes, 180 data points) is used to analyze metrics, calculating the mean and standard deviation. For example, CPU utilization has a mean of 50% and a standard deviation of 10%. The dynamic threshold is set to the mean plus 1.5 times the standard deviation, i.e., 65%. If the load increases (utilization > 70% for 5 minutes), the adjustment factor is adjusted to 1.2, and the threshold drops to 62%; if the load decreases (utilization < 30%), the factor is adjusted to 2.0, and the threshold rises to 70%. The threshold is logged ( / var / log / bmc / threshold.log) and displayed via a web interface (URL: https: / / 192.168.1.100 / threshold). To improve accuracy, BMC updates the sliding window periodically (refreshing hourly) to ensure the threshold reflects the latest load patterns.

[0035] Step 4.3: Status Analysis and Anomaly Detection The system compares real-time data with dynamic thresholds. For example, if CPU utilization exceeds the threshold of 65% (70%), it is marked as high load, and a log entry "CPU Over Threshold: 70% > 65%" is generated. If multi-dimensional anomalies are detected (e.g., temperature > 50°C and storage I / O < 10MB / s), it is marked as "potential hard drive failure," and an alarm is sent via SNMP (trap: system_alert). Anomaly events are stored in SEL ( / var / log / bmc / sel.log), which can be queried by administrators via a web interface.

[0036] Step 4.4: Dynamic Threshold Optimization Assess the false alarm rate (no adjustment triggered after an alarm) and the missed alarm rate (no alarm but a fault occurs), for example, by counting the number of false alarms in the past 24 hours. If the false alarm rate > 10%, increase the adjustment coefficient to 1.8; if the missed alarm rate > 5%, decrease it to 1.3. Optimization results are stored in the configuration file ( / etc / bmc / threshold.conf) and synchronized to the management device via the REST API (POST / api / threshold / update). To support self-learning, BMC analyzes historical data weekly and automatically adjusts the initial coefficient value (e.g., from 1.5 to 1.6).

[0037] 5. Predictive load balancing and resource adjustment Step 5.1: Load Forecasting Future load is predicted using exponential smoothing, for example, CPU utilization over the past 5 minutes [50%, 55%, 60%, 65%, 70%], with a smoothing factor of 0.3, predicting a utilization of 65.45% for the next minute. The prediction period is set to 5-10 minutes, and the data is stored in the BMC cache ( / tmp / bmc / forecast). To verify accuracy, the mean squared error between the predicted and actual values ​​is calculated; if the error is less than 5%, the prediction is considered reliable. The prediction process is implemented using an embedded Python script (HYPERLINK "http: / / forecast.py"\h forecast.py), running in BMC's lightweight Python environment (Python 3.8). If the prediction fails (e.g., due to insufficient data), it reverts to the default value (current utilization).

[0038] Step 5.2: Adaptive Resource Adjustment If the predicted utilization (65.45%) exceeds the threshold (65%), increase power allocation via BMC (from 250W to 300W) or enable the backup GPU module (NVIDIA Tesla V100). Adjustment commands are executed via IPMI, such as "ipmitoolchassis power control up". If the predicted utilization is <30%, reduce power to 200W or disable the redundant module by executing "ipmitool raw 0x0c 0x02". Adjust the priority configuration to optimize storage I / O first (switch to high-speed SSD), then adjust computing resources (increase CPU frequency to 3.0GHz). Adjust the log storage to " / var / log / bmc / resource_adjust.log".

[0039] Step 5.3: Exception Handling and Fault Tolerance If the predicted temperature exceeds 75°C (based on a trend of rising from 45°C to 70°C over the past 10 minutes), automatically switch to the redundant storage module (4TB HDD, model: WD Gold) by executing "echo 1 > / sys / block / sdb / device / rescan". If the failure persists, trigger a reboot (ipmitool chassis power reset). Fault tolerance mechanisms include checking the status of the redundant module ( / proc / mdstat) to ensure a successful switchover. Abnormal events are notified to the administrator via email and SMS (smtp: HYPERLINK "mailto:alert@example.com" \h alert@example.com, Twilio API).

[0040] Step 5.4: Effect Verification and Feedback Verify the adjustment effect, for example, check if the CPU utilization has dropped to 50% and the temperature has dropped to 50°C. The results are obtained using the BMC command "ipmitool sensor". If the expected results are not achieved, record it as a bottleneck ("Resource AdjustmentFailed: CPU 60%"). Feedback is used to optimize the smoothing coefficient; for example, if the prediction is too high, it is reduced to 0.2, and the configuration file ( / etc / bmc / forecast.conf) is updated.

[0041] Example 2 The server management and automated hardware resource expansion system of the present invention includes the following core steps and components, which are broken down in detail below: Step 1: Hardware resource initialization and data acquisition By utilizing the server's built-in monitoring tools and sensors, real-time hardware operation data is collected, providing a comprehensive data foundation for subsequent intelligent identification, threshold adjustment, and load prediction. The data collection process is broken down into the following sub-steps: Step 1.1: Deployment and Configuration of Monitoring Tools Deploy a Baseboard Management Controller (BMC) in the server system, select hardware that supports the IPMI 2.0 protocol (such as the ASPEED AST2600 chip), and realize remote access function through an independent management network interface (such as a dedicated NIC with IP address configured as 192.168.1.100 and port 623).

[0042] Configure BMC's sensor scanning function to automatically identify core components in server nodes, including processors (such as Intel Xeon), memory modules (such as DDR4), storage devices (such as NVMe SSDs), and network interfaces (such as 40GbE network cards). Set the data acquisition frequency to every 10 seconds to ensure the real-time and continuous nature of the data, while avoiding excessively high acquisition frequencies that could burden system performance.

[0043] Step 1.2: Target Definition and Data Acquisition Define the data collection targets, covering the following key metrics: CPU utilization: expressed as a percentage (e.g., 50%) or in units of cores (e.g., 2.5 cores). Memory usage: in GB (e.g., 16GB used / 32GB total capacity); Storage I / O: Read and write speeds (e.g., 500MB / s read, 300MB / s write); Temperature: Real-time temperatures of the processor, hard drive, and motherboard (e.g., CPU 45°C, SSD 40°C). Network traffic: Input and output bandwidth (e.g., 2Gbps input, 1.5Gbps output). Power supply status: voltage (e.g., 12V), power consumption (e.g., 250W).

[0044] Data is acquired from hardware sensors via the I2C bus or PCIe interface and stored in time-series format, formally represented as: ; in, For time datasets, For timestamps (accurate to milliseconds, such as 2025-04-06 10:00:00.123). This refers to the corresponding indicator value.

[0045] Step 1.3: Data Validation and Persistent Storage The collected data undergoes integrity verification, such as checking for null or outlier values. If an anomaly is detected, it is recorded in the BMC event log and the corresponding data point is discarded. The data is compressed using the BMC's built-in storage engine, uses a time-series database, and has a data retention period of 30 days, supporting historical data analysis and model training.

[0046] Step 1.4: Scalability Support for Data Acquisition The data collection metrics can be dynamically expanded according to business needs. For example, additional sensors can be used to monitor fan speed (unit: RPM, such as 2000 RPM), power efficiency (unit: watts / hour), or ambient humidity (percentage, such as 40%).

[0047] It supports user-defined data collection rules, such as adding new data collection endpoints (e.g., " / sensors / fan_speed") via the BMC web interface, and incorporating the expanded data into subsequent analysis and prediction processes.

[0048] Step 2: Intelligent identification of hardware resources Based on the collected data, the system intelligently identifies newly added hardware resource modules, extracts their key parameters, and verifies compatibility to provide support for subsequent configuration.

[0049] Step 2.1: Hardware Detection The server's expansion slots (such as PCIe slots or SAS interfaces) are periodically scanned via the I2C interface to detect hot-plug events, such as the insertion of new storage or network modules. Information from the module identification unit (EEPROM) of the hardware modules is read, including the manufacturer ID (such as "Samsung"), product ID (such as "SSD-001"), version number (such as "v1.2"), and serial number (such as "SN123456").

[0050] Step 2.2: Parameter Extraction and Compatibility Verification Parse the EEPROM data to extract hardware specifications, such as the storage module's capacity (4TB), interface type (SAS 12Gb / s), maximum power consumption (15W), or the network module's bandwidth (40GbE). Verify hardware compatibility with server nodes, such as checking PCIe version (Gen4 vs Gen3) or power requirements. If incompatible, generate an error log (e.g., "PCIe version mismatch") and trigger an alert via the BMC.

[0051] Step 2.3: Storage and Sharing of Recognition Results Format the recognition results as structured data, such as a JSON object: Data is stored in the BMC's memory cache and shared to the server management device via the IPMI protocol, supporting subsequent automated configuration and remote viewing.

[0052] Step 3: Automated Resource Configuration Based on the identification results, hardware resources are automatically configured and seamlessly integrated into the system to ensure plug-and-play functionality.

[0053] Step 3.1: Resource Allocation Automatically allocate PCIe addresses, for example, assign "0000:01:00.0" to the storage module and "0000:02:00.0" to the network module. Configure interrupt requests (IRQs), such as assigning IRQ 10 to the storage module and IRQ 11 to the network module, ensuring no resource conflicts. Set up direct memory access (DMA) channels, for example, assigning DMA channel 0 to the storage module to support efficient data transfer.

[0054] Step 3.2: Driver Loading and Initialization Load the corresponding driver based on the module type. For example, load the Linux kernel module "mpt3sas" for the SAS storage module and "mlx5_core" for the 40GbE network module. Dynamically insert the driver using the kernel module manager (modprobe) and check the loading status (use the "lsmod" command to verify the module's existence). Perform hardware initialization, such as allocating a logical volume (LVM) for the storage module or configuring an IP address (e.g., 10.0.0.1) for the network module.

[0055] Step 3.3: Configuration Verification and Error Handling Perform a hardware self-test (Power-On Self-Test, POST) to confirm that the modules are operating normally. For example, the storage module returns an "online" status, and the network module returns "link up".

[0056] If the configuration fails (e.g., driver loading timeout), the automatic rollback operation (unloading the module and releasing the PCIe address) will be performed, and a detailed error report (e.g., "driver version mismatch") will be generated through the BMC.

[0057] Step 3.4: Resource Status Update Update the system resource table, for example, add the newly added 4TB storage to the available storage pool and mark it as "active"; add the 40GbE network card to the network interface list and mark it as "available".

[0058] The server management device is notified via the IPMI protocol to update the global resource view.

[0059] Step 4: Dynamic threshold adjustment and operational status monitoring (innovation point) The system's operating status is monitored in real time through an intelligent management module, and a dynamic threshold adjustment mechanism is introduced to more intelligently adapt to load changes and hardware status fluctuations.

[0060] Step 4.1: Data Acquisition and Preprocessing Obtain real-time data from step 1, such as CPU utilization, memory usage, and temperature, and generate a time series vector (e.g., a CPU utilization sequence for the past 10 minutes: [40%, 45%, 50%, 55%]).

[0061] The data is smoothed using a 5-second moving average algorithm to reduce transient noise interference. For example, the smoothed sequence is: [42%, 47%, 52%].

[0062] Step 4.2: Dynamic threshold calculation Based on historical data and real-time load fluctuations, monitoring thresholds are dynamically adjusted to avoid the limitations of traditional fixed thresholds.

[0063] Calculation method: A sliding window (window size: approximately 180 data points over the past 30 minutes) is used to analyze the statistical characteristics of the metrics, including the mean (e.g., mean CPU utilization of 50%) and standard deviation (e.g., 10%). ; The dynamic threshold is defined as: ; in, The adjustment factor (default value is 1.5, can be adjusted through the management interface) adapts to load fluctuations. In high-load scenarios (such as CPU utilization >70% for 5 minutes), the sensitivity is reduced to 1.2. In low-load scenarios (such as CPU utilization <30%), the scale is increased to 2.0 to reduce false alarms.

[0064] Example: If the average CPU utilization within the sliding window is 50%, the standard deviation is 10%, and the load is normal, then the threshold is 1.5, and the dynamic threshold is 50% + 1.5 × 10% = 65%. If the load increases, the threshold is 1.2, and the threshold is adjusted to 62%.

[0065] Compared to traditional fixed thresholds (such as a constant 80%), dynamic thresholds can more accurately reflect the system status and avoid frequent alarms or missing critical anomalies.

[0066] Step 4.3: Status Analysis and Anomaly Detection Compare real-time data with dynamic thresholds. For example, if the current CPU utilization is 70% and the dynamic threshold is 65%, mark it as "high load state". Detect multi-dimensional anomalies, such as temperature > dynamic threshold (assuming 50°C) and storage I / O < 10MB / s, mark it as "potential hard drive failure". Output anomaly event logs, such as "2025-04-06 10:05:00 - CPU utilization exceeds threshold 65%, current value 70%".

[0067] Step 4.4: Dynamic Threshold Optimization Regularly evaluate the effectiveness of the threshold, for example, by calculating the false alarm rate (no adjustment triggered after an alarm) and the missed alarm rate (no alarm but a fault occurred) over the past 24 hours. The specific calculation method is as follows: ; If the false positive rate is greater than 10%, it will automatically increase (e.g., from 1.5 to 1.8); if the false negative rate is greater than 5%, it will decrease (e.g., to 1.3), thus achieving self-learning optimization.

[0068] Step 5: Predictive Load Balancing and Resource Adjustment (Innovation) Based on predictive models and dynamic thresholds, resource allocation is adjusted in advance to achieve load balancing and optimize performance and energy consumption.

[0069] Step 5.1: Load Forecasting A lightweight time series forecasting algorithm (exponential smoothing) is used to predict future load changes and anticipate resource demands in advance. The calculation method is as follows: ; in, This is the current predicted value. This is the predicted value from the previous period. This is the current actual value. This is the smoothing coefficient (default 0.3, which can be adjusted through historical data error analysis).

[0070] Example: If the CPU usage over the past 5 minutes was 50%, 55%, 60%, 65%, and 70% respectively, and = 0.3, then: Prediction for the 1st minute: 50%; Prediction for the 2nd minute: 50%×0.7+55%×0.3=51.5%; Prediction for the 5th minute: 63.5%×0.7+70%×0.3=65.45%. Prediction for the 6th minute: ≈ 65.45%.

[0071] Prediction period: 5-10 minutes in the future. The time window can be configured according to business needs (such as shortening to 5 minutes in high-concurrency scenarios).

[0072] Verify the accuracy of the prediction: Calculate the mean square error (MSE) between the predicted value and the actual value. If the MSE is less than 5%, the prediction is considered reliable.

[0073] Step 5.2: Adaptive Resource Adjustment Based on the forecast results and dynamic thresholds, adjust resource allocation in advance: If the predicted CPU utilization is >65% (dynamic threshold), increase the power distribution via BMC (e.g., from 250W to 300W) or enable a standby computing module (e.g., NVIDIA Tesla GPU).

[0074] If the predicted load drops to <30%, reduce the power allocation (e.g., reduce to 200W) or shut down redundant modules to save energy.

[0075] The adjustment strategy supports priority configuration, such as prioritizing the increase of storage I / O capabilities (switching to high-speed SSDs) and then adjusting computing resources (increasing CPU frequency).

[0076] Step 5.3: Exception Handling and Fault Tolerance If a failure risk is predicted, such as a continuous rise in temperature (from 45°C to 70°C in the past 10 minutes, and predicted to exceed 75°C in the next 5 minutes), fault-tolerant operations are automatically executed: Switch to redundant storage modules (e.g., switch from a failed SSD to a spare HDD). Restore services by sending a reboot command via IPMI (ipmitool chassis power reset).

[0077] If the prediction accuracy is insufficient (MSE>10%), automatically switch to a conservative strategy (such as fixed threshold mode) to ensure system stability.

[0078] Step 5.4: Effect Verification and Feedback Monitor the adjusted status, such as CPU utilization dropping from 70% to 50% and temperature dropping from 70°C to 55°C, to verify the load balancing and cooling effects. Feed the adjustment results back to the prediction model to optimize the smoothing coefficient (e.g., reduce it to 0.2 if the prediction is too high).

[0079] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A method for intelligently identifying and expanding heterogeneous server resources for dynamic load analysis, characterized in that, include: S1: Construct an intelligent heterogeneous server resource expansion system for dynamic load analysis, the system including servers. Nodes, hardware resource expansion modules, intelligent management modules, and management devices; S2: The intelligent management module identifies newly added hardware resource modules through the I2C interface and automatically configures the resources; S3: The intelligent management module dynamically adjusts the trigger thresholds for resource management based on historical data and real-time load fluctuations; S4: The intelligent management module uses time series forecasting algorithms to predict future loads and adaptively adjusts resource allocation based on the forecast results to achieve load balancing.

2. The method for intelligent identification and expansion of heterogeneous server resources for dynamic load analysis according to claim 1, characterized in that, The dynamic threshold adjustment in step S3 includes the following steps: S3.1: Use a sliding window to analyze the indicator data over the past 30 minutes and calculate the mean and standard deviation; S3.2: Adjust the regulation coefficient according to load fluctuations to generate a dynamic threshold; S3.3: Compare real-time data with dynamic thresholds to detect abnormal states and record abnormal events to the event log.

3. The method for intelligent identification and expansion of heterogeneous server resources for dynamic load analysis according to claim 1, characterized in that, The predictive load balancing uses exponential smoothing to predict load trends over the next 5 to 10 minutes and adjusts power distribution or enables backup computing modules based on the predicted values.

4. The method for intelligent identification and expansion of heterogeneous server resources for dynamic load analysis according to claim 1, characterized in that, The intelligent identification includes scanning the expansion slot through the I2C interface, reading the EEPROM information of the hardware module, extracting the manufacturer ID, product ID and specification parameters, and verifying compatibility with the server node.

5. The method for intelligent identification and expansion of heterogeneous server resources for dynamic load analysis according to claim 1, characterized in that, The automatic configuration includes allocating PCIe addresses, configuring interrupt request (IRQ) and direct memory access (DMA) channels, and loading the corresponding drivers and initializing the hardware modules.

6. The method for intelligent identification and expansion of heterogeneous server resources for dynamic load analysis according to claim 1, characterized in that, The management device provides a visual interface through OpenBMC, displaying real-time monitoring data, dynamic threshold curves, and predicted load trends, and supports remote configuration and alarm management.

7. The method for intelligent identification and expansion of heterogeneous server resources for dynamic load analysis according to claim 2, characterized in that, The dynamic threshold adjustment further includes an optimization step, which adjusts the adjustment coefficient by calculating the false alarm rate and the false negative rate to improve the accuracy of the threshold.

8. The method for intelligent identification and expansion of heterogeneous server resources for dynamic load analysis according to claim 3, characterized in that, The predictive load balancing includes verifying the prediction error; if the error exceeds a preset threshold, it switches to a conservative strategy to ensure system stability.

Citation Information

Cited By

  • Module identification method and device based on hybrid detection mechanism and storage medium

    CN121387676A

  • Module identification method and device based on hybrid detection mechanism, and storage medium

    CN121387676B