A method and device for perceiving three-dimensional features of server load based on BMC

CN122064562BActive Publication Date: 2026-08-21ANQING (TIANJIN) COMPUTER CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610524796.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-20
Publication Date
2026-08-21
Estimated Expiration
2046-04-20

AI Technical Summary

Technical Problem

[0005]有鉴于此,本申请提供一种基于BMC的服务器负载三维特征感知方法和装置,可解决现有BMC负载感知方法因感知维度单一、判定逻辑零散,无法全面准确地反映服务器负载的多维动态特征,导致后续散热和功耗调控策略与真实负载状态不匹配,影响服务器稳定性与能效比的问题

Benefits of technology

[0021]本申请提供的基于BMC的服务器负载三维特征感知方法和装置,通过BMC采集CPU、GPU及IO带宽等多维度硬件负载参数,并基于这些参数构建能够量化资源竞争关系与整体计算负荷的专属指标,进而从负载类型、负载水平、斜率特征三个维度对服务器负载状态进行全面识别,最终将识别结果编码为标准化感知标签输出,整体上实现了服务器负载的多维度、全特征、高精度感知。这样的三维感知机制使得负载感知结果更加全面、精准且具备统一的数据格式,能够为下游散热与功耗调控系统提供可靠、适配的决策依据,有效提升了负载感知对硬件调控的指导价值。具体的,通过BMC采集CPU负载率、GPU负载率和IO带宽占比三类实时工作参数,依托BMC独立于服务器主系统运行、不占用计算资源的特性,可在不影响主系统性能的前提下,实现核心硬件负载参数的实时采集,保证了参数获取的及时性与采集过程的无干扰性。基于采集到的参数计算得到计算资源竞争比、IO资源竞争比、总资源竞争比以及计算负载参数,这些量化指标能够客观反映服务器内部计算资源与IO资源之间的竞争关系以及整体计算负荷水平,为后续的负载特征识别奠定了精确的量化基础。在此基础上,依次识别服务器的负载类型、负载水平和斜率特征,从资源瓶颈方向、负载绝对强度、动态变化趋势三个维度完整拆解服务器负载状态,改变了传统单一维度感知的局限性,使得感知结果能够真实反映服务器的实际运行状况。最后,将三类识别结果按照预设编码规则生成标准化感知标签并输出,实现了多维度感知结果的统一结构化封装,便于下游系统快速读取和解析,实现了负载感知与散热、功耗调控系统之间的高效对接。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064562B_ABST
    Figure CN122064562B_ABST
Patent Text Reader

Abstract

The application provides a BMC-based server load three-dimensional feature perception method and device, and belongs to the field of intelligent management of servers. The method comprises the following steps: collecting real-time working parameters through a BMC; calculating a computing resource competition ratio and an IO resource competition ratio based on the real-time working parameters; calculating a total resource competition ratio based on the computing resource competition ratio and the IO resource competition ratio, and calculating a computing load parameter based on a CPU load rate and a GPU load rate; identifying a server load type based on the computing load parameter and the computing resource competition ratio, the IO resource competition ratio and the total resource competition ratio; identifying a server load level; identifying a slope feature of the server load; generating a standardized perception tag according to a preset coding rule based on the identification results of the load type, the load level and the slope feature, and outputting a server load three-dimensional feature perception result. The method and device provided by the application can solve the problems of incomplete server load perception, lagging regulation and control and low precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server intelligent management technology, and in particular to a method and apparatus for server load three-dimensional feature perception based on BMC. Background Technology

[0002] Server energy consumption has become a core cost component of data centers. Server load exhibits multi-dimensional, highly volatile, and scenario-specific characteristics: compute-intensive tasks prioritize computing power stability, I / O-intensive tasks require ensuring data transmission without packet loss, and high-load balancing tasks need to prevent overheating while avoiding energy waste. The Baseboard Management Controller (BMC), as a hardware management core operating independently of the main server system, directly determines the effectiveness of heat dissipation and power regulation through its load sensing accuracy, making it a crucial element in ensuring server stability and energy efficiency.

[0003] Current server load monitoring technologies based on load balancing (BMC) primarily focus on collecting single-dimensional hardware parameters and applying simple thresholds. Mainstream solutions only collect parameters such as CPU load rate, GPU load rate, or I / O bandwidth percentage. While some solutions collect multiple parameters simultaneously, they only apply independent thresholds to each parameter without quantitatively analyzing the relationships between them. Furthermore, existing technologies often output discrete values ​​or simple high / low load indicators without a standardized encoding format. They also rely on static judgments based on single data collections, failing to capture dynamic load trends and making them susceptible to distortion caused by momentary fluctuations.

[0004] Therefore, on the one hand, a single-dimensional perception method cannot quantify the competitive relationship between computing resources and I / O resources, making it difficult to identify load types and locate resource bottlenecks. There is also no unified quantitative indicator to classify load intensity, resulting in a lack of comprehensiveness and refinement in the perception results. On the other hand, the lack of a standardized output format increases the integration cost of downstream heat dissipation and power consumption control systems. Furthermore, the lack of a static judgment method for dynamic trend perception can easily lead to frequent adjustments in the control system, causing problems such as lagging, over- or under-control, which not only affects the stability of server operation but also easily leads to energy waste and shortens the lifespan of hardware. Therefore, there is an urgent need for a BMC load perception solution that can achieve multi-dimensional, full-feature, and high-precision perception of server load. Summary of the Invention

[0005] In view of this, this application provides a method and apparatus for three-dimensional feature perception of server load based on BMC, which can solve the problem that the existing BMC load perception method cannot fully and accurately reflect the multi-dimensional dynamic features of server load due to the single perception dimension and fragmented judgment logic, resulting in the mismatch between subsequent heat dissipation and power consumption control strategies and the actual load state, thus affecting the stability and energy efficiency of the server.

[0006] Specifically, this application is implemented through the following technical solution:

[0007] The first aspect of this application provides a method for perceiving three-dimensional features of server load based on BMC, the method comprising:

[0008] By collecting the server's CPU load rate, GPU load rate, and IO bandwidth ratio through BMC, the server's real-time operating parameters can be obtained.

[0009] The computational resource contention ratio and I / O resource contention ratio are calculated based on the real-time operating parameters; the total resource contention ratio is calculated based on the computational resource contention ratio and I / O resource contention ratio; and the computational load parameters are calculated based on the CPU load rate and GPU load rate.

[0010] The server load type is identified based on the computing load parameters and the computing resource contention ratio, IO resource contention ratio, and total resource contention ratio.

[0011] Identify the server load level based on the calculated load parameters;

[0012] The slope characteristics of the server load are identified based on the slope direction and absolute value of the calculated load parameters.

[0013] The identification results of the load type, load level, and slope features are used to generate standardized perception labels according to preset coding rules, and the server load three-dimensional feature perception results are output based on the standardized perception labels.

[0014] A second aspect of this application provides a server load three-dimensional feature perception device based on BMC, the device comprising an acquisition module, a calculation module, a recognition module and a processing module;

[0015] The acquisition module is used to acquire the server's CPU load rate, GPU load rate, and IO bandwidth ratio through the BMC to obtain the server's real-time operating parameters.

[0016] The computing module is used to calculate the computing resource contention ratio and the IO resource contention ratio based on the real-time working parameters; and to calculate the total resource contention ratio based on the computing resource contention ratio and the IO resource contention ratio, and to calculate the computing load parameters based on the CPU load rate and the GPU load rate.

[0017] The identification module is used to identify the server load type based on the computing load parameters and the computing resource contention ratio, IO resource contention ratio, and total resource contention ratio.

[0018] The identification module is also used to identify the server load level based on the calculated load parameters;

[0019] The identification module is also used to identify the slope characteristics of the server load based on the slope direction and absolute value of the slope of the calculated load parameters.

[0020] The processing module is used to generate standardized perception labels from the identification results of the load type, load level, and slope features according to preset coding rules, and output the three-dimensional feature perception results of the server load based on the standardized perception labels.

[0021] This application provides a server load 3D feature perception method and device based on BMC. It collects multi-dimensional hardware load parameters such as CPU, GPU, and I / O bandwidth using BMC, and constructs specific indicators based on these parameters to quantify resource competition and overall computing load. This allows for comprehensive identification of the server load status from three dimensions: load type, load level, and slope characteristics. Finally, the identification results are encoded into standardized perception labels for output, achieving multi-dimensional, full-feature, and high-precision perception of server load. This 3D perception mechanism makes the load perception results more comprehensive, accurate, and has a unified data format, providing reliable and adaptable decision-making basis for downstream heat dissipation and power consumption control systems, effectively enhancing the guiding value of load perception for hardware control. Specifically, by collecting three types of real-time operating parameters—CPU load rate, GPU load rate, and I / O bandwidth ratio—using BMC, and leveraging the characteristic that BMC runs independently of the main server system and does not occupy computing resources, real-time acquisition of core hardware load parameters can be achieved without affecting the performance of the main system, ensuring the timeliness of parameter acquisition and the non-interference of the acquisition process. Based on the collected parameters, the computational resource contention ratio, I / O resource contention ratio, total resource contention ratio, and computational load parameters are calculated. These quantitative indicators objectively reflect the competitive relationship between computational and I / O resources within the server and the overall computational load level, laying a precise quantitative foundation for subsequent load feature identification. Building upon this, the server's load type, load level, and slope characteristics are identified sequentially. The server load status is comprehensively analyzed from three dimensions: resource bottleneck direction, absolute load intensity, and dynamic change trend. This overcomes the limitations of traditional single-dimensional perception, enabling the perception results to truly reflect the server's actual operating status. Finally, the three types of identification results are generated into standardized perception tags according to preset encoding rules and output, achieving a unified structured encapsulation of multi-dimensional perception results. This facilitates rapid reading and parsing by downstream systems and enables efficient integration between load perception and heat dissipation and power consumption control systems. Attached Figure Description

[0022] Figure 1 A flowchart of an embodiment of a server load 3D feature perception method based on BMC provided in this application;

[0023] Figure 2This is a schematic diagram of a second embodiment of a server load three-dimensional feature sensing device based on BMC provided in this application. Detailed Implementation

[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.

[0025] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0027] The following specific embodiments are given to illustrate the technical solution of this application in detail.

[0028] Example 1

[0029] Figure 1 This is a flowchart of an embodiment of a server load 3D feature perception method based on BMC provided in this application. Please refer to... Figure 1 The method provided in this embodiment may include:

[0030] S101. Collect the server's CPU load rate, GPU load rate, and IO bandwidth ratio through BMC to obtain the server's real-time operating parameters.

[0031] It should be noted that the collection of server load-related parameters in this application is based on the BMC, which stands for Baseboard Management Controller. The BMC can operate independently of the main server system and has the ability to monitor and collect data on the server hardware status in real time. By completing the parameter collection through the BMC, it is not necessary to occupy the server's CPU, GPU and other computing resources, which can ensure that the collection process does not affect the performance of the main system and at the same time guarantee the real-time performance of the data.

[0032] By collecting the server's CPU load rate, GPU load rate, and IO bandwidth ratio through BMC, the server's real-time operating parameters can be obtained. These three types of parameters correspond to the load status of the server's general computing resources, graphics / parallel computing resources, and data transmission resources, respectively, which can comprehensively reflect the overall hardware operation of the server.

[0033] During the data collection process, BMC matches the corresponding collection interface, collection cycle, and applicable scenarios based on the hardware attributes of different characteristic data to achieve targeted parameter collection. Specifically, CPU load rate is collected through an Intel PECI+APML and AMD SMBus+APML compatible interface, with a collection cycle set to, for example, 10ms. This interface is compatible with both Intel and AMD, the two major mainstream server CPU architectures, covering the full-scenario needs of pure CPU servers and CPU+GPU hybrid architecture servers. The short 10ms collection cycle accurately captures instantaneous fluctuations in CPU load, adapting to various business scenarios sensitive to changes in computing resources and ensuring real-time load monitoring. For GPU load rate collection, the SMBus BBI interface is used for enterprise-level GPUs, while a Host-side + IPMI backhaul method is used for ordinary GPUs, with a collection cycle also set to 10ms. This collection logic covers different levels of GPU hardware and adapts to GPU-intensive business scenarios such as AI inference and graphics rendering. The consistent 10ms collection cycle with CPU ensures the time synchronization of computing resource load monitoring, facilitating accurate calculation of subsequent computing resource contention ratios. In addition, the IO bandwidth ratio is collected through the network card / RAID controller management interface, with the collection period also set to 10ms. This interface directly connects to the core data transmission hardware of the server, which can accurately reflect the bandwidth usage of the IO link and is suitable for IO-intensive business scenarios such as databases and storage servers. The 10ms collection period can capture changes in IO bandwidth in a timely manner and keep the time dimension aligned with CPU and GPU load data.

[0034] It should be noted that the specific values ​​of the collection cycle for CPU load rate, GPU load rate, and IO bandwidth ratio can be flexibly selected according to actual business needs, server hardware performance, and BMC communication load. For example, it can be adjusted to 5ms, 20ms, or 50ms depending on the scenario requirements. However, it is essential to ensure that the collection cycles of the three parameters remain synchronized. If the collection cycles are inconsistent, the collection time of the three parameters will be misaligned, making it impossible to obtain complete CPU, GPU, and IO load data sets at the same time dimension. This will lead to distortion in the calculation of resource competition relationships and will fail to accurately reflect the true load status of the server.

[0035] After BMC continuously collects the three types of parameters according to the above rules, the raw data has problems such as inconsistent format, misaligned collection time, and insufficient hardware compatibility. If it is directly used for subsequent calculations, it is easy to cause result deviation or process errors. Therefore, the raw collected data needs to be preprocessed. Specifically, the CPU load rate, GPU load rate, and IO bandwidth ratio are converted into an integer format between 0 and 100; timestamps are added to the CPU load rate and GPU load rate data, and the timestamps of the IO bandwidth ratio data are reused from adjacent load data that already have timestamps; when no GPU device is detected on the server, the GPU load rate collection process is automatically skipped.

[0036] It should be noted that the raw collected data may be presented in various formats such as decimals, percentage strings, or raw hardware values, and the numerical range is not uniform. Converting it to an integer format of 0-100 can completely eliminate format differences and avoid subsequent calculation errors. At the same time, the integer format is more efficient in calculation and better suited to the lightweight computing capabilities of BMC, without adding extra load to BMC. In addition, millisecond-accurate collection timestamps are added to the CPU load rate and GPU load rate data, and the timestamps of adjacent CPU load rate and GPU load rate data with timestamps are reused for the IO bandwidth ratio data. This is because the collection interfaces and hardware interaction logic of the three types of parameters are different, which can easily lead to slight misalignments in collection time. The subsequent calculation of indicators such as resource contention ratio requires complete parameter data in the same time dimension as support. Timestamp reuse can achieve precise alignment of the three types of parameters, ensuring that complete CPU, GPU, and IO load data are available at the same timestamp. Among them, the timestamps of CPU load rate or GPU load rate are chosen to reuse the IO bandwidth ratio data because the collection interface of CPU load rate and GPU load rate has higher time accuracy.

[0037] It should also be noted that when no GPU device is detected on the server, the GPU load rate collection process is automatically skipped. Before starting the collection, BMC will first read the server hardware configuration information to complete the GPU hardware presence detection; if it detects that the server does not have a GPU configured, or the GPU is faulty or not enabled, the GPU load rate collection step will be automatically blocked, and only the CPU load rate and IO bandwidth ratio will be collected. This avoids collection errors caused by hardware absence, ensures the robustness of the collection process, reduces invalid collection operations, and improves the overall collection efficiency.

[0038] S102. Calculate the computing resource contention ratio and IO resource contention ratio based on the real-time working parameters; calculate the total resource contention ratio based on the computing resource contention ratio and IO resource contention ratio; and calculate the computing load parameters based on the CPU load rate and GPU load rate.

[0039] After obtaining real-time working parameters with a uniform format and time alignment, these working parameters can reflect the load status of different resources, but they are independent of each other and cannot directly reflect the competition relationship between computing resources and I / O resources inside the server. They are also difficult to quantify the resource bottleneck type of the current load. Therefore, it is necessary to further calculate indicators based on CPU load rate, GPU load rate and I / O bandwidth ratio to obtain quantitative indicators that can objectively reflect the resource occupancy relationship inside the server. Specifically, these include computing resource contention ratio, I / O resource contention ratio, total resource contention ratio, and computing load parameters that can comprehensively reflect the overall computing load of the server.

[0040] The computing resource contention ratio characterizes the relative scarcity of computing resources (CPU+GPU) compared to I / O resources in a server, reflecting that the current system load is more biased towards computing pressure. The I / O resource contention ratio characterizes the relative scarcity of I / O resources compared to computing resources in a server, reflecting that the current system load is more biased towards data transmission pressure. These two contention ratios effectively distinguish whether the server is currently experiencing a computing resource bottleneck, an I / O resource bottleneck, or a relatively balanced state between the two. Based on this, the total resource contention ratio is further calculated, which is the higher of the two ratios. Thus, using the more strained resource metric as a representative of the overall resource contention level, regardless of whether the current bottleneck is computing or I / O resources, the total resource contention ratio captures the more intense competition, allowing for a quick and intuitive assessment of whether the server exhibits a significant resource imbalance.

[0041] Specifically, the sum of the CPU load rate and GPU load rate, plus a zero-prevention baseline value, is used as the dividend, and the sum of the IO bandwidth ratio, plus the zero-prevention baseline value, is used as the divisor. The resulting division yields the computing resource contention ratio. Similarly, the sum of the IO bandwidth ratio, plus the zero-prevention baseline value, is used as the dividend, and the sum of the CPU load rate and GPU load rate, plus the zero-prevention baseline value, is used as the divisor. The resulting division yields the IO resource contention ratio.

[0042] Based on the above description, the computational resource contention ratio = (CPU load rate + GPU load rate + zero-prevention baseline) / (IO bandwidth ratio + zero-prevention baseline); the IO resource contention ratio = (IO bandwidth ratio + zero-prevention baseline) / (CPU load rate + GPU load rate + zero-prevention baseline). It can be seen that the calculation formulas for the computational resource contention ratio and the IO resource contention ratio are mathematically reciprocals. Furthermore, it should be noted that a computational resource contention ratio greater than 1 indicates that the load pressure on computational resources is higher than that on IO resources, and the server may currently be facing a computing power bottleneck; an IO resource contention ratio greater than 1 indicates that the load pressure on IO resources is higher than that on computational resources, and the server may currently be facing a data transmission bottleneck. These two ratios transform the originally isolated CPU, GPU, and IO load data into comparable indicators of resource contention intensity. Similarly, a larger total resource contention ratio indicates more intense resource competition and that the server is closer to resource saturation; a total resource contention ratio close to 1 indicates that the load on computational and IO resources is relatively balanced, with no obvious bottleneck.

[0043] It's important to note that a zero-division prevention baseline is introduced during the calculation of resource contention ratio and I / O resource contention ratio. This baseline is an extremely small positive number, such as 0.01. Its main function is to avoid division by zero errors when the denominator is zero, ensuring the rationality of the calculation process. During actual server operation, CPU load rate, GPU load rate, or I / O bandwidth percentage may momentarily drop to zero. Directly using zero as the divisor will lead to calculation anomalies or program crashes. By introducing the zero-division prevention baseline, the calculation of the contention ratio can be stably executed under any extreme scenario. Furthermore, because this baseline value is extremely small, its actual impact on the calculation results is negligible and will not change the physical meaning and numerical characteristics of the contention ratio. Specifically, the specific value of the zero-division prevention baseline can be flexibly adjusted according to actual accuracy requirements, such as setting it to 0.001 or 0.1, but it must be ensured to be much smaller than the numerical range of normal load parameters to ensure the accuracy of the calculation results.

[0044] It's important to note that this step also requires defining the computational load parameters. While both CPU and GPU load rates fall under the category of computing resources, their weights differ across different business scenarios. For instance, in AI inference scenarios, fluctuations in GPU load have a greater impact on overall performance; in traditional web service scenarios, CPU load is more representative. The computational load parameter is used to uniformly characterize the overall utilization level of the server's computing resources. It integrates the loads of both CPU and GPU computing resources into a single comprehensive indicator. By weighted summing of CPU and GPU load rates, a single quantitative value that comprehensively reflects the server's overall computing load is obtained. This parameter retains the real-time nature of the original load data while adapting to different hardware platforms through weighting coefficients.

[0045] Specifically, the calculation of the computational load parameters involves a first weighting coefficient corresponding to the CPU load rate and a second weighting coefficient corresponding to the GPU load rate. These two weighting coefficients are not fixed but can be dynamically adjusted based on the server's hardware configuration and the business operation scenario. For example, in CPU-intensive business scenarios, the first weighting coefficient can be set larger to make the computational load parameters more sensitive to changes in CPU load; in GPU-intensive business scenarios, the second weighting coefficient can be set larger. This flexible configuration of the weighting coefficients allows the computational load parameters to adapt to the needs of different hardware platforms and business types, avoiding perceptual biases caused by hardware differences or scenario changes. In practical applications, the specific values ​​of the weighting coefficients can be determined through offline calibration, online learning, or expert experience, and this application does not impose specific limitations on this. For example, if the first weighting coefficient is 0.4 and the second weighting coefficient is 0.6, then the computational load parameter = CPU load rate × 0.4 + GPU load rate × 0.6.

[0046] S103. Identify the server load type based on the computing load parameters and the computing resource contention ratio, IO resource contention ratio, and total resource contention ratio.

[0047] It should be noted that the indicators obtained above can only reflect the magnitude of computing pressure and the intensity of resource competition, but cannot intuitively reflect the overall load type of the server. Therefore, it is necessary to construct load type identification rules based on the above parameters, and map the scattered quantitative values ​​into load type labels with clear business meanings.

[0048] Specifically, identifying the server load type based on the computing load parameters and the computing resource contention ratio, IO resource contention ratio, and total resource contention ratio includes:

[0049] (1) Determine the load type judgment threshold and resource contention judgment threshold respectively based on the server's hardware configuration parameters and business operation requirements.

[0050] It should be noted that the load type judgment threshold is used to distinguish whether the overall server load is low or medium-high. Its value is directly related to the server's hardware configuration parameters such as CPU model, number of cores, memory specifications, and hardware rated load capacity, and is also set in conjunction with the actual load requirements of business operations. The resource contention judgment threshold is used to determine whether there is a significant resource contention imbalance on the server. Its value is determined based on the business's tolerance for resource fluctuations, system stability requirements, and actual operating experience. Both types of thresholds are set according to the adaptability of actual hardware and business scenarios, rather than using fixed values. For example, the load type judgment threshold can be 30%, and the resource contention judgment threshold can be 1.5. Specifically, for AI training servers configured with high-performance CPUs and GPUs, whose computing resources are abundant, the load type judgment threshold can be appropriately increased to avoid misjudging normal business load as low load. For IO-intensive database servers, whose business is sensitive to IO latency, the resource contention judgment threshold can be appropriately decreased to identify IO bottleneck risks earlier. The methods for determining the thresholds may include offline calibration, online learning, expert experience, or business SLA requirements, etc., and this application does not specifically limit them.

[0051] (2) When the calculated load parameter is less than the load type determination threshold, the server load type is determined to be low load balancing type.

[0052] If the calculated load parameters are below the load type determination threshold, it indicates that the server's overall computing pressure is currently low, and system resources are idle or lightly loaded. In this case, regardless of the competition between computing and I / O resources, the server has no significant resource bottlenecks and is therefore uniformly classified as a low-load balanced type. This type indicates that the server is operating under low load and can readily handle more business workloads.

[0053] (3) When the computing load parameter is greater than or equal to the load type determination threshold, the total resource contention ratio is greater than the resource contention determination threshold, and the computing resource contention ratio is greater than the IO resource contention ratio, the server load type is determined to be compute-intensive.

[0054] When the computing load parameters have reached a medium-high load level, and the total resource contention ratio exceeds the resource contention judgment threshold, it indicates that there is a significant resource contention phenomenon on the server. If the computing resource contention ratio is greater than the IO resource contention ratio, according to the definition of contention ratio in S102, the load pressure of computing resources (CPU+GPU) is significantly higher than that of IO resources. The server's operating bottleneck is concentrated at the computing power level, which often occurs in computing power-consuming business scenarios such as AI inference, numerical calculation, video encoding, and scientific computing. Therefore, this state is judged as computing-intensive.

[0055] (4) When the computing load parameter is greater than or equal to the load type determination threshold, the total resource contention ratio is greater than the resource contention determination threshold, and the IO resource contention ratio is greater than the computing resource contention ratio, the server load type is determined to be IO intensive.

[0056] If the computing load parameters are at a medium-high load level and the total resource contention ratio exceeds the resource contention threshold, and the IO resource contention ratio is greater than the computing resource contention ratio, it indicates that the load pressure of IO resources is significantly higher than that of computing resources. The server's operational bottleneck is concentrated in IO-related links such as data reading and writing, network transmission, and storage interaction. This often occurs in IO-dependent business scenarios such as database services, file storage, big data transmission, and cloud disk reading and writing. Therefore, this state is judged as IO-intensive.

[0057] (5) When the calculated load parameter is greater than or equal to the load type determination threshold and the total resource contention ratio is less than or equal to the resource contention determination threshold, the server load type is determined to be balanced.

[0058] When the computational load parameters are at a medium-high load level, but the total resource contention ratio does not exceed the resource contention threshold, it indicates that although the overall load of the server is high, there is no obvious imbalance in competition between computational resources and IO resources. The load levels of the two are well matched, and the resource usage is relatively coordinated. There is no single resource bottleneck, and the overall server operation is stable and efficient. Therefore, this state is judged as balanced.

[0059] It should be noted that the determination criteria for the four load types mentioned above are independent and non-overlapping, and cover all possible load scenarios. The low-load balancing type covers scenarios with low loads; the compute-intensive and I / O-intensive types respectively cover scenarios with high loads and significant resource contention, and the direction of contention is accurately distinguished by comparing the compute resource contention ratio and the I / O resource contention ratio; the balanced type covers scenarios with high loads but no significant resource contention. This classification system ensures that any set of input parameters can be uniquely mapped to a load type, providing a reliable foundation for the accurate matching of subsequent control strategies.

[0060] S104. Identify the server load level based on the calculated load parameters.

[0061] After identifying the server load type, this step further categorizes the server load levels based on computational load parameters to achieve a tiered quantification of server operating pressure. It's important to note that load type characterizes the server's current resource bottlenecks and business characteristics, while load level characterizes the server's current computing resource utilization and congestion. These two aspects describe the server's load status from different dimensions. Therefore, by tiering and identifying load levels, the current load on the server can be intuitively reflected.

[0062] Specifically, identifying the server load level based on the calculated load parameters includes:

[0063] (1) Determine the first load level threshold and the second load level threshold respectively based on the server's hardware performance indicators and the load baseline requirements of business operation.

[0064] It should be noted that both the first and second load level thresholds are adaptively set based on the actual hardware performance indicators of the server and the baseline requirements of the business operation load, rather than using fixed values. The hardware performance indicators include parameters that reflect the server's hardware capacity, such as CPU architecture, number of cores, clock speed, and GPU computing power. The baseline requirements of the business operation load are determined based on factors such as business type, real-time requirements, stability requirements, and business SLAs, ensuring that the threshold settings align with the actual operational needs of different hardware platforms and business scenarios. The specific methods for determining the thresholds can include offline calibration, expert experience configuration, and historical operational data statistics; this application does not further limit these methods.

[0065] It should be noted that the first load level threshold in this step can be the same as the load type determination threshold used in S103, or it can be set to a different value according to the actual scenario requirements. The two are clearly distinct in function. The load type determination threshold distinguishes between low-load balancing and other high-load types, serving the classification of load types; the first load level threshold divides low load levels into stable load levels, serving the grading of load pressure levels. In practical applications, to simplify system configuration and improve the consistency of the judgment logic, both can be set to the same value, for example, both set to 30%. In scenarios requiring finer distinction between load type and load level, they can also be configured independently, without interference or conflict.

[0066] (2) When the calculated load parameter is less than or equal to the first load level threshold, the server load level is determined to be low load level.

[0067] When the computing load parameters are at or below the first load level threshold, it indicates that the combined computing resource usage of the server's CPU and GPU is low, the overall system operating pressure is low, and resources are sufficient. At this time, the server can quickly respond to new business requests without computing resource congestion or increased processing latency, and is therefore judged as a low load level.

[0068] (3) When the calculated load parameter is greater than the first load level threshold and less than or equal to the second load level threshold, the server load level is determined to be a stable load level.

[0069] When the calculated load parameters are between the first load level threshold and the second load level threshold, it indicates that the server has entered a normal business carrying state. There is a certain amount of computing resources occupied, but it has not reached the overload critical state. The system resource usage is reasonable and the operating state is stable. It can meet the current business processing needs and retain a certain amount of resource redundancy to cope with instantaneous fluctuations. It belongs to a normal and healthy operating range, and is therefore judged as a stable load level.

[0070] (4) When the calculated load parameter is greater than the second load level threshold, the server load level is determined to be a high load level.

[0071] When the computational load parameters exceed the second load level threshold, it indicates that the server's CPU and GPU computing resources are under high utilization, the overall system load is high, and the remaining resource space is limited. If it continues to be in this state, there may be risks such as increased processing latency and slower response. It is also a key judgment condition for starting strategies such as enhanced heat dissipation, power consumption adjustment, and resource scheduling. Therefore, it is judged as a high load level.

[0072] It should be noted that the above three load level determination criteria adopt a left-closed and right-closed interval division rule (i.e., ≤ first load level threshold, > first load level threshold and ≤ second load level threshold, > second load level threshold), ensuring that any calculated load parameter value can be uniquely mapped to a load level without overlap. This hierarchical system, independent yet complementary to the load type determination in S103, reveals the directional characteristics of resource competition, while the load level quantifies the absolute intensity of the load. Together, they constitute the scene dimension and absolute value dimension in the three-dimensional feature perception system. Through this step, we obtain the current server load level label, including low load level, stable load level, and high load level. This label, combined with the load type label obtained in S103, enables the downstream control system to not only know the current load type but also its absolute intensity level, providing a basis for subsequently formulating differentiated control strategies.

[0073] S105. Identify the slope characteristics of the server load based on the slope direction and absolute value of the slope of the calculated load parameters.

[0074] After identifying the server load type and load level, this step further captures the dynamic trend of server load changes. Specifically, it identifies the slope characteristics of the server load by calculating the direction and absolute value of the slope of the load parameter changes. The slope characteristics characterize the dynamic trend of the load. Based on the preceding description, it is impossible to determine whether the load is in a stable state, a slowly changing state, or a sudden fluctuation state simply by knowing the load type and load level. Identifying the slope characteristics allows for a comprehensive understanding of both the static attributes and dynamic trends of the server load.

[0075] Specifically, identifying the slope characteristics of the server load based on the slope direction and absolute value of the calculated load parameters includes:

[0076] (1) Determine the slope statistical period and obtain the initial value of the calculated load parameter at the start time and the final value of the calculated load parameter at the end time of the statistical period.

[0077] It should be noted that the representation of slope characteristics relies on the quantification of the changing trend of load parameters within a specific time window. Therefore, it is first necessary to determine a reasonable slope statistical period. The length of this period is not fixed but can be determined comprehensively based on the server's business type, sensitivity to load fluctuations, and control response requirements. For businesses with frequent load fluctuations and high response speed requirements, the slope statistical period can be set to a shorter duration to quickly capture the instantaneous trend of load changes. For businesses with gentle load fluctuations and higher stability requirements, the slope statistical period can be set to a longer duration to avoid misjudgment of slope characteristics caused by instantaneous fluctuations. After determining the slope statistical period, the system will synchronously collect the initial value of the calculated load parameters at the start of the period and the final value of the calculated load parameters at the end of the period.

[0078] (2) Divide the difference between the final value and the initial value of the calculated load parameter by the duration of the slope statistical period to obtain the slope of the calculated load parameter. The positive or negative attribute of the slope is the slope direction.

[0079] The slope calculation follows the basic mathematical definition of slope, quantifying the rate of load change by the change in quantity / time. Specifically, the difference between the final and initial values ​​of the load parameters reflects the magnitude of load change within the statistical period. A positive difference indicates an upward trend in load within that period; a negative difference indicates a downward trend; and a zero difference indicates no change in load. Dividing this difference by the duration of the slope statistical period yields the rate of load change per unit time, i.e., the slope. It's important to note that the positive or negative attribute of the slope determines the direction of the load change, which is the basis for determining whether the load is increasing or decreasing: a positive slope corresponds to an increasing load, and a negative slope corresponds to a decreasing load. The larger the absolute value of the slope, the faster the rate of load change.

[0080] (3) Perform an absolute value operation on the slope of change to obtain the slope intensity, wherein the slope intensity is the absolute value of the slope of change.

[0081] It's important to note that the slope direction only indicates the trend of load change, not the rate of change. Slope intensity, however, the absolute value of the slope, directly reflects the magnitude of the load change rate. A larger slope intensity indicates a greater and more drastic change in load per unit time; a smaller slope intensity indicates a smoother load change, closer to a steady state. By separating the slope into direction and intensity, we can more accurately and comprehensively describe the dynamic characteristics of the load.

[0082] (4) Identify the slope characteristics of the server load based on the numerical characteristics of the slope direction and slope intensity.

[0083] By combining the slope direction (rising / falling) and slope intensity (degree of change), and using preset thresholds, the slope characteristics of server load can be subdivided into different types to match different load change scenarios. Specifically, a first slope threshold and a second slope threshold need to be determined based on the server's business type and sensitivity to load fluctuations. The values ​​of these two thresholds can be flexibly selected according to the scenario. For businesses that are highly sensitive to load fluctuations and require rapid response to sudden load surges, the first and second slope thresholds can be set to relatively small values ​​to identify load change trends earlier. For businesses that are less sensitive to load fluctuations and prioritize stability, both thresholds can be appropriately increased to avoid misjudgments due to minor fluctuations. Specifically, this application does not impose specific limitations on the method of determining the thresholds.

[0084] The specific identification rules are as follows: When the slope intensity is less than or equal to the first slope threshold, the slope characteristic of the server load is determined to be a stable-steady characteristic. At this time, the slope intensity is extremely small, indicating that the rate of change of the load within the statistical period is very slow, with almost no obvious fluctuations. The server load is in a stable state and no additional control intervention is required. This characteristic often occurs in scenarios with stable business volume. When the slope intensity is greater than the first slope threshold and less than or equal to the second slope threshold, it indicates that there is a significant change in the load, but the rate of change is moderate and has not reached the level of sudden fluctuations. At this time, if the slope direction is positive, it is determined to be a stable-slowly rising characteristic, indicating that the server load is gradually increasing at a smooth rate. The load growth trend can be predicted in advance, reserving buffer time for subsequent control. If the slope direction is negative, it is determined to be a stable-slowly falling characteristic, indicating that the server load is gradually decreasing at a smooth rate. The system resource consumption will gradually decrease, and the control strategy can be adjusted in a timely manner to save energy. When the slope intensity is greater than the second slope threshold, it indicates that the load change rate is extremely fast, belonging to a sudden fluctuation scenario. At this time, if the slope direction is positive, it is determined to be a sudden-rapid rise characteristic, indicating that the server load has suddenly surged, and may face the risk of instantaneous overload. It is necessary to immediately start enhanced control strategies, such as improving heat dissipation efficiency and increasing hardware operating frequency to ensure system stability. If the slope direction is negative, it is determined to be a sudden-rapid fall characteristic, indicating that the server load has suddenly dropped, and system resources are instantly abundant. Control strategies can be adjusted in time, such as reducing fan speed and reducing frequency to save energy, in order to reduce resource waste.

[0085] The identified slope features, along with the load type in S103 and the load level in S104, complement each other to form the scene dimension (load type), absolute value dimension (load level), and trend dimension (slope features) in the three-dimensional feature perception system.

[0086] S106. Generate standardized perception labels from the identification results of the load type, load level, and slope features according to preset coding rules, and output the three-dimensional feature perception results of the server load based on the standardized perception labels.

[0087] It should be noted that after obtaining the recognition results in the three dimensions, these results exist in the form of discrete labels and have not yet formed a unified data structure, making it inconvenient for downstream systems to directly call and parse them. Therefore, it is necessary to integrate and encapsulate the recognition results of these three dimensions according to preset encoding rules to generate a structured and standardized perceptual label as the final output of 3D feature perception.

[0088] Specifically, generating standardized perception tags according to preset encoding rules includes: encoding the identification results of the load type, load level, and slope feature using one byte; setting the seventh and sixth bits of the byte as load type identifier bits; setting the fifth and fourth bits of the byte as load level identifier bits; and setting the third to zeroth bits of the byte as slope feature identifier bits; configuring corresponding binary values ​​for the load type identifier bit, load level identifier bit, and slope feature identifier bit; uniquely identifying the identification results of the load type, load level, and slope feature; and generating the standardized perception tag.

[0089] It should be noted that one byte (8 bits) is used for encoding, taking into account factors such as the number of recognition results and system transmission and storage efficiency. One byte has a small storage space, which can quickly realize data transmission and processing, and adapt to the needs of real-time server control; at the same time, 8 bits can be divided into different identifier bit regions, which can just meet the unique identification requirements of the three-dimensional recognition results. The load type identifier occupies 2 bits (the seventh and sixth bits). According to S103, there are 4 load types. 2 bits can represent 4 different combinations (00, 01, 10, 11), which can just achieve the unique encoding of 4 load types. The load level identifier also occupies 2 bits (the fifth and fourth bits), corresponding to the 3 load levels in S104. The 4 combinations of 2 bits can reserve 1 redundancy to adapt to the expansion needs of subsequent load level types. The slope feature identifier occupies 4 bits (the third to the zeroth bits), corresponding to the 5 slope features in S105. 4 bits can represent 16 combinations, which can not only meet the encoding needs of the current 5 slope features, but also provide sufficient space for the subsequent subdivision and expansion of slope features.

[0090] For example, low load balancing can be coded as 00, compute-intensive as 01, I / O-intensive as 10, and balanced as 11; low load level can be coded as 00, stable load level as 01, and high load level as 10; stable-stable characteristics can be coded as 0000, stable-slowly rising characteristics as 0001, stable-slowly falling characteristics as 0010, burst-rapidly rising characteristics as 0100, and burst-rapidly falling characteristics as 0101. Specific coding values ​​can be flexibly adjusted according to the actual system configuration, and this application does not impose specific limitations on them. Through this coding rule, the three dimensions of load characteristics can be integrated into a standardized sensing label, achieving a unified representation of load characteristics, facilitating rapid reading and parsing by downstream heat dissipation and power consumption control systems.

[0091] It should be noted that server load parameters may fluctuate instantaneously during actual operation. Generating standardized perception labels directly based on a single identification result could lead to frequent label switching, causing frequent adjustments by the downstream control system and affecting server stability. Furthermore, instantaneous fluctuations could distort the perception results. Therefore, the load type identification result, load level identification result, and slope feature identification result need to be averaged using a first sliding window, a second sliding window, and a third sliding window, respectively; the number of collection points in the first, second, and third sliding windows are different for each.

[0092] The sliding window's function is to smooth and denoise continuous recognition results. By averaging multiple sampling points within the window, it weakens the impact of instantaneous fluctuations on the perception results, improving the stability and reliability of standardized perception labels. Three sliding windows with different numbers of sampling points are used because the fluctuation characteristics of load type, load level, and slope feature differ. Load type is a qualitative feature with a relatively low frequency of change, so a sliding window with more sampling points, such as 10 points, can be used to ensure the stability of the load type label and avoid misclassification due to minor fluctuations. Load level is a quantitative feature with a moderate frequency of change, so a sliding window with a moderate number of sampling points, such as 8 points, can be used to ensure smoothing while maintaining real-time perception. Slope feature represents the dynamic trend of load changes and has high real-time requirements, needing to quickly capture trend changes. Therefore, a sliding window with fewer sampling points, such as 5 points, can be used to avoid lag in trend perception due to an excessively large window. Specifically, the number of sampling points in each sliding window can be flexibly adjusted according to the server's business type and load fluctuation characteristics.

[0093] The load type smoothing result, load level smoothing result, and slope feature smoothing result obtained after calculating the average of each sliding window are regenerated into standardized perception labels according to the preset encoding rules. The regenerated standardized perception labels are then output as the three-dimensional feature perception result of the server load. After sliding window smoothing, the standardized perception labels can more realistically and stably reflect the actual state of the server load, avoiding interference caused by instantaneous fluctuations and providing a more reliable decision-making basis for downstream control systems.

[0094] Furthermore, it should be noted that after the server outputs the 3D feature perception results based on the standardized perception labels, the process includes:

[0095] (1) The three-dimensional feature perception results of the server load are sent to the server's heat dissipation control system and power consumption control system.

[0096] The server's thermal management system and power consumption control system are hardware management systems that ensure the server's stable and efficient operation. Both need to formulate targeted control strategies based on the server's current load status. Sending the load's three-dimensional feature perception results (i.e., smoothed standardized perception labels) to these two systems in real time ensures that the control system obtains complete load information in a timely manner.

[0097] (2) The heat dissipation control system adjusts the fan speed and heat dissipation strategy according to the load type, load level and slope characteristics.

[0098] The goal of a thermal control system is to match the heat generated by the load, preventing server hardware damage or performance degradation due to overheating. Load type, load level, and slope characteristics directly determine the server's heat intensity and trend. Therefore, for compute-intensive loads, where the CPU and GPU are under high load and generate significant and continuous heat, the thermal control system needs to increase fan speeds and employ enhanced cooling strategies to ensure timely heat dissipation. For I / O-intensive loads, where heat primarily originates from the I / O controller and network card, the heat generation is relatively low, allowing for conventional cooling strategies and appropriately reducing fan speeds to save energy. For low-load, balanced loads, where heat generation is minimal, fan speeds can be further reduced for energy-efficient cooling. Simultaneously, predictive control is performed based on slope characteristics. If the slope is stable with a slow increase, it indicates a gradual increase in load and heat generation, allowing for advance fan speed increases to prevent heat accumulation. If the slope is sudden with a rapid increase, the highest-level cooling strategy must be activated immediately to quickly address any sudden surge in heat. If the slope is stable with a slow decrease or sudden with a rapid decrease, fan speeds can be gradually reduced to balance cooling performance and energy efficiency.

[0099] (3) The power consumption control system adjusts the operating frequency and voltage of the CPU and GPU according to the load type, load level and slope characteristics.

[0100] The goal of the power consumption control system is to optimize power consumption while ensuring the normal operation of server services, reducing energy consumption and extending hardware lifespan. Therefore, for compute-intensive loads, it is necessary to ensure high computing power output of the CPU and GPU. The power consumption control system needs to increase the operating frequency and voltage of the CPU and GPU to ensure that the computing power requirements of the business are met. For I / O-intensive loads, the CPU and GPU loads are relatively low, and their operating frequency and voltage can be appropriately reduced to reduce power consumption without affecting business operations. For low-load balanced loads, the operating frequency and voltage of the CPU and GPU can be further reduced to achieve deep energy saving. At the same time, predictive control is performed based on slope characteristics. If the slope characteristic is stable to slowly rising, the operating frequency and voltage can be slightly increased in advance to reserve computing power space for load growth and avoid business delays due to insufficient computing power. If the slope characteristic is sudden to rapidly rising, the operating frequency and voltage need to be increased immediately to quickly respond to instantaneous computing power demands. If the slope characteristic is stable to slowly decreasing or sudden to rapidly decreasing, the operating frequency and voltage can be gradually reduced to achieve dynamic optimization of power consumption.

[0101] The method provided in this embodiment has been adapted to the specific interfaces and collection cycles for BMC to collect CPU load rate, GPU load rate, and IO bandwidth ratio. It also clarifies the preprocessing rules for the raw collection parameters. Through standardized formatting, time alignment, and hardware adaptation preprocessing operations, it eliminates various interference factors in the raw parameters while ensuring comprehensive and real-time collection of core hardware load parameters, laying a high-quality data foundation for the accurate calculation of subsequent quantitative indicators. The calculation logic for computing resource contention ratio, IO resource contention ratio, total resource contention ratio, and computing load parameters is limited. Anomalies are avoided by introducing a zero-based benchmark. A weighted summation method is designed for the computing load parameters, ensuring that various quantitative indicators can accurately and stably quantify the internal resource contention relationship and overall computing load of the server, while also adapting to the actual needs of different hardware configurations and business operation scenarios. The identification rules for load type, load level, and slope characteristics are specifically designed in a layered and multi-dimensional manner, combining server hardware configuration parameters and business operation... The system sets adaptive thresholds to ensure that the 3D feature recognition results better reflect the actual operating state of the server. This allows for precise differentiation of different load scenarios, intensity levels, and dynamic trends, achieving refined and comprehensive perception of server load status. A specific one-byte encoding rule was designed for the standardized perception tags, rationally dividing the identifier bits for each dimension of features. This achieves a unified, structured encapsulation of multi-dimensional perception results, facilitating rapid reading and parsing by downstream systems, while also reserving ample space for future feature type expansion. Furthermore, by using sliding windows with varying numbers of sampling points to smooth the recognition results for each dimension, the system effectively reduces the interference of instantaneous load fluctuations on the perception results, improving the stability of the standardized perception tags. This ensures that the final output of the server load 3D feature perception results more closely reflects the actual server operating state, enabling efficient linkage with downstream thermal control and power consumption regulation systems. This ensures the accuracy and timeliness of regulation strategies, ultimately improving the server's energy efficiency ratio and reducing energy consumption while guaranteeing stable server operation.

[0102] Example 2

[0103] Corresponding to the aforementioned embodiment of a BMC-based server load three-dimensional feature perception method, this application also provides an embodiment of a BMC-based server load three-dimensional feature perception device.

[0104] Figure 2 This is a schematic diagram of a second embodiment of a server load 3D feature sensing device based on BMC provided in this application. Please refer to... Figure 2 The device provided in this embodiment includes a data acquisition module 210, a calculation module 220, an identification module 230, and a processing module 240.

[0105] The acquisition module 210 is used to acquire the CPU load rate, GPU load rate and IO bandwidth ratio of the server through the BMC to obtain the real-time working parameters of the server.

[0106] The computing module 220 is used to calculate the computing resource contention ratio and the IO resource contention ratio based on the real-time working parameters; and to calculate the total resource contention ratio based on the computing resource contention ratio and the IO resource contention ratio, and to calculate the computing load parameters based on the CPU load rate and the GPU load rate.

[0107] The identification module 230 is used to identify the server load type based on the computing load parameters and the computing resource contention ratio, IO resource contention ratio and total resource contention ratio.

[0108] The identification module 230 is also used to identify the server load level based on the calculated load parameters;

[0109] The identification module 230 is also used to identify the slope characteristics of the server load based on the slope direction and absolute value of the slope of the calculated load parameters.

[0110] The processing module 240 is used to generate standardized perception labels from the identification results of the load type, load level, and slope features according to preset coding rules, and output the three-dimensional feature perception results of the server load based on the standardized perception labels.

[0111] The apparatus of this embodiment can be used to perform... Figure 1 The steps of the method embodiment shown are similar in principle and process, and will not be repeated here.

[0112] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0113] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0114] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for perceiving three-dimensional features of server load based on BMC, characterized in that, The method includes: By collecting the server's CPU load rate, GPU load rate, and IO bandwidth ratio through BMC, the server's real-time operating parameters can be obtained. The computational resource contention ratio and I / O resource contention ratio are calculated based on the real-time operating parameters; the total resource contention ratio is calculated based on the computational resource contention ratio and I / O resource contention ratio; and the computational load parameters are calculated based on the CPU load rate and GPU load rate. The server load type is identified based on the computing load parameters and the computing resource contention ratio, IO resource contention ratio, and total resource contention ratio. Identify the server load level based on the calculated load parameters; The slope characteristics of the server load are identified based on the slope direction and absolute value of the calculated load parameters. The identification results of the load type, load level, and slope features are used to generate standardized perception labels according to preset coding rules, and the server load three-dimensional feature perception results are output based on the standardized perception labels. The computing resource contention ratio is obtained by: adding the sum of the CPU load rate and the GPU load rate to a zero-prevention baseline value as the dividend, adding the IO bandwidth ratio to a zero-prevention baseline value as the divisor, and dividing by the sum to obtain the computing resource contention ratio. The calculation of the IO resource contention ratio includes: adding the IO bandwidth ratio to a zero-prevention baseline value as the dividend, adding the sum of the CPU load rate and GPU load rate to the zero-prevention baseline value as the divisor, and dividing by the sum to obtain the IO resource contention ratio.

2. The method according to claim 1, characterized in that, The step of calculating the total resource contention ratio based on the computing resource contention ratio and the I / O resource contention ratio includes: taking the maximum value of the computing resource contention ratio and the I / O resource contention ratio as the total resource contention ratio.

3. The method according to claim 1, characterized in that, The process of identifying server load type based on the computing load parameters and the computing resource contention ratio, IO resource contention ratio, and total resource contention ratio includes: Based on the server's hardware configuration parameters and business operation requirements, determine the load type judgment threshold and resource contention judgment threshold respectively; When the calculated load parameter is less than the load type determination threshold, the server load type is determined to be low load balancing. When the computing load parameter is greater than or equal to the load type determination threshold, the total resource contention ratio is greater than the resource contention determination threshold, and the computing resource contention ratio is greater than the IO resource contention ratio, the server load type is determined to be compute-intensive. When the computational load parameter is greater than or equal to the load type determination threshold, the total resource contention ratio is greater than the resource contention determination threshold, and the IO resource contention ratio is greater than the computational resource contention ratio, the server load type is determined to be IO intensive. When the calculated load parameter is greater than or equal to the load type determination threshold, and the total resource contention ratio is less than or equal to the resource contention determination threshold, the server load type is determined to be balanced.

4. The method according to claim 1, characterized in that, The process of identifying the server load level based on the calculated load parameters includes: Based on the server's hardware performance indicators and the load baseline requirements of business operations, the first load level threshold and the second load level threshold are determined respectively. When the calculated load parameter is less than or equal to the first load level threshold, the server's load level is determined to be low. When the calculated load parameter is greater than the first load level threshold and less than or equal to the second load level threshold, the server's load level is determined to be a stable load level. When the calculated load parameter is greater than the second load level threshold, the server's load level is determined to be high.

5. The method according to claim 1, characterized in that, The step of identifying the slope characteristics of server load based on the slope direction and absolute value of the calculated load parameters includes: Determine the slope statistical period, and obtain the initial value of the calculated load parameter at the start time and the final value of the calculated load parameter at the end time of the statistical period; The difference between the final value and the initial value of the calculated load parameter is divided by the duration of the slope statistical period to obtain the slope of the calculated load parameter. The positive or negative attribute of the slope is the slope direction. The slope intensity is obtained by taking the absolute value of the slope change; Based on the numerical characteristics of the slope direction and slope intensity, the slope characteristics of the server load are identified.

6. The method according to claim 1, characterized in that, The process of generating standardized perception tags according to preset encoding rules includes: The identification results of the load type, load level, and slope feature are encoded using one byte. The seventh and sixth bits of the byte are set as load type identifier bits, the fifth and fourth bits of the byte are set as load level identifier bits, and the third to zero bits of the byte are set as slope feature identifier bits. Configure corresponding binary values ​​for the load type identifier, load level identifier, and slope feature identifier to uniquely identify the recognition results of the load type, load level, and slope feature, and generate the standardized perception label.

7. The method according to claim 1, characterized in that, After the server outputs the 3D feature perception results based on the standardized perception labels, the process includes: The server load 3D feature perception results are sent to the server's heat dissipation control system and power consumption regulation system; The heat dissipation control system adjusts the fan speed and heat dissipation strategy according to the load type, load level and slope characteristics; The power consumption control system adjusts the operating frequency and voltage of the CPU and GPU according to the load type, load level and slope characteristics.

8. A server load three-dimensional feature perception device based on BMC, characterized in that, The device includes an acquisition module, a calculation module, an identification module, and a processing module; The acquisition module is used to acquire the server's CPU load rate, GPU load rate, and IO bandwidth ratio through the BMC to obtain the server's real-time operating parameters. The computing module is used to calculate the computing resource contention ratio and the IO resource contention ratio based on the real-time working parameters; and to calculate the total resource contention ratio based on the computing resource contention ratio and the IO resource contention ratio, and to calculate the computing load parameters based on the CPU load rate and the GPU load rate. The identification module is used to identify the server load type based on the computing load parameters and the computing resource contention ratio, IO resource contention ratio, and total resource contention ratio. The identification module is also used to identify the server load level based on the calculated load parameters; The identification module is also used to identify the slope characteristics of the server load based on the slope direction and absolute value of the slope of the calculated load parameters. The processing module is used to generate standardized perception labels from the identification results of the load type, load level, and slope features according to preset coding rules, and output the three-dimensional feature perception results of the server load based on the standardized perception labels. The computing resource contention ratio is obtained by: adding the sum of the CPU load rate and the GPU load rate to a zero-prevention baseline value as the dividend, adding the IO bandwidth ratio to a zero-prevention baseline value as the divisor, and dividing by the sum to obtain the computing resource contention ratio. The calculation of the IO resource contention ratio includes: adding the IO bandwidth ratio to a zero-prevention baseline value as the dividend, adding the sum of the CPU load rate and GPU load rate to the zero-prevention baseline value as the divisor, and dividing by the sum to obtain the IO resource contention ratio.

Citation Information

Patent Citations

  • Software load computing resource virtualization allocation method and device, equipment and medium

    CN121144045A

  • Edge computing resource scheduling management system of Internet of Things equipment

    CN121879948A