Server system and power consumption control method

By directly connecting the microcontroller unit to the accelerator card unit and combining it with a lightweight convolutional neural network model, multi-dimensional data is collected in real time, and control parameters are dynamically generated. This solves the real-time and accuracy problems of GPU power consumption control and achieves flexible and accurate power management.

CN121879549APending Publication Date: 2026-04-17INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2026-03-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies for GPU power consumption control suffer from poor real-time performance and low precision, failing to adapt to changes in GPU load in a timely and accurate manner, resulting in performance loss or power waste.

Method used

By directly connecting the microcontroller unit to the accelerator card unit, multi-dimensional operating data is collected in real time. Combined with a lightweight convolutional neural network model, control parameters are dynamically generated to achieve flexible and accurate power consumption control.

Benefits of technology

It improves the real-time performance and accuracy of GPU power consumption control, enabling it to adapt to different load conditions in a timely and accurate manner, reducing peak power consumption and avoiding performance loss and power waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121879549A_ABST
    Figure CN121879549A_ABST
Patent Text Reader

Abstract

The invention discloses a server system and a power consumption control method, and relates to the technical field of servers, the server system adopts a micro-control unit to collect real-time operation data of an accelerator card unit, the micro-control unit is directly connected with the accelerator card unit, high-speed collection of the real-time operation data can be achieved, and the power consumption of the accelerator card unit can be controlled through the real-time operation data. In combination with the preset regulation and control model, the regulation and control parameters adaptive to the running state of the current accelerator card unit are dynamically generated in an artificial intelligence mode, and compared with a traditional regulation and control mode adopting a fixed threshold value, the regulation and control parameters can better adapt to the running state of the current accelerator card unit; and flexible and accurate power consumption regulation and control can be realized according to different operation states. By means of the mode, the real-time performance and the regulation and control precision of power consumption regulation and control of the accelerator card unit in the server system are improved, and power consumption regulation and control of the accelerator card unit can be achieved timely and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server technology, and in particular to a server system and a power consumption control method. Background Technology

[0002] With the surge in demand for large-scale AI model training and high-performance computing, the Graphics Processing Unit (GPU) has become a core power source for servers. To reduce the overall power consumption of servers, GPU power consumption control is necessary. In related technologies, the Baseboard Management Controller (BMC) collects real-time GPU power consumption data through sensors and presets a fixed power consumption threshold. When the GPU power consumption exceeds the threshold, the BMC sends a frequency reduction command to the GPU or directly limits the power supply, forcing the power consumption to be suppressed below the threshold.

[0003] However, the relevant technologies have poor real-time performance and low precision in GPU control. How to achieve timely and accurate power consumption control of GPUs has become an urgent technical problem to be solved. Summary of the Invention

[0004] This application provides a server system and a power consumption control method to at least address the problem in related technologies of how to improve the topology identification efficiency of server systems, thereby improving fault location efficiency.

[0005] This application provides a server system, including:

[0006] Accelerator card unit;

[0007] A microcontroller unit is connected to the accelerator card unit. The microcontroller unit is configured to: collect real-time operating data of the accelerator card unit and generate control parameters based on a preset control model and the real-time operating data.

[0008] A substrate management unit is connected to the microcontroller unit, and the substrate management unit is configured to receive control parameters sent by the microcontroller unit.

[0009] A control unit is connected to the baseboard management unit, and the control unit is configured to: receive control parameters sent by the baseboard management unit; and control the operating state of the accelerator card unit according to the control parameters.

[0010] This application also provides a power consumption control method applied to a microcontroller unit of a server system, comprising:

[0011] Collect real-time operating data from the accelerator card unit;

[0012] Based on the preset control model and the real-time operating data, control parameters are generated;

[0013] The control parameters are sent to the control unit so that the control unit can control the operating state of the accelerator card unit according to the control parameters.

[0014] This application describes a server system that uses a microcontroller unit (MCU) to collect real-time operating data from the accelerator card unit. The MCU is directly connected to the accelerator card unit, enabling high-speed data acquisition. Compared to the traditional method of data acquisition using a baseboard management controller, this improves data acquisition efficiency and real-time performance. After acquiring high-real-time operating data, the MCU in this embodiment dynamically generates control parameters adapted to the current operating state of the accelerator card unit using artificial intelligence, based on the real-time operating data and a preset control model. Compared to the traditional method of using fixed thresholds for control, this method better adapts to the current operating state of the accelerator card unit, achieving flexible and accurate power consumption control for different operating states. This approach improves the real-time performance and accuracy of power consumption control for the accelerator card unit in the server system, enabling timely and accurate power consumption control. Attached Figure Description

[0015] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 Schematic diagram of the server system provided in the embodiments of this application Figure 1 ;

[0017] Figure 2 Schematic diagram of the server system provided in the embodiments of this application Figure 2 ;

[0018] Figure 3 Schematic diagram of the server system provided in the embodiments of this application Figure 3 ;

[0019] Figure 4 A flowchart illustrating the power consumption control method provided in an embodiment of this application;

[0020] Figure 5 This application provides a functional diagram of each unit in a server system.

[0021] Figure 6This is a schematic diagram of the power consumption control device provided in the embodiments of this application;

[0022] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0024] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0025] With the surge in demand for large-scale AI model training and high-performance computing, graphics processing units (GPUs) have become a core power source for servers. A single high-end GPU can consume over 400W under full load, and due to task load fluctuations, such as matrix multiplication and communication phase switching, the peak-to-valley power consumption difference often exceeds 30%. To reduce the overall power consumption of servers, GPU power consumption control is necessary. Current key technologies in GPU power management include: hardware-level power path control and capacitor / battery power buffering; software-level operating system (OS) dynamic task scheduling and GPU core frequency reduction; and heterogeneous architecture-level BMC monitoring and MCU edge chip-assisted control. The core technology aims to maximize GPU power consumption reduction while ensuring AI task throughput and inference latency performance. The challenge lies in balancing "real-time response speed" and "control precision." GPU load fluctuations can reach the microsecond level, making it difficult for traditional control schemes to capture instantaneous changes and easily leading to excessive performance loss. GPU power consumption control schemes have significant flaws. While capacitors and batteries buffer and smooth power fluctuations, they only address power stability without reducing actual energy consumption. Dual power path designs only achieve "on / off control" and cannot adapt to dynamic loads. Reliance on dynamic scheduling by the OS or server results in millisecond-level latency in data acquisition and instruction execution, making it difficult to capture microsecond-level load changes in the GPU. Although the BMC has system monitoring capabilities, it lacks real-time data processing hardware. A heterogeneous architecture system with "real-time monitoring, intelligent decision-making, and precise execution" is needed to achieve a dynamic balance between GPU power consumption and performance.

[0026] Among related technologies, the GPU fixed threshold power consumption capping scheme based on BMC is currently the most widely used GPU power consumption control technology in the server field. Its core scheme is as follows: BMC collects real-time power consumption data of GPU through sensors and presets a fixed power consumption threshold; when the GPU power consumption exceeds the threshold, BMC sends a frequency reduction command to the GPU or directly limits the power supply to force the power consumption to be suppressed below the threshold; at the communication layer, a private protocol or simple command pass-through is used to transmit only single control signals such as "power consumption exceeded" and "execute frequency reduction". However, the relevant technologies suffer from the following main problems: Lag in regulation: The power consumption sampling cycle of the BMC is in the millisecond range, which cannot capture the microsecond-level load fluctuations of the GPU. This results in regulation being triggered only after the power consumption peak has already occurred, leading to limited power suppression effects. Insufficient accuracy: The use of inflexible fixed thresholds fails to consider GPU load types, tensor core activity ratios, task characteristics, large language model (LLM) inference, and image rendering, making over-regulation prone to occur. Forced frequency reduction at low loads leads to performance loss or insufficient regulation, while excessively high thresholds at high loads still result in wasted power. Protocol closedness: Relying on proprietary communication protocols, compatibility is poor, unable to adapt to different manufacturers' OSs and GPU models, and can only transmit single instructions, failing to achieve fine-grained parameter regulation. Lack of intelligent decision-making: Lacking learning of the "load-power-performance" mapping relationship, the regulation logic is fixed and cannot dynamically adapt to the complex load changes of AI tasks. The poor real-time performance and low precision of the relevant technologies for GPU regulation make timely and accurate GPU power consumption regulation an urgent technical problem to be solved.

[0027] To address the aforementioned issues, this application provides a server system and power consumption control method. The server system includes a microcontroller unit directly connected to the accelerator card unit. This microcontroller unit efficiently collects real-time operating data from the accelerator card unit and, in conjunction with a pre-defined control model, dynamically generates control parameters adapted to the current operating state of the accelerator card unit using artificial intelligence. This enables timely and accurate control of the accelerator card unit's power consumption.

[0028] Optionally, in view of the shortcomings of the prior art, the embodiments of this application aim to solve the following core technical problems: slow control response: unable to capture GPU microsecond-level load fluctuations, resulting in the inability to suppress power consumption peaks in a timely manner; low control accuracy: relying on a single indicator such as power consumption or utilization rate or a fixed threshold, failing to achieve a dynamic balance of "load-power consumption-performance", easily causing performance loss or power waste; poor compatibility and scalability: using a proprietary protocol, with limited adaptability, and unable to adapt to the load characteristics of different AI tasks and different GPU models; no intelligent self-optimization capability: the control logic is fixed, unable to dynamically optimize decisions based on actual operating data, and adaptable to a single scenario.

[0029] Optionally, embodiments of this application provide a heterogeneous architecture system for reducing GPU power consumption in servers. Through heterogeneous collaboration between the BMC and MCU, combined with AI models to generate control parameters in real time, the OS is guided to dynamically adjust GPU utilization, thereby reducing power consumption while ensuring task performance.

[0030] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] Optionally, Figure 1 Schematic diagram of the server system provided in the embodiments of this application Figure 1 ,like Figure 1 As shown, the server system provided in this application embodiment includes:

[0032] Accelerator card unit 11.

[0033] The microcontroller unit 12 is connected to the accelerator card unit 11.

[0034] The substrate management unit 13 is connected to the microcontroller unit 12.

[0035] The control unit 14 is connected to the substrate management unit 13.

[0036] It is understood that the accelerator card unit 11, microcontroller unit 12, substrate management unit 13, and control unit 14 can be chip units from any manufacturer and of any type capable of performing the above functions. This application embodiment does not impose any specific restrictions on this.

[0037] The microcontroller unit 12 is configured to: collect real-time operating data of the accelerator card unit 11, and generate control parameters based on the preset control model and the real-time operating data.

[0038] The substrate management unit 13 is configured to receive control parameters sent by the microcontroller unit 12.

[0039] The control unit 14 is configured to receive control parameters sent by the baseboard management unit 13 and control the operating state of the accelerator card unit 11 according to the control parameters.

[0040] Optionally, the microcontroller unit 12 is connected to the accelerator card unit 11 via a single-channel interconnection link through peripheral components.

[0041] Optionally, the Peripheral Component Interconnect Express (PCIE) single-channel link can be a PCIe 4.0 lane. Its high bandwidth and low latency ensure real-time performance. Combined with the direct connection architecture, it enables the microcontroller unit to stably acquire real-time operating data from the graphics processing unit at microsecond intervals. It also supports advanced semantic data acquisition for precise control. With a single physical channel, it meets performance requirements while achieving zero occupation of the server system's main system computing resources, minimal hardware complexity, and power consumption, demonstrating the efficiency and economy of the system design.

[0042] Optionally, the accelerator card unit 11 may include any one or more of a GPU, a Tensor Processing Unit (TPU), a Neural Processing Unit (NPU), an Artificial Intelligence Accelerator (AI Accelerator), a Field-Programmable Gate Array (FPGA), a Data Processing Unit (DPU), and an Application-Specific Integrated Circuit (ASIC). In this embodiment, for ease of explanation, the accelerator card unit includes a GPU as an example.

[0043] Optionally, when the accelerator card unit 11 includes a GPU, the accelerator card unit 11 is equipped with a high-performance GPU and has at least one of a power consumption sensor, a utilization counter, and a tensor core activity monitoring unit built in, for collecting real-time operating data of the accelerator card unit 11.

[0044] Optionally, the accelerator card unit 11 supports outputting multi-dimensional operating data through a management interface used to acquire internal operating status data of the graphics processing unit, in response to real-time data reading requests from the microcontroller unit 12. This management interface can be the driver interface of the accelerator card unit 11.

[0045] Optionally, the substrate management unit 13 is a substrate management controller.

[0046] Optionally, the control unit 14 is an operating system running on the host.

[0047] Optionally, the control unit 14 is deployed on a server host based on a central processing unit (CPU).

[0048] Optionally, the control unit 14 receives control parameters through the GPU driver, performs operations such as GPU task allocation adjustment, core frequency control, and task fragmentation, and sends the actual execution results back to the BMC. Optionally, the actual execution results include at least one of actual utilization, power consumption change, and performance loss.

[0049] Optionally, the microcontroller unit 12 is directly connected to the accelerator card unit 11 via PCIe 4.0x1 Lane, and the microcontroller unit 12 communicates with the baseboard management unit 13 via Universal Serial Bus (USB) to achieve parameter transmission; the baseboard management unit 13 is connected to the control unit 14 via Ethernet to achieve standardized parameter pass-through.

[0050] The microcontroller unit 12 is responsible for real-time data acquisition and AI inference, the substrate management unit 13 is responsible for protocol conversion and data transmission, and the control unit 14 is responsible for load execution and result feedback, forming a closed-loop control architecture of "acquisition-decision-execution-feedback".

[0051] Optionally, the preset control model is a quantized lightweight convolutional neural network (TinyCNN) model.

[0052] Optionally, the quantization process here can be 8-bit integer quantization (INT8) to reduce the memory footprint of the preset control model, thereby reducing the power consumption of the microcontroller unit and ensuring the performance of power control.

[0053] Optionally, the real-time operational data includes at least one of the following: accelerator card unit real-time utilization, tensor core active percentage, core voltage, real-time power consumption, load fluctuation frequency, and task type.

[0054] The real-time utilization rate of the accelerator card unit can be expressed as GPU utilization, reflecting the overall busyness of the accelerator card unit.

[0055] Tensor core active percentage is a key metric designed specifically for AI workloads to differentiate computation types.

[0056] Core voltage and real-time power consumption are direct electrical status indicators of the accelerator card unit.

[0057] Load fluctuation frequency characterizes how fast the load changes and is key to capturing microsecond-level fluctuations.

[0058] Task types can be represented by task type encoding to distinguish different tasks such as LLM inference and image rendering.

[0059] Through multi-dimensional real-time operational data, the model can accurately perceive the operating status and power consumption status of the accelerator card unit 11, thereby enabling the model to achieve accurate modeling, accurate mapping, and accurate inference in multiple dimensions. The high-frequency collected data such as fluctuation frequency is a prerequisite for the system to capture and respond to microsecond-level load changes, thus improving the accuracy of power consumption control.

[0060] Optionally, the control parameters include at least one of the target utilization rate, core frequency reduction magnitude, and task fragmentation threshold.

[0061] The target utilization rate can guide the task scheduler of the control unit 14 to achieve the expected level of computing resource usage of the accelerator card unit 11.

[0062] The core clock frequency reduction can guide the accelerator card unit 11 driver to adjust the core clock frequency by a specific percentage or absolute value.

[0063] Task sharding thresholds can guide applications or runtimes to break large tasks down into smaller, more granular sizes to match the current optimal energy efficiency point.

[0064] By configuring the aforementioned various control parameters, coordinated intervention is achieved at three levels: system scheduling, hardware frequency, and application design. This enables multi-level and refined joint optimization, resulting in a better energy efficiency balance. Specifically, target utilization and task sharding are directly related to task completion time, while the frequency reduction magnitude is directly related to instantaneous power consumption. The combination of these three factors constitutes the direct control variables for achieving dynamic and precise balance, further improving the accuracy of power consumption control.

[0065] In this embodiment, the server system uses a microcontroller unit to collect real-time operating data from the accelerator card unit. The microcontroller unit is directly connected to the accelerator card unit, enabling high-speed acquisition of real-time operating data. Compared to the traditional method of data acquisition using a baseboard management controller, this improves data acquisition efficiency and real-time performance. After acquiring high-real-time operating data, the microcontroller unit in this embodiment dynamically generates control parameters adapted to the current operating state of the accelerator card unit using artificial intelligence, based on the real-time operating data and a preset control model. Compared to the traditional method of using fixed thresholds for control, this method better adapts to the current operating state of the accelerator card unit, achieving flexible and accurate power consumption control for different operating states. Through this approach, the real-time performance and accuracy of power consumption control for the accelerator card unit in the server system are improved, enabling timely and accurate power consumption control of the accelerator card unit.

[0066] Optionally, the baseboard management unit 13 is used to transmit control parameters between the microcontroller unit 12 and the control unit 14, thereby realizing transparent data transmission. Based on standard protocol communication, it can be adapted to mainstream server operating systems of different manufacturers and versions without modification, breaking the compatibility barrier caused by proprietary protocols. The system reuses the existing hardware and management interface of the baseboard management unit 13 on the server, without the need to install special drivers or agent programs in the server system. It is easy to deploy and has low operating costs. Communication isolation is achieved through the baseboard management unit 13, decoupling the real-time tasks of the microcontroller unit 12 from the complex system environment of the control unit 14, avoiding mutual interference between functions and improving the stability of the system.

[0067] Optionally, the substrate management unit 13 is mainly responsible for transmitting the data parameters calculated by AI to the OS.

[0068] Optionally, the baseboard management unit 13 can adopt a heterogeneous multi-core or master-slave multi-core BMC chip with a "main core + coprocessor" structure. This type of BMC can simultaneously, efficiently and reliably handle two completely different tasks. It can not only efficiently transmit the data sent by the microcontroller unit 12 to the control unit 14, but also return the data sent by the control unit 14 to the microcontroller unit 12, thereby improving the timeliness and accuracy of power consumption processing.

[0069] Optionally, the substrate management unit 13 is configured to: perform data encapsulation processing on the control parameters to obtain encapsulated data in a first preset data format; and send the encapsulated data to the control unit.

[0070] Optionally, the first preset data format here is a standardized server management protocol format.

[0071] Specifically, the first preset data format here can be JavaScript Object Notation (JSON) format conforming to the Redfish protocol standard. It is compatible with mainstream server OS and GPU manufacturers, allowing for compatibility with different devices without requiring hardware design modifications.

[0072] In one possible implementation, the baseboard management unit 13 receives raw control parameters from the microcontroller unit 12. The Redfish protocol module of the baseboard management unit 13, following the Redfish v1.10 schema definition, populates these parameters into a structured JSON object. A standards-compliant JSON document, parsable by any Redfish-enabled client, is generated and sent to the control unit 14 on the host operating system via Hypertext Transfer Protocol (HTTP) / Hypertext Transfer Protocol Secure (HTTPS).

[0073] Here, the baseboard management unit is used to standardize and mediate the communication protocol, encapsulate the control parameters into data, thereby obtaining encapsulated data that meets the transmission requirements and is compatible with a variety of control units, ensuring data transmission compatibility, applicable to a variety of server systems, and improving the flexibility and reliability of power consumption control.

[0074] Optionally, the substrate management unit 13 is configured to transmit packaged data to the control unit 14 according to the communication link pre-established by the control unit 14.

[0075] Optionally, the control unit 14 and the baseboard management unit 13 can interact through a communication interface based on a standard network protocol, including but not limited to a Keyboard Controller Style (KCS) interface or a LAN over Universal Serial Bus (LAN over USB) interface.

[0076] In one possible implementation, the baseboard management unit 13 comprises two main sub-modules: a Redfish protocol module and a data pass-through module.

[0077] Optionally, the Redfish protocol module encapsulates the control parameters output by the microcontroller unit 12 into the Redfish v1.10 standard JSON format to ensure cross-vendor OS compatibility.

[0078] Optionally, the data transmission module establishes a dedicated communication link between the substrate management unit 13 and the control unit 14, which can ensure that the parameter transmission delay is ≤10μs, and at the same time receive the execution result feedback from the control unit 14 to achieve accurate model optimization and precise control.

[0079] Through the aforementioned communication link, the substrate management unit and the control unit can achieve efficient data transmission, further improving the real-time performance of power consumption control.

[0080] Through a dedicated universal link, the embodiments of this application can ensure the reliability and determinism of communication, adapt to different server hardware and operating systems, improve compatibility, and both KCS and LAN over USB are part of the server out-of-band management network, which is physically or logically isolated from the in-band network that processes business data, improving the efficiency of data transmission and further improving the efficiency and reliability of power control.

[0081] Through the high-speed acquisition and high-speed data transmission link of the aforementioned microcontroller unit, the real-time response speed is fast, and the total latency of the control closed loop is <20µs. It can capture GPU microsecond-level load fluctuations, avoiding the lag and power waste of traditional solutions, and improving the efficiency of power peak suppression. Moreover, the additional power consumption is low, with the average power consumption of the microcontroller unit 12 being <5mW, which is far lower than the energy-saving benefits brought by GPU power control and will not increase the overall energy consumption burden of the system.

[0082] like Figure 2 As shown, Figure 2 Schematic diagram of the server system provided in the embodiments of this application Figure 2 ,like Figure 2 As shown, the microcontroller unit 12 includes a data acquisition module 121 and an inference module 122. The data acquisition module 121 and the inference module 122 are connected.

[0083] The acquisition module 121 is configured to acquire real-time operating data of the acceleration card unit 11 and transmit the real-time operating data to the inference module 122.

[0084] The inference module 122 is configured to: acquire real-time operating data and generate control parameters based on the preset control model and the real-time operating data.

[0085] Here, the microcontroller unit in this embodiment includes an acquisition module for performing real-time data acquisition and an inference module for performing inference. Through these two different modules, the microcontroller unit functions can be processed in parallel, improving data acquisition efficiency and data processing efficiency, increasing response speed and processing efficiency, further improving the real-time performance of power consumption control, and improving the operating performance of the server system.

[0086] Optionally, the microcontroller unit 12 is a dual-core or heterogeneous multi-core MCU chip.

[0087] Optionally, the microcontroller unit includes a first processor core and a second processor core that run in parallel; the second processor core is equipped with an acquisition module 121, which is configured to acquire real-time operating data; the first processor core is equipped with an inference module 122, which is configured to run a preset control model and generate control parameters.

[0088] Optionally, the first processor core and the second processor core are different types of processor cores, forming a heterogeneous multi-core architecture.

[0089] Optionally, the first processor core is the main core, and the second processor core is the auxiliary core.

[0090] Optionally, the first processor core runs the TinyCNN advanced model, while the second processor core is responsible for communication and interaction with multiple GPUs and BMCs. Tasks are partitioned and processed in parallel to solve the computing power bottleneck of a single core in a multi-GPU scenario.

[0091] The above configuration enables the rational use of hardware resources, improves resource utilization, and enhances the efficiency of power consumption control.

[0092] Optionally, the inference module 122 is configured as follows:

[0093] Obtain a preset control model; input real-time operating data into the preset control model to determine the control parameters based on the output of the preset control model.

[0094] The preset control model is obtained by training the initial control model using running data samples and the corresponding control parameter labels.

[0095] Here, the model in this application makes decisions based on the complex mapping relationship of "load-power consumption-performance", which breaks through the limitations of traditional solutions that rely on a single indicator or fixed threshold, and achieves dynamic and accurate balance, effectively reducing power consumption while ensuring the performance of the server system.

[0096] Optionally, a preset control model is deployed on the main core of the microcontroller unit 12.

[0097] Optionally, the instruction set and memory subsystem of the main core's digital signal processor (DSP) can be optimized to achieve microsecond-level inference, further improving the efficiency of power consumption control.

[0098] Optionally, the inference module 122 is configured as follows:

[0099] Obtain the initial control model; obtain the running data samples and the corresponding control parameters; train the initial control model based on the running data samples and the corresponding control parameters to obtain the preset control model.

[0100] Optionally, the model can be trained by fine-tuning to effectively save computing resources, improve power consumption control performance, and enhance the processing performance of the server system.

[0101] Alternatively, training can be performed using supervised learning methods on the cloud or a high-performance server.

[0102] Optionally, the running data sample is historical or simulated load data collected on the target GPU, which can be the multidimensional real-time running data in the above embodiments; the control parameter label is the optimal or better control parameter under the load obtained by optimization algorithms such as reinforcement learning or expert strategies.

[0103] Training objective: The training process aims to enable the model to learn the mapping relationship between complex load features and optimal control actions. The loss function can be the mean squared error.

[0104] By training the above model with targeted data, it is possible to train dedicated models for different types of accelerator card units or different typical load scenarios, thereby improving the system's adaptability and accuracy, and further enhancing the accuracy of power consumption control.

[0105] Optionally, the microcontroller unit 12 also includes a data cache module.

[0106] The data caching module is connected to the acquisition module 121 and the inference module 122 respectively.

[0107] The acquisition module 121 is configured as follows:

[0108] The system collects real-time operating data from the accelerator card unit and transmits the real-time operating data to the cache of the microcontroller unit 12.

[0109] The data caching module is configured as follows:

[0110] The system obtains real-time running data through caching; performs data quantization on the real-time running data to obtain quantized running data; and sends the quantized running data to the inference module.

[0111] Inference module 122 is configured as follows:

[0112] Based on the preset control model and the quantified operating data, control parameters are generated.

[0113] Optionally, quantization is performed using INT8.

[0114] In one possible implementation, the auxiliary core of the microcontroller unit 12 is mainly responsible for caching the acquired data, acquiring 6-dimensional data of the GPU through the PCIe interface, and transmitting it to the main core of the microcontroller unit 12 after completing INT8 quantization. The auxiliary core interacts with the baseboard management unit 13 through the USB bus, transmitting control parameters / receiving feedback data, monitoring the inference state of the main core of the microcontroller unit, and switching to the backup model when an anomaly is triggered to ensure system stability.

[0115] This application embodiment buffers high-speed data streams through a caching mechanism, avoiding data loss and timing discrepancies, providing stable and reliable data input for the model, and ensuring the correctness of control decisions. The raw data is converted to low-precision formats such as INT8 in real time, and data quantization preprocessing improves system processing efficiency, reduces internal data transmission bandwidth and memory usage, and achieves low inference latency and low power consumption operation.

[0116] Optionally, the substrate management unit 13 is configured to: receive the control result sent by the control unit 14; and send the control result to the inference module 122 so that the inference module 122 can incrementally learn the preset control model based on the control result.

[0117] Here, this application embodiment provides a result feedback function, and the inference module can perform optimization learning based on the feedback results, thereby improving the accuracy of the model and further improving the accuracy of power consumption control.

[0118] Optionally, the inference module 122 is configured to: receive the control results sent by the substrate management unit 13; perform filtering processing on the control results to obtain multiple valid control results; if the number of valid control results is greater than a preset incremental learning threshold, then perform incremental updates on the preset control model based on the multiple valid control results to obtain the updated control model.

[0119] It is understood that the preset incremental learning threshold here can be determined according to the actual situation, and the embodiments of this application do not impose specific restrictions on it.

[0120] Through incremental learning, the model can dynamically adapt to the load characteristics of different AI tasks and different GPU models, eliminating the need for offline retraining and reducing deployment costs.

[0121] The embodiments of this application perform filtering processing on the control results, which reduces the amount of data, effectively saves computing resources, and improves the accuracy of incremental model updates by filtering out inaccurate data.

[0122] Optionally, the inference module 122 is configured to: determine the running data samples and corresponding control parameter labels for incremental updates based on multiple valid control results; and perform incremental learning processing on the preset control model based on the running data samples and corresponding control parameter labels for incremental updates to update the weights of the fully connected layers of the preset control model, thereby obtaining the updated control model.

[0123] Here, by only updating the weights of the fully connected layers, accurate model optimization can be achieved with less computational resources, further improving power consumption control performance.

[0124] Optionally, the substrate management unit 13 is configured to: identify the operating system type of the control unit 14 that sends the control result; perform format conversion processing on the control result according to the operating system type to obtain the converted control result in a second preset format; and send the converted control result to the inference module 122 so that the inference module 122 can perform incremental learning on the preset control model according to the converted control result.

[0125] The second preset format here is determined by the operating system type, which is used to ensure compatibility with a wide variety of server hosts, improves the accuracy and compatibility of power consumption control, and enhances the stability and reliability of the server system.

[0126] In one possible implementation, the baseboard management unit 13 adds a "parameter adaptation field" to support automatic adjustment of parameter format according to the operating system type corresponding to the control unit 14, thereby improving compatibility.

[0127] In one possible implementation, the model incrementally learns and optimizes as follows:

[0128] After the microcontroller unit 12 accumulates a preset number of sets of valid data fed back by the control unit 14, it triggers an incremental learning process, which only updates the weights of the fully connected layer of the preset control model and does not affect real-time control.

[0129] In some embodiments, the weight parameters of the fully connected layer account for 40% of the model, and its optimization training time is ≤10ms. This ensures real-time control while improving control efficiency and guaranteeing the performance of power consumption control.

[0130] Optionally, the preset data here is the data after filtering outliers.

[0131] The preset quantity can be determined according to the actual situation, and this application embodiment does not impose specific restrictions on it.

[0132] In some embodiments, incremental learning employs a stochastic gradient descent (SGD) optimizer with a learning rate of 0.0005 and a mean squared error (MSE) loss function to ensure the model dynamically adapts to the load characteristics of different inference tasks, thereby improving optimization efficiency. While ensuring learning effectiveness, it boasts advantages such as low memory consumption and low computational complexity, perfectly adapting to the resource constraints of lightweight incremental learning on microcontroller units and guaranteeing real-time power consumption control.

[0133] Optionally, the microcontroller unit 12 is configured to:

[0134] Save the preset control model to the spare model area of ​​the non-volatile memory unit of the microcontroller unit 12; obtain the incremental learning result of the preset control model; if the incremental learning result is a successful update, save the updated control model to the main model area of ​​the non-volatile memory unit of the microcontroller unit 12; if the incremental learning result is a failed update, roll back to the preset control model according to the spare model area.

[0135] The spare model area is used for model rollback after incremental learning fails.

[0136] The main model area is used to perform inference of the control parameters.

[0137] Optionally, the non-volatile memory cell is the on-chip flash memory of the MCU or an integrated circuit used to implement the storage function.

[0138] Optionally, during the firmware design phase, the flash memory is divided into at least two logical regions:

[0139] Main model area: Stores the validated and effective control models currently used for real-time inference.

[0140] Backup model area: Stores a known stable model version that was successfully verified in the last time (i.e., the "preset control model" or the model that was successfully updated in the last time).

[0141] Here, if the model update fails or abnormal parameters are found in the optimized model during the model verification process, the model can be quickly rolled back through the backup model area, which enhances the fault tolerance of the system and ensures the stability of power consumption control.

[0142] Optionally, the acquisition module 121 is configured as follows:

[0143] Collect real-time operating data and corresponding encoding type data from the accelerator card unit, and transmit the real-time operating data and encoding type data to the inference module;

[0144] Inference module 122 is configured as follows:

[0145] Based on the coded data, determine the weight coefficients of the preset control model.

[0146] The encoding type data includes accelerator card unit model encoding and / or task type encoding.

[0147] In one possible implementation, the embodiments of this application can achieve multi-GPU / multi-task adaptation optimization:

[0148] During the GPU data acquisition phase, new features such as "GPU model encoding" and "task type encoding" are added (e.g., the first type of GPU model is encoded as 1, the second type of GPU model is encoded as 2; LLM inference is encoded as 1, and image rendering is encoded as 2). The preset control model dynamically adjusts the weight coefficients according to the encoding.

[0149] Traditional solutions require training and deploying different dedicated models for different GPU models or AI tasks, resulting in complex model management, high storage overhead, and inflexible scalability. The embodiments of this application, by introducing encoded type data into the input features, enable the same lightweight TinyCNN model to adapt to different scenarios, dynamically adjust its inference logic, and uniformly support multiple heterogeneous hardware and diverse computational loads with a single model, achieving global optimization with a single deployment. When a new GPU model or task is introduced, only the encoding needs to be extended and corresponding incremental learning performed, without changing the system architecture or deploying a new model. The system has extremely high scalability and improves power consumption control efficiency and accuracy.

[0150] To ensure the reliability of power consumption control and improve the stability and fault tolerance of the server system, this application embodiment is configured with a security mode. In the security mode, the power consumption of the accelerator card unit 11 is directly controlled by the microcontroller unit 12.

[0151] Optionally, the microcontroller unit 12 is configured to:

[0152] Monitor the real-time power consumption of the accelerator card unit 11; if the real-time power consumption is greater than the preset maximum power consumption threshold or less than the preset minimum power consumption threshold, send a first control command to the accelerator card unit 11 according to the preset safety parameters to control the operating status of the accelerator card unit 11.

[0153] Optionally, the microcontroller unit 12 is configured to:

[0154] The system monitors whether there is a transmission abnormality in the server system. If it is determined that there is a link abnormality in the server system, a second control command is sent to the acceleration card unit 11 according to the control parameters to control the operating status of the acceleration card unit 11.

[0155] Optionally, the microcontroller unit 12 is configured to:

[0156] The status parameters of the control unit 14 are collected, and based on the status parameters, it is determined whether the server system is experiencing overload of the control unit 14 and / or hardware failure of the control unit 14; and / or, a link detection request is sent to the control unit 14, and within a preset detection time, based on whether a response to the link detection request is heard, it is determined whether the server system is experiencing a link abnormality between the microcontroller unit 12 and the baseboard management unit 13 and / or a link abnormality between the baseboard management unit 13 and the control unit 14.

[0157] It is understood that the preset detection time, preset minimum power consumption threshold, and preset maximum power consumption threshold can all be determined according to the actual situation, and the embodiments of this application do not impose specific restrictions on them.

[0158] Optionally, the transmission channel for the first control command is a preset backup transmission link.

[0159] Optionally, the transmission channel for the second control command is a preset backup transmission link.

[0160] Optionally, the preset backup transmission link can reuse the data transmission link of the microcontroller unit 12 used to acquire the data of the accelerator card unit 11, or it can be another link between the microcontroller unit 12 and the accelerator card unit. The triggering method can be a dedicated sideband signal pin or a preset PCIe emergency message.

[0161] Optionally, after detecting a complete failure of the main intelligent control loop, the microcontroller unit 12 automatically switches to the aforementioned hardware emergency channel and sends pre-stored conservative control commands. The main intelligent control loop is the transmission link between the microcontroller unit 12, the substrate management unit 13, and the control unit 14.

[0162] Optionally, the microcontroller unit 12 continuously and actively detects the health status of the control unit 14, the substrate management unit 13, and each communication link through a heartbeat mechanism and status polling.

[0163] In one implementation, when the GPU power consumption exceeds the normal range (>400W or <70W), the TinyCNN model pauses inference, outputs preset safety parameters (e.g., target utilization 50%, frequency reduction 20%), and sends the preset safety parameters directly to the GPU to avoid hardware damage.

[0164] This application embodiment establishes a high-response hardware security firewall, providing final protection before software control completely fails, and absolutely preventing physical damage to the accelerator card due to overload.

[0165] The embodiments of this application realize system function degradation and fault isolation. When the advanced intelligent control fails, the system can automatically switch to the backup hardware control layer to ensure that the basic safety functions are not interrupted.

[0166] The embodiments of this application provide real-time self-diagnostic capabilities across the entire chain, enabling the system to predict and quickly locate faults, and significantly improving the system's observability, maintainability, and overall reliability.

[0167] Optionally, Figure 3 Schematic diagram of the server system provided in the embodiments of this application Figure 3 As shown in the figure

[0168] As shown in Figure 3, Figure 3 The control unit 14 is deployed in the host 31, which is connected to the accelerator card unit 11 and is used to configure the accelerator card unit 11 according to parameters. The accelerator card unit 11 here can be a GPU module.

[0169] Optionally, Figure 3 The microcontroller unit 12 includes a main core 32 and a secondary core 33. The secondary core 33 is used to connect to the accelerator card unit 11 to collect real-time operating data. The real-time operating system (PTOS) of the main core 32 is used to carry the data monitoring engine 320, which can be a GPU data monitoring engine (Engine). The Engine can be configured with a GPU data monitoring submodule, an AI inference submodule for carrying lightweight neural network models, and a parameter tuning module.

[0170] Figure 3 The two interactive lines between the intermediate baseboard management unit 13 and the host 31 represent the transmission of control parameters and the feedback of control results, respectively. These two interactive lines can be implemented through different connection links or through the same connection link.

[0171] It is understandable that the interaction lines indicated by the arrows in the diagram represent the direction of signal / data transmission. Any two interaction lines between two identical transmission entities can be implemented through different connection links or through the same connection link.

[0172] Optionally, the interaction lines indicated by the arrows in the diagram only represent part of the data / signal interaction methods. Figure 3 There are more ways to interact between any two units or modules.

[0173] Figure 4This is a flowchart illustrating the power consumption control method provided in this application embodiment. The execution entity of the power consumption control method provided in this application embodiment can be the aforementioned microcontroller unit 12, such as... Figure 4 As shown, embodiments of this application provide a power consumption control method, which is described in detail below:

[0174] 401. Collect real-time operating data of the accelerator card unit.

[0175] 402. Generate control parameters based on the preset control model and real-time operating data.

[0176] 403. Send the control parameters to the control unit so that the control unit can control the operating status of the accelerator card unit according to the control parameters.

[0177] Optionally, Figure 5 This application provides a functional diagram of each unit in a server system, as illustrated in the embodiments of this application. Figure 5 As shown, the microcontroller unit 12 includes a main core 32 and a secondary core 33. The microcontroller unit 12 implements inter-core division of labor, hardware acceleration, and lightweight deployment. The main core 32 and the secondary core 33 communicate with each other via a mailbox mechanism. The microcontroller unit 12 also communicates with the baseboard management unit 13 via a mailbox mechanism. The baseboard management unit 13 is used for protocol conversion, data transmission, and status monitoring. The baseboard management unit 13 is mainly used for parameter transmission. The baseboard management unit 13 communicates with the control unit 14 via the Redfish protocol. The control unit 14 serves as the execution terminal and feedback hub, used for control.

[0178] In one possible implementation, combining Figure 5 The specific processing stages of the power consumption control method are as follows:

[0179] Data acquisition phase: The GPU data monitoring engine collects the GPU's 6-dimensional core data in real time through the PCIe interface, formats it according to the tuning parameters .json, and caches it in the MCU's on-chip random access memory (RAM).

[0180] AI inference stage: The TinyCNN model in the main core loads the collected data and, based on the trained "load-power-performance" mapping relationship, infers and generates 3D control parameters. The parameter output format is INT8, which occupies less storage space and has low inference latency.

[0181] Parameter pass-through stage: The auxiliary core sends the control parameters to the BMC via the USB bus. The Redfish protocol module in the BMC encapsulates the parameters into standard JSON format and passes them through the Ethernet to the OS control unit.

[0182] Execution phase: The OS control unit parses parameters, calls the GPU driver to adjust load distribution and reduce core frequency by 15%, achieving precise control of GPU utilization.

[0183] Feedback phase: The OS sends the actual execution results (78% utilization, 235W power consumption, and 0.2s increase in inference latency) back to the BMC, which then synchronizes them to the MCU for incremental learning optimization of the TinyCNN model.

[0184] Through the above embodiments, the embodiments of this application can achieve the following technical effects:

[0185] Heterogeneous architecture collaboration: Proposes a heterogeneous architecture of "MCU integrated dedicated monitoring engine + BMC protocol pass-through" to achieve microsecond-level data acquisition and low-latency parameter transmission, solving the pain point of slow response in traditional solutions.

[0186] AI-driven precise control: Employing a lightweight TinyCNN model, control parameters are generated based on 6-dimensional multi-feature inference, breaking through the traditional control logic of "single indicator + fixed threshold" and achieving a dynamic balance of "load-power consumption-performance".

[0187] Redfish Protocol Standardization Adaptation: AI control parameters are encapsulated in Redfish standard JSON format for transparent transmission, solving the problems of "closed protocols and poor compatibility" in existing technologies, and supporting fine-grained parameter transmission.

[0188] Online incremental learning: The MCU supports online lightweight incremental learning of the TinyCNN model, dynamically optimizing mapping relationships, adapting to different scenarios, and improving the long-term accuracy of control.

[0189] It should be noted that the functions that the above-mentioned server system can achieve are all applicable to the power consumption control method in the embodiments of this application, and will not be elaborated here.

[0190] For a description of the features in the embodiment corresponding to the power consumption control method, please refer to the relevant description of the embodiment corresponding to the server system, which will not be repeated here.

[0191] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0192] Figure 6 This is a schematic diagram of the power consumption control device provided in an embodiment of this application. Figure 6 As shown, embodiments of this application also provide a power consumption control device 60, including:

[0193] The acquisition module 601 is used to collect real-time operating data of the accelerator card unit;

[0194] The generation module 602 is used to generate control parameters based on the preset control model and real-time operating data;

[0195] The sending module 603 is used to send the control parameters to the control unit so that the control unit can control the operating status of the accelerator card unit according to the control parameters.

[0196] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes at least one processor 701 and a memory 702. Optionally, the electronic device 70 further includes a communication component 703. The processor 701, memory 702, and communication component 703 are connected via a bus.

[0197] In a specific implementation, at least one processor 701 executes computer execution instructions stored in memory 702, causing at least one processor 701 to execute the above-described power consumption control method embodiment.

[0198] The specific implementation process of processor 701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0199] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0200] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0201] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0202] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described power consumption control method embodiments when running.

[0203] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0204] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described power consumption control method embodiments.

[0205] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described power control method embodiments.

[0206] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0207] The power consumption control method provided in this application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A server system, characterized in that, include: Accelerator card unit; A microcontroller unit is connected to the accelerator card unit. The microcontroller unit is configured to: collect real-time operating data of the accelerator card unit and generate control parameters based on a preset control model and the real-time operating data. A substrate management unit is connected to the microcontroller unit, and the substrate management unit is configured to receive control parameters sent by the microcontroller unit. A control unit is connected to the substrate management unit, and the control unit is configured to receive control parameters sent by the substrate management unit. The operating state of the accelerator card unit is adjusted according to the control parameters.

2. The server system according to claim 1, characterized in that, The substrate management unit is configured as follows: The control parameters are encapsulated to obtain encapsulated data in a first preset data format; The encapsulated data is sent to the control unit.

3. The server system according to claim 2, characterized in that, The substrate management unit is configured as follows: The packaged data is transmitted to the control unit according to the pre-established communication link between the substrate management unit and the control unit.

4. The server system according to any one of claims 1 to 3, characterized in that, The microcontroller unit is configured to: Monitor the real-time power consumption of the accelerator card unit; If the real-time power consumption is greater than the preset maximum power consumption threshold or less than the preset minimum power consumption threshold, then according to the preset safety parameters, a first control command is sent to the accelerator card unit to control the operating state of the accelerator card unit.

5. The server system according to any one of claims 1 to 3, characterized in that, The microcontroller unit is configured to: Monitor the server system for any transmission anomalies; If a link anomaly is detected in the server system, a second control command is sent to the accelerator card unit according to the control parameters to control the operating status of the accelerator card unit.

6. The server system according to claim 5, characterized in that, The microcontroller unit is configured to: Collect the status parameters of the control unit, and determine whether the server system is experiencing overload of the control unit and / or hardware failure of the control unit based on the status parameters; And / or, A link detection request is sent to the control unit, and within a preset detection time, based on whether a response to the link detection request is received, it is determined whether the server system experiences a link anomaly between the microcontroller unit and the baseboard management unit and / or a link anomaly between the baseboard management unit and the control unit.

7. The server system according to any one of claims 1 to 3, characterized in that, The microcontroller unit includes a data acquisition module and an inference module; the data acquisition module and the inference module are connected. The acquisition module is configured to: acquire the real-time operating data of the accelerator card unit and transmit the real-time operating data to the inference module; The inference module is configured to: acquire the real-time operating data, and generate control parameters based on the preset control model and the real-time operating data.

8. The server system according to claim 7, characterized in that, The inference module is configured as follows: Obtain a preset control model; wherein the preset control model is obtained by training an initial control model using running data samples and the control parameter labels corresponding to the running data samples; The real-time operating data is input into the preset control model to determine the control parameters based on the output of the preset control model.

9. The server system according to claim 8, characterized in that, The inference module is configured as follows: Obtain the initial regulation model; Obtain the operational data sample and the corresponding control parameters of the operational data sample; The initial control model is trained based on the operational data sample and the corresponding control parameters to obtain the preset control model.

10. The server system according to claim 9, characterized in that, The substrate management unit is configured as follows: Receive the control result sent by the control unit; The control results are sent to the inference module so that the inference module can incrementally learn the preset control model based on the control results.

11. The server system according to claim 10, characterized in that, The inference module is configured as follows: Receive the control results sent by the substrate management unit; The control results are then filtered to obtain multiple effective control results; If the number of effective control results is greater than the preset incremental learning threshold, then based on the multiple effective control results, an incremental update is performed on the preset control model to obtain the updated control model.

12. The server system according to claim 11, characterized in that, The inference module is configured as follows: Based on the multiple effective control results, determine the operating data samples and corresponding control parameter labels for incremental updates; Based on the running data samples used for incremental updates and the corresponding control parameter labels, the preset control model is subjected to incremental learning processing to update the weights of the fully connected layers of the preset control model, thereby obtaining the updated control model.

13. The server system according to claim 10, characterized in that, The substrate management unit is configured as follows: Identify the operating system type of the control unit that sends the control result; Based on the operating system type, the control result is converted to a new format to obtain the converted control result in a second preset format. The converted control result is sent to the inference module so that the inference module can incrementally learn the preset control model based on the converted control result.

14. The server system according to any one of claims 10 to 13, characterized in that, The microcontroller unit is configured to: The preset control model is saved to the spare model area of ​​the non-volatile storage unit of the microcontroller; wherein, the spare model area is used for model rollback after incremental learning failure; Obtain the incremental learning results of the preset control model; If the incremental learning result is a successful update, the updated control model is saved to the main model area of ​​the non-volatile storage unit of the microcontroller; wherein, the main model area is used to perform the inference of the control parameters; If the incremental learning result is an update failure, then the system is rolled back to the preset control model according to the backup model area.

15. The server system according to any one of claims 8 to 13, characterized in that, The acquisition module is configured as follows: The system collects real-time operating data and corresponding encoding type data of the accelerator card unit, and transmits the real-time operating data and the encoding type data to the inference module; wherein, the encoding type data includes the accelerator card unit model code and / or task type code; The inference module is configured as follows: Based on the encoded data, the weight coefficients of the preset control model are determined.

16. The server system according to claim 7, characterized in that, The microcontroller unit also includes a data cache module; the data cache module is connected to the acquisition module and the inference module respectively; The acquisition module is configured as follows: The system collects real-time operating data from the accelerator card unit and transmits the real-time operating data to the cache of the microcontroller unit. The data caching module is configured as follows: The real-time running data is obtained through the cache; The real-time running data is subjected to data quantization processing to obtain quantized running data; The quantized running data is sent to the inference module; The inference module is configured as follows: Based on the preset control model and the quantized operating data, control parameters are generated.

17. The server system according to claim 16, characterized in that, The preset control model is a lightweight convolutional neural network model after quantization processing.

18. The server system according to any one of claims 1 to 3, characterized in that, The real-time operating data includes at least one of the following: accelerator card unit real-time utilization rate, tensor core active percentage, core voltage, real-time power consumption, load fluctuation frequency, and task type.

19. The server system according to claim 18, characterized in that, The control parameters include at least one of the following: target utilization rate, core frequency reduction magnitude, and task fragmentation threshold.

20. A power consumption control method, characterized in that, The method, applied to a microcontroller unit of a server system as described in any one of claims 1 to 19, comprises: Collect real-time operating data from the accelerator card unit; Based on the preset control model and the real-time operating data, control parameters are generated; The control parameters are sent to the control unit so that the control unit can control the operating state of the accelerator card unit according to the control parameters.

Citation Information

Patent Citations

  • Server power consumption dynamic optimization and cooperative heat dissipation control system based on AI

    CN121326537A

  • Server monitoring method and system adaptive to power consumption regulation and control of domestic server

    CN121501603A