Voltage determination method of processor, electronic equipment, storage medium and product
By training the load prediction model to dynamically adjust the processor voltage, the energy loss problem caused by high voltage at low load is solved by traditional processors, and more efficient energy management is achieved.
Patent Information
- Application Number
- CN202510964322.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The voltage determination technology of traditional processors relies on fixed threshold control, lacks dynamic response to actual load changes, and maintains high voltage in low load scenarios, resulting in high energy losses.
By obtaining the operating status information and historical load information of the server cluster, the load prediction model is trained, the target model is used to determine the operating voltage of the processor, and the voltage is dynamically adjusted to adapt to load changes.
It realizes the avoidance of over-power supply in low-load scenarios, reduces energy losses, and improves energy utilization efficiency.
Smart Images

Figure CN120469562A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of server technology, and in particular to a method for determining processor voltage, an electronic device, a storage medium, and a product. Background Art
[0002] Related processor voltage determination schemes typically use fixed threshold-based voltage regulation strategies, which lack the ability to respond in real time to dynamic changes in the actual load of server clusters. Even in low-load scenarios, the processor may still maintain a high operating voltage, resulting in high energy loss. Summary of the Invention
[0003] The present application provides a processor voltage determination method, electronic device, storage medium and product to at least solve the problem in the related art of relying on fixed threshold control, lacking dynamic response to actual load changes, maintaining high voltage in low-load scenarios, and resulting in high energy loss.
[0004] The present application provides a method for determining a voltage of a processor, comprising: Obtaining first operating status information of the server cluster; Training a load prediction model based on the first operating status information and historical load information of the server cluster to obtain a target model; Acquire second operating status information of the server cluster, and determine a target load corresponding to the second operating status information using the target model; An operating voltage of the processor is determined according to the target load.
[0005] The present application also provides a processor voltage determination device, comprising: An acquiring unit, configured to acquire first operating status information of a server cluster; a training unit, configured to train a load prediction model based on the first operating status information and historical load information of the server cluster to obtain a target model; a first determining unit, configured to obtain second operating status information of the server cluster, and determine a target load corresponding to the second operating status information by using the target model; The second determining unit is configured to determine an operating voltage of the processor according to the target load.
[0006] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned processor voltage determination methods when executing the computer program.
[0007] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods for determining the voltage of the processor are implemented.
[0008] The present application also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned methods for determining the voltage of a processor when the computer program is executed by a processor.
[0009] Through this application, the first operating status information of the server cluster is obtained; based on the first operating status information and the historical load information of the server cluster, the load prediction model is trained to obtain a target model; the second operating status information of the server cluster is obtained, and the target model is used to determine the target load corresponding to the second operating status information; according to the target load, the operating voltage of the processor is determined, which solves the technical problem that the voltage determination technology of the traditional processor relies on fixed threshold control, lacks dynamic response to actual load changes, maintains high voltage in low load scenarios, and leads to high energy loss, thereby achieving the technical effect of avoiding overpowering and reducing energy loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1 A flowchart of a method for determining voltage of a processor provided in an embodiment of the present application; Figure 2 A schematic diagram of a server-based energy recovery and dynamic power management system provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a voltage determination device for a processor provided in an embodiment of the present application. DETAILED DESCRIPTION
[0012] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0013] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0014] In order to facilitate those skilled in the art to better understand the technical solutions described in the embodiments of the present disclosure, the technical terms in the embodiments of the present disclosure are explained as follows before introducing the embodiments of the present disclosure.
[0015] Computing power has leapt forward, doubling the scalability of central processing units (CPUs) and overcoming the constraints of existing software and hardware resources at an unprecedented rate. This is the result of a combination of factors, with the rise of artificial intelligence (AI) and the surge in demand for computing power playing a key role. First, the advent of a multidimensional era of virtual and real life has led to a surge in computing power. The rapid development of AI technology, particularly the widespread application of deep learning, has exponentially increased the demand for computing power. Training larger neural network models and implementing more complex algorithms and applications requires more powerful computing power. Algorithms and computing power mutually reinforce each other. Algorithm innovation and optimization often require more powerful computing power, while increased computing power in turn drives algorithm innovation and application. This mutually reinforcing relationship creates a virtuous cycle between the development of AI technology and the improvement of computing power. Second, with the continuous advancement of cutting-edge technologies such as quantum computing and photonic computing, one-dimensional computing architectures are no longer sufficient, and entirely new computing models will emerge in the future. These new computing technologies offer higher computational efficiency and lower energy consumption, potentially breaking through the bottlenecks of existing computing technologies. Thirdly, the introduction of heterogeneous computing architectures enables the CPU to collaborate with other processor types, such as graphics processing units (GPUs), data processing units (DPUs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs), achieving more efficient computing. These processors each have their own strengths and can be flexibly scheduled and combined based on task characteristics, further improving overall server performance and scalability.
[0016] Current server systems face a triple challenge when running tasks such as large-scale model training and real-time inference: energy utilization and heat loss. First, there's the energy efficiency bottleneck: traditional dynamic power management (DPM) technology relies on fixed threshold control, resulting in over 30% ineffective power loss. Second, there's thermal management imbalance: GPU clusters can experience instantaneous heat flux densities of up to 300W / cm², while traditional liquid cooling systems experience temperature fluctuations of ±5°C. Finally, resource fragmentation: power supply, cooling, and computing units lack coordinated optimization, resulting in an overall Power Usage Effectiveness (PUE) rating generally exceeding 1.5.
[0017] The following briefly introduces three solutions in related technologies: Solution A dynamically monitors the real-time load of heterogeneous computing nodes, such as CPUs, GPUs, and FPGAs, and optimizes task allocation strategies using reinforcement learning algorithms to maximize computing resource utilization. It automatically adjusts processor frequency ranges based on task priority (e.g., the high-performance range of 2.4-3.8 GHz for matrix operations and the energy-efficient range of 1.2-2.0 GHz for preprocessing tasks), reducing energy consumption by 10-15%.
[0018] Solution B integrates heterogeneous data from multiple sources, including power load, meteorological, and economic data, and uses transfer learning to extract deep features to improve forecast accuracy. The hybrid forecasting model combines ARIMA periodicity correction with elastic network regression to address nonlinear issues in medium- and long-term forecasts. An adaptive particle swarm algorithm adjusts feature weights in real time, enabling minute-level forecast responses. For the first time, nonlinear temperature correction is introduced into power load forecasting, enhancing model robustness in extreme weather conditions.
[0019] Solution C embeds a bismuth telluride / graphene composite thermoelectric module in the battery pack, achieving gradient thermal energy capture at 75-85°C (primary), 55-65°C (secondary), and 40-50°C (tertiary), with an overall conversion efficiency of 12%. It prioritizes power to the 48V DC bus through a bidirectional (Direct Current / Direct Current, DC / DC) conversion circuit, and connects to the flywheel energy storage system through a suboptimal path. This is the first time that thermoelectric recovery and active heat dissipation functions have been integrated within the battery pack, extending battery life and improving energy recycling rates.
[0020] The above scheme still has the following defects: Solution A uses a dynamic scheduling strategy based on reinforcement learning. However, its load forecasting model relies solely on a linear weighting of historical loads and fails to account for the nonlinear effects of sudden changes in ambient temperature on load. For example, current temperature fluctuations rely on offline co-simulation using COMSOL and ANSYS (a single simulation takes >30 minutes), making it impossible to implement online dynamic optimization of the scheduling strategy and struggling to cope with sudden load spikes.
[0021] Solution B uses a convolutional neural network for transfer learning, but fails to optimize the network structure for the timing characteristics of power loads, resulting in insufficient capture of the temporal and spatial correlations between the data and historical load data. The patent also lacks a real-time fuse mechanism for sudden load changes, relying solely on post-fault corrections, which can lead to cascading failures.
[0022] Solution C uses a bismuth telluride / graphene composite module, but the conversion efficiency is only 3-5% in low temperature difference scenarios (ΔT<15°C), and low-grade thermal energy cannot be effectively recovered. When the battery cell temperature is >85°C, the liquid cooling system is triggered, but predictive thermal management is not integrated.
[0023] In order to solve the above technical problems, the present application obtains the first operating status information of the server cluster; trains the load prediction model based on the first operating status information and the historical load information of the server cluster to obtain a target model; obtains the second operating status information of the server cluster, and uses the target model to determine the target load corresponding to the second operating status information; determines the operating voltage of the processor according to the target load, thereby solving the technical problem that the voltage determination technology of the traditional processor relies on fixed threshold control, lacks dynamic response to actual load changes, maintains high voltage in low load scenarios, and leads to high energy loss, thereby achieving the technical effect of avoiding overpowering and reducing energy loss.
[0024] An embodiment of the present application provides a method for determining the voltage of a processor. The method can be applied to high-energy consumption scenarios such as ultra-large-scale artificial intelligence training, cloud computing platforms, and edge computing nodes. The method can be executed by a power management unit or controller built into the processor.
[0025] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0026] Figure 1 A flowchart of a method for determining processor voltage provided by an embodiment of the present disclosure is provided.
[0027] like Figure 1 As shown, the method comprises the following steps: Step 101: Obtain first operating status information of a server cluster; In some embodiments, a server cluster refers to a computing resource pool consisting of multiple servers. Server clusters are typically used to process large-scale parallel tasks such as AI training and big data analysis.
[0028] In some embodiments, the first operating status information refers to the server cluster operating data collected in real time by the edge computing node, including the server load Pload(t) at each moment, the processor temperature Tcore(t) at each moment, CPU utilization, network throughput, etc.
[0029] In some embodiments, the first operating status information of the server cluster collected by the edge computing node can be obtained through a monitoring system, such as a baseboard management controller (BMC) or an operating system interface.
[0030] Step 102: training a load prediction model based on the first operating status information and historical load information of the server cluster to obtain a target model; In some embodiments, the historical load information of the server cluster refers to the amount of tasks or computing load records carried by the server cluster in the past period of time, such as the CPU / GPU load at the past n time points.
[0031] In some embodiments, time series features including mean, variance, and trend items may be extracted from historical load information of the server cluster.
[0032] In some embodiments, the load prediction model is used to predict future load change trends based on input operating status information. The target model refers to a load prediction model that has been trained and optimized and has good generalization ability and prediction accuracy.
[0033] In some embodiments, the first operating status information and the historical load information of the server cluster can be preprocessed and used as input features of the load prediction model, wherein the preprocessing includes at least normalization processing and outlier processing to improve the efficiency and accuracy of load prediction model training.
[0034] Step 103: Acquire second operating status information of the server cluster, and determine a target load corresponding to the second operating status information using a target model; In some embodiments, the second operating status information of the server cluster refers to the operating status of the server cluster obtained at the current moment.
[0035] In some embodiments, the target load corresponding to the second operating state information refers to the amount of tasks that the processor needs to process in the current or short future time period predicted based on the second operating state information and the target model.
[0036] Step 104: Determine the operating voltage of the processor according to the target load.
[0037] In some embodiments, the operating voltage of a processor refers to the voltage value provided to the core circuit of the processor. The magnitude of the voltage value directly affects the power consumption and performance of the processor. The higher the voltage, the better the performance, but also the greater the energy consumption.
[0038] In some embodiments, based on the target load, the operating voltage of the processor can be determined from the mapping relationship between the server cluster load and the voltage, wherein the mapping relationship between the server cluster load and the voltage can be expressed in the form of a mapping relationship table or in the form of a mapping relationship array, which is not limited in this application.
[0039] In some embodiments, after determining the operating voltage of the processor according to the target load, the operating voltage of the processor may be dynamically adjusted using a dynamic voltage and frequency scaling (DVFS) mechanism.
[0040] In some embodiments, by dynamically adjusting the processor voltage based on the server cluster operating status and load prediction model, by sensing the cluster status, training the prediction model, inferring the load in real time and accurately controlling the voltage, on-demand energy supply and energy saving and consumption reduction can be achieved.
[0041] Through this application, the first operating status information of the server cluster is obtained; based on the first operating status information and the historical load information of the server cluster, the load prediction model is trained to obtain a target model; the second operating status information of the server cluster is obtained, and the target model is used to determine the target load corresponding to the second operating status information; according to the target load, the operating voltage of the processor is determined, which solves the technical problem that the voltage determination technology of the traditional processor relies on fixed threshold control, lacks dynamic response to actual load changes, maintains high voltage in low load scenarios, and leads to high energy loss, thereby achieving the technical effect of avoiding overpowering and reducing energy loss.
[0042] In some embodiments, training a load prediction model based on the first operating status information and the historical load information of the server cluster to obtain a target model includes: Obtaining ambient temperature information of the server cluster; In some embodiments, the ambient temperature information of the server cluster refers to the external ambient temperature of the computer room or cabinet where the server is located, which can usually be obtained through temperature and humidity sensors deployed in the rack or room to reflect the heat dissipation conditions of the equipment in the server cluster.
[0043] In some embodiments, multiple high-precision temperature and humidity sensors can be deployed in a server cabinet or computer room to collect ambient temperature data through monitoring systems such as the Intelligent Platform Management Interface (IPMI) and BMC. The data frequency can be from every 5 seconds to once per minute to ensure timing consistency.
[0044] determining training data based on the first operating state information, historical load information of the server cluster, and ambient temperature information of the server cluster; In some embodiments, based on the first operating status information, the historical load information of the server cluster, and the ambient temperature information of the server cluster, the three types of data are aligned by timestamp to construct training data with a unified time series.
[0045] The load prediction model is trained using the training data to obtain a target model.
[0046] In some embodiments, the load prediction model can be a random forest model or an eXtreme Gradient Boosting (XGBoost) model. The aforementioned training data is divided into a training set, a validation set, and a test set. After determining the loss function, the model is trained by minimizing the loss and adjusting the parameters to obtain the final target model, where the loss function can be a cross entropy loss function, etc.
[0047] In some embodiments, the ambient temperature information of the server cluster is obtained; training data is determined based on the first operating status information, the historical load information of the server cluster, and the ambient temperature information of the server cluster; the load prediction model is trained using the training data to obtain a target model, the ambient temperature information of the server cluster is introduced, and a training set is constructed in combination with the operating status and historical load data to train a load prediction model with higher prediction accuracy and generalization capability, thereby achieving a more intelligent and efficient processor voltage control strategy, with significant energy saving and system stability improvement effects.
[0048] In some embodiments, using the training data to train the load prediction model to obtain a target model includes: Using the training data, training the load prediction model; Determine a first error for each training session, terminate the training session when the first error satisfies a preset condition, and obtain a target parameter combination, wherein the first error is a prediction error, and the target parameter combination includes a weight coefficient and a temperature correction coefficient; Based on the target parameter combination, a target model is obtained.
[0049] In some embodiments, the mathematical expression of the load prediction model is as follows:
[0050] in, P i history Indicates the historical load of the server cluster, α i is the weight coefficient, β k is the temperature correction coefficient, T k is the kth ambient temperature information of the server cluster, m represents the number of training times, and n is the number of historical loads. Indicates the predicted load value.
[0051] In some embodiments, the first error refers to the deviation between the model's predicted output and the actual output, and is typically measured using metrics such as mean squared error (MSE) and mean absolute error (MAE). In this application, the mathematical expression for determining the first error is as follows:
[0052] in, represents the predicted load value at the tth iteration training, It represents the actual load value during the t-th iteration training, and m represents the number of training times.
[0053] In some embodiments, terminating training when the first error satisfies a preset condition includes: terminating training when the first error satisfies a preset error threshold, or terminating training when the first error satisfies a preset number of iterations, or terminating training when the first error does not significantly decrease after several consecutive rounds.
[0054] In some embodiments, the target parameter combination refers to a set of optimal model parameters obtained after the load prediction model training is completed, including weight coefficients of each feature in the model and a temperature correction coefficient for adjusting temperature effects.
[0055] In some embodiments, the target model determined based on the aforementioned target parameter combination can be deployed in an actual system for online reasoning and voltage control decision-making.
[0056] In some embodiments, the load prediction model is trained using the training data; the first error of each training is determined, and the training is terminated when the first error meets a preset condition to obtain a target parameter combination, wherein the first error is a prediction error, and the target parameter combination includes a weight coefficient and a temperature correction coefficient; based on the target parameter combination, a target model is obtained, and the model training process is dynamically terminated by introducing an error control mechanism to obtain a target parameter combination including a weight coefficient and a temperature correction coefficient, thereby constructing a target model with high precision, strong generalization capability and support for temperature perception, which can provide data for subsequent processor voltage adjustment.
[0057] In some embodiments, determining a first error for each training session, terminating the training session when the first error satisfies a preset condition, and obtaining a target parameter combination includes: Determine a first error and a second error for each training, where the second error is a regularization error; In some embodiments, the regularization error refers to a penalty term for the model error, which is used to prevent overfitting and improve the generalization ability of the model.
[0058] Based on the first error and the second error of each training, a target error of each training is determined, and the training is terminated when the target error meets a preset condition to obtain a target parameter combination.
[0059] In some embodiments, the target error refers to the weighted sum of the prediction error and the regularization error, which is used to determine whether the model training is terminated.
[0060] In some embodiments, based on the first error and the second error of each training, a target error for each training is determined. The mathematical expression of the target error is as follows:
[0061] in, λ is the penalty coefficient, represents the predicted load value at the tth iteration training, It represents the actual load value during the t-th iteration training, and m represents the number of training times.
[0062] In some embodiments, the aforementioned weight coefficient and temperature correction coefficient can be adjusted dynamically. In practical applications, the weight coefficient will decay over time, so a higher weight can be given to recent data. For the temperature correction coefficient, it can be determined according to the first few terms of the Lagrange polynomial expansion, such as β k ×T k The product of can be expanded to a polynomial β1×T1+β2×T2 to capture the nonlinear relationship between temperature and load.
[0063] In some embodiments, by determining the first error and the second error of each training, the second error is the regularization error; based on the first error and the second error of each training, the target error of each training is determined, and the training is terminated when the target error meets the preset conditions to obtain the target parameter combination. By introducing a dual evaluation mechanism of prediction error and regularization error during the training process, the target error is constructed as the basis for terminating the model training, and finally the target parameter combination including the weight coefficient and the temperature correction coefficient is obtained, which can improve the generalization ability of the load prediction model.
[0064] In some embodiments, determining the operating voltage of the processor according to the target load includes: Acquire a frequency domain list of the processor, the frequency domain list comprising at least one of a high-performance domain, an energy-efficiency domain, a standby domain, and an offline domain; In some embodiments, the frequency domain list refers to a set of frequency operating ranges supported by the processor, each range corresponding to a specific performance level and power consumption level, among which the high-performance domain (2.4-3.8GHz) is used to execute matrix operation cores; the energy efficiency domain (1.2-2.0GHz) is used to process data preprocessing tasks; the standby domain (0.8GHz) is used to maintain cache consistency; the offline domain (<0.5V) is used for hardware-level power-off isolation, reducing energy consumption but still maintaining responsiveness. The frequency domains can also be divided more and more finely according to actual needs.
[0065] In some embodiments, the frequency domain list of the processor may be queried and obtained through an operating system interface.
[0066] determining a target frequency domain from the frequency domain list according to a load requirement of the processor; An operating voltage of the processor in the target frequency domain is determined according to the target load.
[0067] In some embodiments, the target frequency domain refers to the most appropriate frequency operating range selected according to the current processor load requirement.
[0068] In some embodiments, the mathematical expression for determining the processor operating voltage is as follows:
[0069] in, V core Indicates the operating voltage of the processor (i.e. core voltage), P max Indicates the maximum load of the server cluster, P predict It represents the predicted load value of the server cluster. Furthermore, the slope of 0.25 in the formula can be adjusted based on the results of multiple tests.
[0070] In some embodiments, by obtaining a frequency domain list of the processor; determining a target frequency domain from the frequency domain list according to the load requirements of the processor; and determining the operating voltage of the processor in the target frequency domain according to the target load, by defining a frequency domain list of the processor and dynamically selecting a suitable target frequency domain and its corresponding voltage according to the load requirements, a refined power management strategy is implemented, which is particularly suitable for intelligent energy efficiency control scenarios in data centers, AI acceleration platforms, and edge computing devices.
[0071] In some embodiments, after determining the operating voltage of the processor according to the target load, the method further includes: In response to an operating voltage of the processor being different from an initial voltage of the processor, determining a temperature difference of the processor; In some embodiments, the initial voltage of the processor refers to a reference voltage value of the processor in a standard or default operating mode.
[0072] In some embodiments, the operating voltage of the processor being different from the initial voltage of the processor indicates that the voltage of the processor has changed.
[0073] In some embodiments, the temperature difference refers to the difference between the actual temperature of the processor caused by the power consumption difference caused by the voltage change and the temperature of the processor in the reference state. The temperature difference of the processor can be determined by comparing the temperatures corresponding to different voltages.
[0074] The temperature difference is converted into corresponding target electrical energy using thermoelectric materials, and the target electrical energy is recovered to the power supply system.
[0075] In some embodiments, the thermoelectric material is a bismuth telluride / graphene composite thermoelectric material.
[0076] In some embodiments, the target electrical energy refers to the reusable electrical energy output generated by the thermoelectric material using the temperature difference, which can be used to supplement energy in the power supply system.
[0077] In some embodiments, the power supply system refers to a system in a server, GPU cluster, or data center that is responsible for providing power support, including a power supply, an uninterruptible power supply (UPS), a battery, and the like.
[0078] In some embodiments, thermoelectric material is attached to the surface of the processor heat dissipation module, with the two ends of the material contacting the high-temperature area (processor) and the low-temperature area (heat sink) respectively. The temperature difference drives the generation of current, and the output DC power is connected to the power supply system after passing through the DC / DC boost module.
[0079] In some embodiments, the module for thermoelectric power generation can be determined based on the temperature difference. The modules for thermoelectric power generation include the primary module: CPU / GPU heat sink surface (the hot end temperature (Thot) of the thermoelectric material is 75 to 85°C), the secondary module: the outer wall of the liquid cooling pipe (Thot is 55 to 65°C), and the tertiary module: the cabinet exhaust duct (Thot is 40 to 50°C). In some embodiments, the temperature difference of the processor is determined in response to the fact that the operating voltage of the processor is different from the initial voltage of the processor; the temperature difference is converted into corresponding target electrical energy using thermoelectric materials, and the target electrical energy is recovered to the power supply system. By sensing the temperature difference caused by the voltage change of the processor, the waste heat is converted into usable electrical energy using thermoelectric materials, and the electrical energy is recovered to the power supply system, thereby achieving energy reuse and energy-saving optimization of the system.
[0080] In some embodiments, recycling the target electric energy to the power supply system includes: The target electric energy is recovered to the power supply system according to a preset recovery path, wherein the preset recovery path includes at least one of a DC bus, a flywheel energy storage system, and a backup battery pack.
[0081] In some embodiments, a bidirectional DC / DC conversion circuit may be used to recover the target electrical energy to the power supply system.
[0082] In some embodiments, the DC bus refers to the high-voltage DC distribution trunk line commonly found in data centers or servers, typically a 48V or 380V DC system, used to efficiently distribute electrical energy; a flywheel energy storage system is a physical energy storage device that stores kinetic energy through a high-speed rotating rotor and converts it into electrical energy when needed; a backup battery pack can be a chemical energy storage device in a UPS system, and the DC bus is usually used preferentially.
[0083] In some embodiments, by recycling the target electrical energy converted from waste heat to the power supply system according to a preset path, energy reuse and intelligent coordinated optimization of the power supply system are achieved.
[0084] In some embodiments, as Figure 2 As shown, Figure 2This is a schematic diagram of a server-based energy recovery and dynamic power management system provided in an embodiment of the present application. It constructs a three-layer "perception-decision-execution" architecture, including: a load-driven dynamic power distribution model, a multidimensional perception network, dynamic frequency island technology, and a gradient thermal power recovery system. The multidimensional perception network deploys a 32-channel high-precision sensor array for real-time acquisition of electrical parameters: processor core voltage Vcore (0.6-1.2V), power rail current Irail (±1% accuracy), power supply unit PSU ripple (<10mV); thermodynamic parameters: junction temperature (0.1°C resolution), fluid pressure difference (0-5kPa); and computational load: instructions per cycle (IPC), cache hit rate, dynamic random access memory (DRAM) bandwidth utilization, and other information.
[0085] In some embodiments, taking an 8kW GPU server cluster as an example, the AI model predicts that the computing power demand will drop by 30% at night, and automatically switches to low-power mode, saving 2.4kWh / h of electricity. Through the three-stage thermoelectric module, 0.82kWh of electricity is recovered per hour, of which 0.5kWh is fed back to the DC bus and 0.32kWh is stored in the flywheel system; when the ambient temperature is >28°C, absorption refrigeration is started, and 45°C waste heat is used to drive the refrigeration unit, with a heat recovery rate of 0.78.
[0086] Through this application, the first operating status information of the server cluster is obtained; based on the first operating status information and the historical load information of the server cluster, the load prediction model is trained to obtain a target model; the second operating status information of the server cluster is obtained, and the target model is used to determine the target load corresponding to the second operating status information; according to the target load, the operating voltage of the processor is determined, which solves the technical problem that the voltage determination technology of the traditional processor relies on fixed threshold control, lacks dynamic response to actual load changes, maintains high voltage in low load scenarios, and leads to high energy loss, thereby achieving the technical effect of avoiding overpowering and reducing energy loss.
[0087] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0088] The embodiment of the present application further provides a processor voltage determination device 300, Figure 3 A schematic diagram of a voltage determination device for a processor according to an embodiment of the present disclosure is shown in FIG. Figure 3 As shown, including: An acquiring unit 301 is configured to acquire first operating status information of a server cluster; A training unit 302 is configured to train a load prediction model based on the first operating state information and the historical load information of the server cluster to obtain a target model; A first determining unit 303 is configured to obtain second operating status information of the server cluster and determine a target load corresponding to the second operating status information using the target model; The second determining unit 304 is configured to determine an operating voltage of the processor according to the target load.
[0089] Furthermore, in a possible implementation of the embodiment of the present disclosure, the training unit 302 is configured to: Obtaining ambient temperature information of the server cluster; determining training data based on the first operating state information, historical load information of the server cluster, and ambient temperature information of the server cluster; The load prediction model is trained using the training data to obtain a target model.
[0090] Furthermore, in a possible implementation of the embodiment of the present disclosure, the training unit 302 is configured to: Using the training data, training the load prediction model; Determine a first error for each training session, terminate the training session when the first error satisfies a preset condition, and obtain a target parameter combination, wherein the first error is a prediction error, and the target parameter combination includes a weight coefficient and a temperature correction coefficient; Based on the target parameter combination, a target model is obtained.
[0091] Furthermore, in a possible implementation of the embodiment of the present disclosure, the training unit 302 is configured to: Determine a first error and a second error for each training, where the second error is a regularization error; Based on the first error and the second error of each training, a target error of each training is determined, and the training is terminated when the target error meets a preset condition to obtain a target parameter combination.
[0092] Furthermore, in a possible implementation of the embodiment of the present disclosure, the second determining unit 304 is configured to: Acquire a frequency domain list of the processor, the frequency domain list comprising at least one of a high-performance domain, an energy-efficiency domain, a standby domain, and an offline domain; determining a target frequency domain from the frequency domain list according to a load requirement of the processor; An operating voltage of the processor in the target frequency domain is determined according to the target load.
[0093] Furthermore, in a possible implementation of the embodiment of the present disclosure, the processor voltage determination device 300 further includes a recovery unit, which is configured to: In response to an operating voltage of the processor being different from an initial voltage of the processor, determining a temperature difference of the processor; The temperature difference is converted into corresponding target electrical energy using thermoelectric materials, and the target electrical energy is recovered to the power supply system.
[0094] Furthermore, in a possible implementation of the embodiment of the present disclosure, the recovery unit is further configured to: The target electric energy is recovered to the power supply system according to a preset recovery path, wherein the preset recovery path includes at least one of a DC bus, a flywheel energy storage system, and a backup battery pack.
[0095] Through this application, the first operating status information of the server cluster is obtained; based on the first operating status information and the historical load information of the server cluster, the load prediction model is trained to obtain a target model; the second operating status information of the server cluster is obtained, and the target model is used to determine the target load corresponding to the second operating status information; according to the target load, the operating voltage of the processor is determined, which solves the technical problem that the voltage determination technology of the traditional processor relies on fixed threshold control, lacks dynamic response to actual load changes, maintains high voltage in low load scenarios, and leads to high energy loss, thereby achieving the technical effect of avoiding overpowering and reducing energy loss.
[0096] For the description of the features in the embodiment corresponding to the processor voltage determination device, reference can be made to the relevant description of the embodiment corresponding to the processor voltage determination method, which will not be repeated here.
[0097] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned processor voltage determination method embodiments.
[0098] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned processor voltage determination method embodiments when running.
[0099] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0100] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned processor voltage determination method embodiments are implemented.
[0101] An embodiment of the present application further provides another computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned processor voltage determination method embodiments are implemented.
[0102] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0103] The above is a detailed introduction to the voltage determination method, electronic device, storage medium and product of a processor provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A method for determining a processor voltage, characterized in that: The method comprises: Obtaining first operating status information of the server cluster; Training a load prediction model based on the first operating status information and historical load information of the server cluster to obtain a target model; Acquire second operating status information of the server cluster, and determine a target load corresponding to the second operating status information using the target model; An operating voltage of the processor is determined according to the target load.
2. The method for determining processor voltage according to claim 1, wherein: The training of the load prediction model based on the first operating status information and the historical load information of the server cluster to obtain a target model includes: Obtaining ambient temperature information of the server cluster; determining training data based on the first operating state information, historical load information of the server cluster, and ambient temperature information of the server cluster; The load prediction model is trained using the training data to obtain a target model.
3. The method for determining processor voltage according to claim 2, wherein: The method of training the load prediction model using the training data to obtain a target model includes: Using the training data, training the load prediction model; Determine a first error for each training session, terminate the training session when the first error satisfies a preset condition, and obtain a target parameter combination, wherein the first error is a prediction error, and the target parameter combination includes a weight coefficient and a temperature correction coefficient; Based on the target parameter combination, a target model is obtained.
4. The method for determining processor voltage according to claim 3, wherein: The determining of a first error for each training session, and ending the training when the first error satisfies a preset condition, to obtain a target parameter combination, includes: Determine a first error and a second error for each training, where the second error is a regularization error; Based on the first error and the second error of each training, a target error of each training is determined, and the training is terminated when the target error meets a preset condition to obtain a target parameter combination.
5. The method for determining processor voltage according to claim 1, wherein: The determining the operating voltage of the processor according to the target load includes: Acquire a frequency domain list of the processor, the frequency domain list comprising at least one of a high-performance domain, an energy-efficiency domain, a standby domain, and an offline domain; determining a target frequency domain from the frequency domain list according to a load requirement of the processor; An operating voltage of the processor in the target frequency domain is determined according to the target load.
6. The method for determining processor voltage according to claim 1, wherein: After determining the operating voltage of the processor according to the target load, the method further includes: In response to an operating voltage of the processor being different from an initial voltage of the processor, determining a temperature difference of the processor; The temperature difference is converted into corresponding target electrical energy using thermoelectric materials, and the target electrical energy is recovered to the power supply system.
7. The method for determining processor voltage according to claim 6, wherein: The step of recovering the target electric energy to the power supply system includes: The target electric energy is recovered to the power supply system according to a preset recovery path, wherein the preset recovery path includes at least one of a DC bus, a flywheel energy storage system, and a backup battery pack.
8. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the voltage determination method of the processor according to any one of claims 1 to 7 when executing the computer program.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the voltage determination method of the processor according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the voltage determination method of the processor according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Server processor frequency adjustment method and device and storage medium
CN117311987A
GPU power consumption control method and system, computer equipment and storage medium
CN120162210A
Computer board card control method and device, electronic equipment and storage medium
CN120276825A