Fan rotating speed determination model training method, network equipment energy saving method and equipment

By training a fan speed determination model using deep reinforcement learning algorithms and combining load flow and temperature, the fan speed and service board chip frequency are dynamically adjusted, solving the energy-saving problem of network equipment under different service loads and achieving precise energy-saving control of the equipment.

CN121614830AActive Publication Date: 2026-03-06FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511815669.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-06
Estimated Expiration
2045-12-04

AI Technical Summary

Technical Problem

Existing technologies struggle to dynamically optimize energy consumption of network equipment under varying workloads, particularly for fan panels and service panels, leading to a continuous increase in energy consumption.

Method used

A deep reinforcement learning algorithm is used to train a fan speed determination model. Combined with load flow and equipment temperature, the fan speed and service board chip frequency are dynamically adjusted. Precise control is achieved by constructing a state space, action space, reward function and experience pool for training.

Benefits of technology

Without affecting business operations, the system dynamically matches the lowest fan speed and chip frequency to achieve precise energy saving and maximize the energy-saving window.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121614830A_ABST
    Figure CN121614830A_ABST
Patent Text Reader

Abstract

The invention provides a fan rotating speed determination model training method, a network equipment energy-saving method and equipment, and belongs to the technical field of network equipment energy-saving optimization. Constructing a state space based on the temperature of the network equipment at the current moment, the fan rotating speed of the network equipment at the current moment, the load flow of the network equipment at the current moment and the average load flow of a future set step length; determining an action space, and taking the fan rotating speed as a control quantity; a state updating mechanism: simulating the thermal behavior of the real network equipment, calculating the temperature of the network equipment at the next moment according to the current state, the current action and the environment temperature, and further updating the state; constructing a reward function: constructing the reward function based on the temperature control precision, the fan energy consumption and the control smoothness; in addition, the fan rotating speed determination model trained based on the environment modeling can determine the optimal fan rotating speed instruction based on the predicted load flow, and the fan energy-saving effect of the network equipment is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of network equipment energy-saving optimization technology, and specifically relates to a fan speed determination model training method, network equipment energy-saving method and equipment. Background Technology

[0002] In recent years, with the increasing demand for internet access, the overall carbon footprint of the information and communication technology (ICT) industry has received widespread attention. While network technology continues to develop, leading to improved network transmission speeds and capacity, it has also resulted in a continuous increase in network energy consumption. Against this backdrop, energy conservation in network equipment has become a key focus of research and practice in the ICT industry.

[0003] Large-scale network equipment is modularly integrated. Taking Passive Optical Network (PON), a mainstream technology for optical access networks, as an example, it includes power supply boards, service boards, fan boards, switching boards, and uplink boards. From an energy consumption perspective, the service board accounts for nearly 75% of the total energy consumption of a fully configured network device. Among these, the main chip within the service board, as the core unit for data processing, dominates the service board's energy consumption and exhibits significant dynamic changes, making it a key entry point for energy-saving optimization of network equipment. Secondly, the maximum power consumption of the fan board can reach over 400W, nearly 10 times higher than its static power consumption (40W). Its energy consumption is significantly affected by factors such as service load, ambient temperature, and network equipment heat dissipation. Other individual boards, such as the power supply board, switching board, uplink board, and main control board, do not change much with service load or ambient temperature, and their energy-saving potential is limited. Therefore, dynamically optimizing the service board or fan board based on user internet habits and different service loads is crucial for network equipment energy conservation. Summary of the Invention

[0004] To address the aforementioned issues, this application provides a fan speed determination model training method, a network device energy-saving method, and a device, thereby improving the energy-saving effect of network devices.

[0005] Firstly, a method for training a fan speed determination model is provided, including: Environmental modeling: Define the state space based on the current temperature of the network device, the current fan speed of the network device, the current load flow of the network device, and the average load flow of the network device at a set future step size; determine the action space, using fan speed as the control variable; state update mechanism: simulate the thermal behavior of real network devices, calculate the temperature of the network device at the next moment based on the current state, current action, and ambient temperature, and update the state using the temperature of the network device at the next moment; construct the reward function: construct the reward function based on temperature control accuracy, fan energy consumption, and control smoothness; An experience pool is constructed to store the interaction data between the policy network and the environment. The policy network interacts with the environment based on the input state information and saves the state, action, reward and the state of the next time step as an experience. Multiple experiences are randomly selected from the experience pool as training samples for the training and learning phase. Training phase: The current value network and policy network are trained using training samples. The policy network outputs actions based on the state and interacts with the environment to generate new experiences. Learning phase: Calculate the current value of the "state-action pair" in the sample using the current value network, and simultaneously calculate the target value corresponding to the sample using the target value network; calculate the loss function based on the difference between the current value and the target value, and then update the parameters of the current value network using the gradient of the loss function; Repeat the iterations of the training and learning phases described above; After each set number of iterations is completed, the parameters of the current value network are synchronized to the target value network to continue the next round of training and learning until the loss function converges to below the preset threshold, or the change in the loss function is less than the set small value. The convergent policy network is used as the model for determining the fan speed after training.

[0006] Furthermore, a reward function is constructed based on temperature control accuracy, fan energy consumption, and control smoothness, including: The reward function is constructed by weighting the negative penalties of the temperature deviation between the target temperature and the actual temperature of the network device, the fan energy consumption, and the fan speed variation.

[0007] Furthermore, simulating the thermal behavior of real network devices, the temperature of the network device at the next moment is calculated based on the current state, current action, and ambient temperature, including: Simulate the thermal behavior of real network devices. Establish a network device heat dissipation process model based on the current temperature of the network device, the current fan speed, the current load flow, the current action, and the ambient temperature. Utilize the temperature change value ΔT of the network device from the current time t to the next time t+1 to obtain the temperature of the network device at the next time, T(t+1) = T(t) + ΔT.

[0008] Furthermore, the established model for the heat dissipation process of network devices is as follows:

[0009] in, The total thermal capacity of the network device is T; T represents the temperature of the network device. The real-time thermal power input to the network device at time t; This represents the static power required to maintain the operation of network devices; This indicates the current load traffic of the network device; and It is a fixed constant; Let t be the heat transfer power from the network device to the environment. Let t be the temperature of the network device. The overall heat transfer coefficient is determined by the fan speed at the current moment. The ambient temperature.

[0010] Secondly, an energy-saving method for devices based on a multi-dimensional variable AI decision network is provided, which, based on the aforementioned trained fan speed determination model, includes: Predict the average load flow over a set future step size; Status information is constructed based on the current temperature of the network device, the current fan speed, the current load traffic, and the predicted average load traffic over a set step size in the future. The constructed state information is input into the trained fan speed determination model to obtain the optimal fan speed command: The optimal fan speed command is sent to the fan driver module.

[0011] Furthermore, it also includes: Establish the correspondence between the frequency of the service disk chip and the maximum service traffic it can carry without packet loss; Based on the established correspondence between the frequency of the service disk chip and the maximum service traffic that it can carry without packet loss, calculate the minimum chip frequency corresponding to the average load traffic of the future set step size, and perform dynamic adjustment of the minimum chip frequency.

[0012] Furthermore, the minimum chip frequency corresponding to the average load flow of the future set step size is calculated, and dynamic adjustment of the minimum chip frequency is performed, including: Using a traffic prediction model, the load traffic of network devices is collected in real time to predict the network device traffic within a set step in the future. Based on the predicted trends and fluctuations in network device traffic, frequency modulation intervals are defined. Based on the load capacity threshold corresponding to each chip frequency, the frequency modulation interval is initially divided for the predicted average load flow of the future set step size. According to the duration and trend of each interval, the intervals with a duration less than the set step size and the intervals with continuous fluctuations are merged with the adjacent intervals to obtain the chip frequency corresponding to the final predicted average load flow of the future set step size.

[0013] Thirdly, an electronic device is provided, comprising: Memory, the memory storing execution instructions; and A processor that executes the execution instructions stored in the memory, causing the processor to perform the method described above.

[0014] Fourthly, a readable storage medium is provided, wherein executable instructions are stored in the readable storage medium, and the executable instructions are executed by a processor to implement the above-described method.

[0015] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the above-described method.

[0016] Compared with the prior art, this application has the following advantages: 1) By applying deep reinforcement learning algorithms to determine fan speed, the future load flow of the equipment and the current temperature of the equipment are taken into account, thus achieving precise control of fan speed.

[0017] 2) Without affecting business operations, based on the predicted load traffic, the lowest service disk main chip frequency and the optimal fan speed are dynamically matched to determine the optimal energy-saving strategy under the current business load. This allows for dynamic control of the equipment's operating status, enabling more precise grasp of energy-saving opportunities and maximizing the energy-saving window.

[0018] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A framework diagram of a reinforcement learning model according to an embodiment of this application is shown; Figure 2 A flowchart of a network device energy-saving method according to an embodiment of this application is shown; Figure 3 A graph showing the relationship between the frequency and the corresponding no-packet-loss traffic threshold according to an embodiment of this application is shown; Figure 4 A graph showing the relationship between power consumption and fan speed of a fan disk according to an embodiment of this application is provided. Figure 5 A flowchart of a chip frequency adjustment method according to an embodiment of this application is shown. Detailed Implementation

[0021] In the solution of this application embodiment, in order to quantitatively determine the fan speed, a reinforcement learning method is introduced to perform environmental modeling, and the reinforcement learning method is applied to the scenario of determining the fan speed of a network device. To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] This application first provides a method for training a fan speed determination model, then provides a network device energy-saving method based on the trained fan speed determination model obtained by this method, and finally provides the corresponding model training device and energy-saving device, as detailed below: A method for training a fan speed determination model; the reinforcement learning model framework diagram of this application is shown below. Figure 1 As shown, the specific training method is as follows: Step 1: Perform environmental modeling.

[0023] Specifically, the state space is defined as follows: The state space is constructed based on the current temperature of the network device, the current fan speed of the network device, the current load flow of the network device, and the average load flow of the network device in the future with a set step size. By integrating information such as current device temperature, fan speed, and predicted load flow, a state space is constructed, and the state vector at time t is... It can be represented as .

[0024] in, This indicates the temperature of the device at time t, reflecting the thermal inertia of the system. This represents the fan speed at time t, used to describe the historical behavior and smoothing characteristics of fan control; This represents the load flow at time t, reflecting the thermal power level corresponding to the equipment's services. Indicates the device in the future The average load flow rate is used to introduce feedforward information, which can increase the heat dissipation intensity in advance before the flow rate increases, thus achieving predictive control.

[0025] Determine the motion space: By clearly defining the possible actions the fan may take, a continuous action space is adopted in the fan speed control. The fan speed will be output as the control variable at each time t.

[0026] State update mechanism: Simulate the thermal behavior of real network devices, based on the current state. ,action Based on the ambient temperature, calculate the temperature of the network device at the next moment, and update the status using the temperature of the network device at the next moment.

[0027] Specifically, simulating the thermal behavior of real network devices, the system calculates the temperature of the network device at the next moment based on the current state (current network device temperature, current fan speed, and current load flow), current actions, and ambient temperature, including: Simulate the thermal behavior of real network devices. Establish a network device heat dissipation process model based on the current temperature of the network device, the current fan speed, the current load flow, the current action, and the ambient temperature. Utilize the temperature change value ΔT of the network device from the current time t to the next time t+1 to obtain the temperature of the network device at the next time, T(t+1) = T(t) + ΔT.

[0028] The established network device heat dissipation process model is as follows:

[0029] in, This refers to the real-time thermal power input to the network device. This represents the static power required to maintain the operation of network devices; Indicates the current load traffic of the network device; and It is a fixed constant; The current temperature of the network device; Comprehensive thermal capacity of network devices; The heat transfer power of network devices to dissipate heat from the environment; The overall heat transfer coefficient is determined by the fan speed at the current moment. The ambient temperature.

[0030] Construct the reward function: A reward function is constructed based on temperature control accuracy, fan energy consumption, and control smoothness.

[0031] Specifically, a reward function is constructed based on temperature control accuracy, fan energy consumption, and control smoothness, including: The reward function is constructed by weighting the negative penalties of the temperature deviation between the target temperature and the actual temperature of the network device, the fan energy consumption, and the fan speed variation.

[0032] The goal of the fan speed control strategy in this application is to suppress excessive energy consumption and avoid control jitter caused by frequent speed adjustments while ensuring that the device temperature remains stable and close to the target temperature. Therefore, the constructed reward function can define the instantaneous reward at time t as a weighted combination of multiple negative penalties, as shown in the following formula:

[0033] in, For temperature deviation, For fan energy consumption, Let be the fan speed at time t. The range of speed change, , , All values ​​are greater than 0 and the weighted sum is 1, allowing for calibration based on actual scenarios to adjust temperature control accuracy, fan energy consumption, and control smoothness preferences.

[0034] Step 2: Build an experience pool.

[0035] The experience pool is used to store the interaction data between the policy network and the environment. The policy network interacts with the environment based on the input state information, saving the state, action, reward, and state of the next time step as an experience.

[0036] Reference Figure 1 At each time step, the system first senses the current state of the environment and outputs a control action, namely the fan speed, based on the policy network. This action acts on the environment, causing a state change, namely, determining the change in equipment temperature based on the equipment heat dissipation model. The environment returns a new state and reward signal to measure whether the control effect has achieved goals such as energy saving, stability, or accuracy. These experiences are stored in the playback pool.

[0037] Specifically, suppose in a scenario where network devices alternate between peak and low load periods. Based on device traffic prediction results, their load traffic sequence { The temperature exhibits periodic fluctuations. At the initial time t, the equipment temperature is... The target temperature is The ambient temperature is The initial fan speed is The state vector at this point can be represented as: Based on the state vector, a policy network outputs continuous actions. Based on the equipment's heat dissipation model, the ambient temperature changes over time after the action 'at' is performed, resulting in... , Temperature deviation Fan energy consumption With speed control quantity Changes collectively determine immediate rewards At each time step, the model sets the state-action-reward-next state quadruple. Stored in the experience replay pool.

[0038] Step 3: Randomly select multiple experiences from the experience pool as training samples to enter the training and learning phase.

[0039] Training phase: The current value network and policy network are trained using training samples. The policy network outputs actions based on the state and interacts with the environment to generate new experiences. Learning phase: Calculate the current value of the "state-action pair" in the sample using the current value network, and simultaneously calculate the target value corresponding to the sample using the target value network; calculate the loss function based on the difference between the current value and the target value, and then update the parameters of the current value network using the gradient of the loss function.

[0040] Step 4: Repeat the above training and learning phases to update the parameters.

[0041] After each set number of iterations is completed, the parameters of the current value network are synchronized to the target value network to continue the next round of training and learning until the loss function converges to below the preset threshold, or the change in the loss function is less than the set small value. The aforementioned iterative process involves evaluating the long-term benefits of behavior through a value network, thereby continuously optimizing the strategy and achieving progressively improved intelligent control performance.

[0042] Step 5: Use the converged policy network as the model for determining the fan speed after training.

[0043] The current value network is updated by randomly sampling samples, and the parameters are periodically synchronized to the target value network. After multiple rounds of training, the model can automatically adjust the fan speed under different traffic fluctuations to achieve dynamic heat dissipation and energy consumption optimization. When the equipment traffic is at its peak, the model will increase the fan speed in advance to suppress temperature rise; during low load periods, it will automatically reduce the speed to achieve energy-saving operation.

[0044] During the policy inference phase, the trained policy network parameters are exported, the network model is loaded onto the device, and the device operating status is collected in real time, including device temperature, ambient temperature, predicted flow rate, etc. The policy network outputs the optimal fan speed command based on the input status and sends the output speed command to the fan drive module, thereby realizing real-time dynamic speed adjustment of the fan.

[0045] The following explains methods for saving energy on the network: A method for saving energy in network devices, such as Figure 2 As shown, it includes: Step 1: Collect real-time network device traffic data to obtain predicted network device traffic.

[0046] Specifically, the average load flow is predicted in the future based on the current load flow at a set step size.

[0047] Step 2: Deploy the policy network.

[0048] The policy network is the policy network trained by the above-mentioned fan speed determination model training method, which will not be described in detail here.

[0049] Step 3: Determine the optimal fan speed.

[0050] The temperature of the network device is constructed based on the current temperature of the network device, the current fan speed, and the current load traffic. The constructed state information is then input into the trained fan speed determination model to obtain the optimal fan speed command.

[0051] Step 4: Send the optimal fan speed command to the fan driver module.

[0052] In actual operation, the power consumption of the service board is positively correlated with the chip frequency. Lowering the frequency reduces power consumption, but lowering the chip frequency will affect the device's service processing capabilities and cause packet loss. Therefore, when performing device frequency switching, the critical service traffic that can ensure no data loss at different frequencies must be considered. Taking PON devices as an example, their frequency and the corresponding no-packet-loss traffic threshold are as follows: Figure 3 As shown. The power consumption of the fan plate mainly depends on the fan speed, such as... Figure 4 As shown, the higher the rotational speed, the higher the power consumption.

[0053] The above-mentioned energy-saving method has achieved relatively precise control of fan speed. Furthermore, in order to achieve maximum energy saving, the above method also includes: calculating the minimum chip frequency corresponding to the load flow of a future set step size based on the correspondence between the service disk chip frequency and the maximum service flow that it can carry without packet loss, and performing dynamic adjustment of the minimum chip frequency.

[0054] It should be noted that in the chip frequency modulation process based on flow prediction, the predicted flow is not averaged.

[0055] The specific method for adjusting the chip frequency is as follows: Figure 5 As shown, it includes the following steps: Step 501: Collect historical traffic data from the current network and perform data cleaning and preprocessing.

[0056] First, relevant time variables are collected through the network management center, and historical traffic data of the devices is organized. Missing values, outliers, and noise are filled, corrected, and filtered to ensure data quality and integrity. Then, the cleaned data is normalized, adjusting the traffic data according to the mean and standard deviation to eliminate the influence of different feature units, facilitating subsequent model training.

[0057] Step 502: Optimize the flow model structure and parameter settings based on the evaluation indicators.

[0058] In step 502, an AI-based machine learning-based traffic prediction model is introduced to facilitate network traffic prediction for devices. By learning from a large amount of historical traffic data from the existing network, key features reflecting the dynamic changes in the time series are automatically captured, including periodic changes, long-term trend evolution, and short-term fluctuation patterns. Representative feature representations are constructed to provide effective information support for the prediction model. Based on the extracted feature information, the model learns the mapping relationship between historical behavior and future states, generating traffic prediction values ​​for a specific future time range. According to relevant evaluation indicators, such as prediction accuracy, the structure and parameters of the prediction model are optimized to improve its performance.

[0059] Step 503: Deploy a traffic prediction model that meets the evaluation criteria, collect the traffic of a single PON board in real time, and predict the device traffic for the next 24 hours.

[0060] Specifically, a traffic prediction model that meets the evaluation criteria is deployed on the OLT device. This model uses historical traffic data from the network device to predict future traffic changes, providing a control basis for dynamic frequency adjustment of the device. The aforementioned 24-hour period is exemplary and not limited to 24 hours; it can be a prediction of future device traffic in a set step size.

[0061] Step 504: Based on the correspondence between the frequency of the service disk chip and the maximum service traffic that it can withstand without packet loss, perform preliminary division of frequency modulation intervals and merging of some adjacent intervals on the traffic prediction results.

[0062] After obtaining the traffic prediction results, based on the traffic threshold that the chip frequency can carry, that is, the maximum service traffic without packet loss at the chip frequency, the predicted traffic is initially divided into frequency adjustment intervals. According to the duration and trend of each interval, the intervals with shorter durations and continuously fluctuating intervals are merged with adjacent intervals. In other words, the future traffic is divided into different levels, thereby achieving more stable control and avoiding the additional power consumption and potential service loss risks caused by frequent switching of chip frequencies.

[0063] Step 505: Based on the maximum predicted load flow of the PON chip in each frequency modulation interval, determine the lowest adjustable chip frequency of the PON chip in that frequency modulation interval.

[0064] The purpose of step 505 is to achieve dynamic adjustment of the chip frequency based on flow prediction, so as to achieve energy saving or performance optimization.

[0065] Through the above five steps, dynamic frequency adjustment of the business disk chip can be achieved based on traffic prediction.

[0066] In the network equipment energy-saving solution of this application, without affecting services, the lowest service disk main chip frequency and the optimal fan speed are dynamically matched based on the predicted load traffic to determine the optimal energy-saving strategy under the current service load. This allows for dynamic control of the equipment's operating status, enabling more precise grasp of energy-saving opportunities and maximizing the energy-saving window.

[0067] This application also provides an electronic device, including: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, causing the processor to perform any of the storage methods described above.

[0068] The hardware architecture of electronic devices / devices can be implemented using a bus architecture. A bus architecture can include any number of interconnect buses and bridges, depending on the specific application and overall design constraints of the hardware. A bus connects various circuits, including one or more processors, memories, and / or hardware modules. A bus can also connect various other circuits such as peripherals, voltage regulators, power management circuits, external antennas, etc. Buses can be Industry Standard Architecture (ISA) buses, Peripheral Component Interconnect (PCI) buses, or Extended Industry Standard Component (EISA) buses, etc. Buses can be categorized as address buses, data buses, control buses, etc.

[0069] For ease of explanation, certain steps of the above method are described in relation to modules. It should be understood that the corresponding module performing one or more steps of the above method may be one or more hardware modules specifically configured to perform the corresponding step, or implemented by a processor configured to perform the corresponding step, or stored in a computer-readable medium for implementation by a processor, or implemented by some combination thereof.

[0070] The specific implementation of each module in the above-mentioned device can refer to the implementation process of the corresponding steps in the above-mentioned method implementation method of this application, and will not be repeated here.

[0071] This application also provides a readable storage medium storing a computer program, which, when executed by a processor, is used to implement the methods described above. A "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples of readable storage media include: an electrical connection with one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM), etc.

[0072] This application also provides a computer program product. The methods of this application can be implemented, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, all or part of the processes or functions of this application are performed.

[0073] Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A fan speed determination model training method, characterized in that, The method comprises the following steps: environment modeling: define the state space, construct the state space based on the temperature of the network device at the current moment, the fan speed of the network device at the current moment, the load flow of the network device at the current moment and the average load flow of the future setting step; determine the action space, take the fan speed as the control quantity; state update mechanism: simulate the real network device thermal behavior, calculate the temperature of the network device at the next moment according to the current state, the current action and the environment temperature, and update the state by using the temperature of the network device at the next moment; construct the reward function: construct the reward function based on the temperature control accuracy, fan energy consumption and control smoothness; construct an experience pool for storing the interaction data of the policy network and the environment, wherein the policy network interacts with the environment based on the input state information, and saves the state, action, reward and next time step state at each time step as an experience; randomly extract multiple experiences from the experience pool as training samples, and enter the training and learning stage: training stage: train the current value network and the policy network by using the training samples, and the policy network outputs the action according to the state and interacts with the environment to generate new experiences; learning stage: calculate the current value of the "state-action pair" in the sample by using the current value network, and calculate the target value corresponding to the sample by using the target value network; calculate the loss function by the difference between the current value and the target value, and update the parameters of the current value network by using the gradient of the loss function; repeat the iteration of the above training stage and learning stage; after completing the iteration of the set number of rounds, synchronize the parameters of the current value network to the target value network, continue the next round of training and learning, and stop until the loss function converges below the preset threshold value or the change amplitude of the loss function is less than the set small value; the trained policy network is used as the trained fan speed determination model.

2. The method of claim 1, wherein, The reward function is constructed based on the temperature control accuracy, fan energy consumption and control smoothness, which comprises: the weighted sum of the negative punishment of the temperature deviation between the target temperature and the actual temperature of the network device, the fan energy consumption and the change amplitude of the speed is taken as the constructed reward function.

3. The method of claim 1, wherein, simulate the real network device thermal behavior, calculate the temperature of the network device at the next moment according to the current state, the current action and the environment temperature, which comprises: simulate the real network device thermal behavior, establish a network device heat dissipation process model according to the temperature of the network device at the current moment, the fan speed at the current moment, the load flow at the current moment, the current action and the environment temperature, and obtain the temperature of the network device at the next moment T(t+1)=T(t)+ΔT by using the temperature change value ΔT of the network device from the current moment t to the next moment t+1.

4. The method of claim 3, wherein, The established network device heat dissipation process model is as follows: wherein, is the comprehensive heat capacity of the network device; T is the temperature of the network device, is the real-time heat power input to the network device at time t; represents the static power for maintaining the operation of the network device; represents the load traffic of the network device at the current time; and is a fixed constant; is the heat transfer power of the network device to the environment at time t; is the temperature of the network device at time t; is the comprehensive heat transfer coefficient, determined by the current speed of the fan, is the ambient temperature.

5. A method for saving energy of an AI decision network device based on multi-dimensional variables, characterized in that, The trained fan speed determination model based on claim 1 comprises: predict the average load flow of the future setting step; construct state information based on the collected temperature of the network device at the current moment, the fan speed at the current moment, the load flow at the current moment and the predicted average load flow of the future setting step; input the constructed state information into the trained fan speed determination model to obtain the optimal fan speed instruction: The optimal fan rotating speed instruction is sent to the fan driving module.

6. The method of claim 5, wherein, Also include: According to the pre-established corresponding relationship between the service disk chip frequency and the maximum service traffic that it can bear without packet loss, the minimum chip frequency corresponding to the load traffic of the future set step is calculated, and the dynamic adjustment of the minimum chip frequency is performed.

7. The method of claim 6, wherein, According to the pre-established corresponding relationship between the service disk chip frequency and the maximum service traffic that it can bear without packet loss, the minimum chip frequency corresponding to the load traffic of the future set step is calculated, and the dynamic adjustment of the minimum chip frequency is performed, including: Using a traffic prediction model, the load traffic of a single PON board card in a network device is collected in real time, and the load traffic of the single PON board card in the future set step is predicted; According to the corresponding relationship between the service disk chip frequency and the maximum service traffic that it can bear without packet loss, the frequency adjustment interval is preliminarily divided and the adjacent interval is combined according to the traffic prediction result; According to the maximum predicted load traffic of the PON chip in each frequency adjustment interval, the minimum chip frequency that the PON chip can adjust in the frequency adjustment interval is determined.

8. An electronic device, comprising: Include: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor executes the method of any one of claims 1 to 7.

9. A readable storage medium, characterized by, The readable storage medium stores execution instructions, and the execution instructions are executed by the processor to implement the method of any one of claims 1 to 7.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Network training method, unmanned aerial vehicle obstacle avoidance method and device

    CN118034355A

  • Subway station environment control wind-water linkage energy-saving optimization control method based on deep learning

    CN119472311A

  • Equipment fan control method based on reinforcement learning, equipment and medium

    CN119641686A

  • Industrial electric cabinet cooling fan energy efficiency control method

    CN119982615A

  • Power regulation method and device based on load feedforward, medium and product

    CN120184940A