Reinforcement learning-based communication base station emergency power generation operation scheduling method and system

By using reinforcement learning to calculate power demand in real time and optimize scheduling strategies, the problems of slow response and poor adaptability in emergency power management of traditional communication base stations are solved, achieving efficient and reliable emergency power management and ensuring the stability and security of communication networks.

CN119784026BActive Publication Date: 2026-01-02CHONGQING THREE GORGES UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411816750.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2026-01-02
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

Traditional emergency power management methods for communication base stations are slow to respond and have poor adaptability, making it difficult to adjust power resource allocation in real time, resulting in resource waste and potential power outages. Furthermore, experience-based scheduling methods lack effective guidance for complex environments.

Method used

By employing reinforcement learning, a state space and action space are established by collecting base station transmission traffic, terrain features, and equipment data. A reward function is designed, and the model is trained using the Q-learning algorithm. The model calculates power demand in real time and optimizes scheduling strategies, including starting/stopping power generation equipment, switching to backup power, and equipment maintenance.

Benefits of technology

It improves the speed of emergency response and the accuracy of dispatch strategies, rationally allocates resources, reduces energy consumption, extends equipment life, and ensures the stability and security of communication networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784026B_ABST
    Figure CN119784026B_ABST
Patent Text Reader

Abstract

The application discloses a communication base station emergency power generation operation scheduling method and system based on reinforcement learning, and relates to the technical field of communication base station emergency power generation. The method comprises the following steps: firstly, collecting the transmission flow and terrain feature data of the base station, calculating the signal receiving and transmitting power, and combining the cooling system power to calculate the real-time power demand. Then, the device data is acquired to generate a state space, including the standby power supply reserve and the device health state, including the remaining amount of the standby power supply and the health state of the power generation device in the communication base station. The defined action space covers the power generation device operation and maintenance options, and the reward function is designed through the power generation cost, power balance and loss. Then, the reinforcement learning model is trained by using the Q-learning algorithm to form a scheduling scheme. Through the intelligent and automatic scheduling process, efficient and reliable emergency power management is realized, and a solid guarantee is provided for the stability and safety of the communication network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of emergency power generation of communication base stations, in particular to a communication base station emergency power generation operation scheduling method and system based on reinforcement learning. BACKGROUND

[0002] In modern communication networks, base stations serve as key nodes, responsible for carrying a large amount of data transmission and user connection. However, when natural disasters, equipment failures or other force majeure events cause power supply interruption, the normal operation of communication base stations will be seriously affected. In this case, how to effectively carry out emergency power generation scheduling to ensure the continuity and reliability of communication services becomes an important research topic.

[0003] Traditional communication base station emergency power management methods usually rely on pre-set operation manuals and manual decision-making, which have low response speed and adaptability, and are difficult to meet the needs of dynamic changes in the environment. This method not only leads to waste of resources due to the difficulty of real-time adjustment of power resource allocation, but also may cause power interruption at critical moments due to human error. In addition, the scheduling method based on experience cannot fully utilize the real-time data and environmental information of communication base stations for fine management, and lacks effective guidance for optimal scheduling strategies in complex environments.

[0004] In emergency situations, the signal transmission of the base station may fluctuate due to changes in environmental factors, equipment conditions, etc., so real-time power demand analysis and reasonable power resource scheduling are needed to maintain the effective operation of the base station. By comprehensively utilizing the multi-dimensional information of the base station's transmission flow data, terrain characteristics, equipment health status and environmental temperature, the reinforcement learning-based method can dynamically adjust the start-stop of the power generation equipment and the power allocation strategy, improving the efficiency and reliability of emergency power management.

[0005] In the prior art, the publication number CN118839991A discloses a communication base station energy management system based on machine learning, which comprises a maintenance optimization unit, a data processing unit, a model training and prediction unit, an energy allocation unit, an energy execution unit and a user interface unit. The maintenance optimization unit, the data processing unit, the model training and prediction unit, the energy allocation unit, the energy execution unit and the user interface unit are connected in turn, and the maintenance optimization unit is connected with the model training and prediction unit and the energy allocation unit, the user interface unit is connected with the energy allocation unit, and the energy execution unit is connected with the data processing unit. This scheme improves the operation efficiency of the communication base station, reduces the dependence on traditional energy, and promotes the sustainable development of the environment. However, due to the change of the operation environment and conditions of the communication base station over time, such as equipment aging, user demand change, environmental factors, etc., the model needs to have sufficient adaptability and robustness to cope with these dynamic changes. If the model update frequency is insufficient or does not have self-adaptive ability, it may lead to poor performance under new or extreme conditions. Therefore, the accuracy and effectiveness of the system are reduced.

[0006] The above information disclosed in the background section is only used to strengthen the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0007] The purpose of the present application is to provide a communication base station emergency power generation operation scheduling method and system based on reinforcement learning to solve the problems raised in the background.

[0008] To achieve the above-mentioned purpose, the present application provides the following technical scheme:

[0009] A communication base station emergency power generation operation scheduling method based on reinforcement learning, comprising the following specific steps:

[0010] In an emergency, the transmission flow data of the communication base station to be monitored is collected, the signal receiving power and the transmission power of the communication base station are calculated in combination with the transmission flow data and the terrain feature data of the communication base station, and the real-time power demand of the communication base station is calculated in combination with the signal receiving power, the transmission power and the power of the cooling system in the communication base station;

[0011] The equipment data in the communication base station to be monitored is collected, and the state space of the communication base station is generated based on the real-time power demand of the communication base station, the equipment data in the communication base station and the average temperature of the environment where the communication base station is located, wherein the equipment data includes the remaining amount of power reserve of the standby power supply and the health status of the power generation equipment in the communication base station;

[0012] define an action space of the communication base station, the action space of the communication base station including starting and stopping specific power generation equipment, switching to a backup power supply, performing equipment maintenance or repair, and adjusting the input power of the power generation equipment, while designing a reward function based on power generation cost, power supply balance index, and equipment loss of the power generation equipment;

[0013] establish a reinforcement learning model based on the state space of the communication base station, the action space of the communication base station, and the reward function, train the reinforcement learning model using a Q-learning algorithm through finite element analysis, and obtain a communication base station emergency power generation operation scheduling model;

[0014] Based on the obtained communication base station emergency power generation operation scheduling model, input the real-time state space data of the monitored communication base station, and the model outputs the optimal action in the action space under the current state space for execution, wherein the optimal action in the action space is the specific action with the maximum expected return in the action space under the current state space.

[0015] Further, the transmission flow data of the monitored communication base station under emergency conditions is collected, and the signal receiving power and the transmission power of the communication base station are calculated based on the transmission flow data of the communication base station and the terrain feature data, wherein the specific formula for calculating the receiving power is:

[0016]

[0017] wherein, is the receiving power of the communication base station at time t, is the reference power, i.e., the power of the communication base station itself when the communication base station is not transmitting data, is the power coefficient required for unit flow, is the downlink data flow at time t, represents the terrain correction value;

[0018] The formula for calculating the transmission power is:

[0019]

[0020] wherein, is the transmission power of the communication base station at time t, is the uplink data flow at time t;

[0021] wherein the terrain correction value is calculated based on the terrain feature data, which includes the average height of the terrain around the communication base station, the average height of the surrounding obstacles, and the distance of the signal transmission path, and the terrain correction value The specific formula for calculation is:

[0022]

[0023] In the formula, The average height of the obstacle. The distance of the signal transmission path. This represents the average elevation of the terrain surrounding the communication base station. and These are the weighting coefficients for obstacle obstruction and terrain height obstruction, respectively. and and All are greater than 0.

[0024] Furthermore, by combining the signal receiving power, transmitting power, and the power of the cooling system within the communication base station, the real-time power demand of the communication base station is calculated. The formula used to determine the real-time power demand of the communication base station is as follows:

[0025]

[0026] In the formula, Let t be the real-time power demand of the communication base station. Let be the power of the cooling system inside the communication base station at time t. Let t be the power requirement of auxiliary equipment within the communication base station, specifically including the power required by routers, switches, and monitoring equipment;

[0027] The power of the cooling system inside the communication base station at time t The formula used for the calculation is:

[0028]

[0029] In the formula, Let t be the cooling power requirement at the reference temperature. For real-time ambient temperature, As the reference temperature, This is the temperature sensitivity constant.

[0030] Furthermore, data from equipment within the monitored communication base station is collected. Based on the real-time power demand of the communication base station, the data from the equipment within the communication base station, and the average temperature of the environment in which the communication base station is located, a state space for the communication base station is generated. This state space includes four parameters: the real-time power demand of the communication base station, the remaining energy reserve of the backup power supply, the health status of the power generation equipment within the communication base station, and the average temperature of the environment in which the communication base station is located. Specifically, these are expressed as follows:

[0031]

[0032]

[0033] In the formula, a state vector at time t, a real-time power demand of the communication base station at time t, a remaining amount of energy reserve of the backup power source at time t, and a health state of the i-th power generation device in the communication base station at time t and an average temperature of the environment in which the communication base station is located at time t, respectively, where i is an index of the power generation device, and n represents the number of power generation devices started at time t.

[0034] Further, the health state of the power generation device in the communication base station at time t is characterized by the actual output power and the actual power generation efficiency of the power generation device, and the formula according to which the health state is calculated is:

[0035]

[0036] In the formula, is the total usage time of the i-th power generation device before time t, is the power generation power index of the i-th power generation device at time t, is the power generation efficiency index of the i-th power generation device at time t; and the formula according to which the power generation power index of the i-th power generation device is calculated is:

[0037]

[0038] In the formula, is the actual output power of the i-th power generation device at time t, is the rated power of the i-th power generation device, where the actual output power of the power generation device is calculated according to the formula:

[0039]

[0040] In the formula, and are the voltage at the output end and the current at the output end of the i-th power generation device at time t, respectively, is the power factor at the output end;

[0041] where the formula according to which the power generation efficiency index of the i-th power generation device is calculated is:

[0042]

[0043] In the formula, is the input power of the i-th power generation device at time t, where the formula according to which the input power is calculated is:

[0044]

[0045] wherein, is the fuel heat value, is the fuel consumption rate of the i-th power generation device at time t, specifically the amount of fuel consumed per unit time, is the conversion efficiency.

[0046] Further, based on the power generation cost of the power generation device, the power supply balance index and the device loss, a reward function is designed, wherein the formula for calculating the reward function is:

[0047]

[0048] wherein, is the reward function, is the power generation cost of the i-th power generation device at time t, is the power supply balance index, is the device loss, wherein , and are the weight coefficients of the power generation cost of the power generation device, the power supply balance index and the device loss, respectively, wherein and , and are all greater than 0;

[0049] Device loss The formula for calculating the device loss is:

[0050]

[0051] wherein, is the number of started power generation devices;

[0052] wherein the power supply balance index The specific formula is:

[0053]

[0054] wherein, is the real-time power demand of the communication base station at time t.

[0055] Further, the reinforcement learning model is trained by finite element analysis using the Q-learning algorithm, and the specific formula is:

[0056]

[0057] wherein, represents the state executes the action , the cumulative weighted reward obtained after represents an action performed in the current state, is a state vector at time t, is a state after performing the action , represents an immediate reward after performing the action , which is obtained by a reward function calculation, is a learning rate for control convergence, is a discount factor, wherein all are actions contained in the action space of the communication base station;

[0058] According to the above formula, the Q value after performing different actions in different states is constantly iteratively updated based on the reward function, and the above process is repeated until convergence, the training of the reinforcement learning model is completed, and the communication base station emergency power generation operation scheduling model is obtained.

[0059] The application also provides a communication base station emergency power generation operation scheduling system based on reinforcement learning, which is used to execute the communication base station emergency power generation operation scheduling method based on reinforcement learning described above, and comprises:

[0060] A power demand analysis module is configured to collect transmission flow data of a communication base station to be monitored under an emergency condition, combine the transmission flow data of the communication base station with terrain feature data, calculate signal receiving power and transmitting power of the communication base station, and combine the signal receiving power, the transmitting power, and power of a cooling system in the communication base station to calculate real-time power demand of the communication base station.

[0061] A state space definition module is configured to collect equipment data in the communication base station to be monitored, generate a state space of the communication base station based on real-time power demand of the communication base station, equipment data in the communication base station, and average temperature of an environment where the communication base station is located, wherein the equipment data includes a remaining amount of electrical energy reserve of a backup power supply and a health status of power generation equipment in the communication base station.

[0062] A reward and action space definition module is configured to define an action space of the communication base station, wherein the action space of the communication base station includes starting and stopping specific power generation equipment, switching to a backup power supply, performing equipment maintenance or repair, and adjusting input power of the power generation equipment, and design a reward function based on power generation cost of the power generation equipment, a power supply balance index, and equipment loss.

[0063] A scheduling model training module is configured to establish a reinforcement learning model based on the state space of the communication base station, the action space of the communication base station, and the reward function, train the reinforcement learning model by using a Q-learning algorithm through finite element analysis, and obtain a communication base station emergency power generation operation scheduling model.

[0064] The power generation operation scheduling module is configured to input real-time state space data of the communication base station to be monitored into the obtained communication base station emergency power generation operation scheduling model, and output an optimal action in the action space under the current state space for execution, wherein the optimal action in the action space under the current state space is a specific action with the maximum expected return in the action space under the current state space.

[0065] Compared with the prior art, the present application has the following advantages:

[0066] The present application introduces advanced reinforcement learning algorithms to address the slow response, poor adaptability, and resource waste of traditional emergency power generation scheduling methods, significantly improving the communication base station's response capability in emergency situations. By collecting and analyzing information such as transmission flow data, terrain features, and equipment data of the communication base station, the scheme can calculate the power demand of the base station in real time, thereby scheduling power and avoiding communication interruptions caused by power shortages. By defining clear state space and action space and designing a reward function based on power generation cost, power supply balance, and equipment wear and tear, the reinforcement learning model can learn and optimize the scheduling strategy in various complex scenarios. The model can quickly output the optimal scheduling action under different environments and states, improving the speed and accuracy of emergency response. In particular, in the case of limited resources, the scheme can reasonably allocate standby power sources and use of power generation equipment, effectively reducing energy consumption and equipment wear and tear, thereby extending the service life and service time of the equipment. In addition, the Q-learning algorithm is used for model training, enabling the scheme to quickly converge to the optimal strategy under limited data and computing resources. Therefore, the reinforcement learning-based communication base station emergency power generation operation scheduling method realizes efficient and reliable emergency power management through intelligent and automated scheduling processes, providing a solid guarantee for the stability and security of the communication network. BRIEF DESCRIPTION OF DRAWINGS

[0067] Figure 1 The present application is a whole method flowchart;

[0068] Figure 2 The present application is a whole system structure schematic diagram. DETAILED DESCRIPTION

[0069] To make the purpose, technical scheme and advantages of the present application clearer, the present application is further described in detail below with specific examples.

[0070] It should be noted that the technical terms or scientific terms used in the present application should be understood as the general meaning understood by those skilled in the art unless otherwise defined. The "first", "second" and similar words used in the present application do not represent any order, quantity or importance, but are only used to distinguish different components. "Include" or "contain" and similar words mean that the elements or objects before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connected" or "connected" and similar words are not limited to physical or mechanical connection, but can include electrical connection, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to represent the relative positional relationship, when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0071] Embodiments:

[0072] Please refer to Figure 1 The present application provides a technical solution:

[0073] A communication base station emergency power generation operation scheduling method based on reinforcement learning, the specific steps comprising:

[0074] Step 1: Collect the transmission flow data of the communication base station to be monitored under emergency conditions, combine the transmission flow data of the communication base station with the terrain feature data, calculate the signal receiving power and transmitting power of the communication base station, and combine the signal receiving power, transmitting power and power of the cooling system in the communication base station to calculate the real-time power demand of the communication base station.

[0075] Wherein, the emergency condition specifically refers to the power failure in natural disasters or emergencies.

[0076] Collect the transmission flow data of the communication base station to be monitored under emergency conditions, combine the transmission flow data of the communication base station with the terrain feature data, calculate the signal receiving power and transmitting power of the communication base station, wherein the specific formula for calculating the receiving power is:

[0077]

[0078] In the formula, is the receiving power of the communication base station at time t, is the reference power, that is, the power of the communication base station itself when the communication base station does not transmit data, is the power coefficient required for unit flow, is the downlink data flow at time t, which refers to the data flow transmitted from the base station to the user equipment, represents the terrain correction value;

[0079] By analyzing the historical power consumption and traffic data of the operator or the device, it can be estimated by regression analysis or other methods , 2G / 3G networks may be tens to hundreds of milliwatts per bit per second (mW / bit / s). These older network technologies are generally less efficient, especially in poor signal conditions, 4G LTE networks are usually a few milliwatts to tens of milliwatts per bit per second (mW / bit / s). LTE networks are more efficient than earlier technologies, especially in good signal conditions, 5G networks may be further reduced to less than one milliwatt to a few milliwatts per bit per second (mW / bit / s).

[0080] where the reference power represents the power consumption of the communication base station itself when there is no data transmission. This is a basic power consumption level related to the energy consumption of the base station in the idle state. Even without user data traffic, the base station still needs to maintain basic operations such as signal maintenance, control channel function operation, etc., and therefore consumes a certain amount of power. This reference power can be regarded as the minimum power required to maintain the basic operation of the base station.

[0081] The reference power is usually dependent on the type, model, supplier, and configuration and operating conditions of the base station, and the power consumption characteristics of the base station, including the power consumption in the idle state, are usually listed in the technical documents or equipment specifications provided by the manufacturer.

[0082] The formula for calculating the transmit power is as follows:

[0083]

[0084] where is the transmit power of the communication base station at time t, is the uplink data traffic at time t, and the uplink data traffic refers to the data traffic transmitted from the user equipment to the base station;

[0085] where the uplink traffic data and the downlink traffic data can be obtained by using network monitoring tools and device logs of the communication base station.

[0086] where the terrain correction value is calculated based on terrain feature data, including the average height of the terrain around the communication base station, the average height of the surrounding obstacles, and the distance of the signal transmission path, and the terrain correction value The specific formula for calculating is as follows:

[0087]

[0088] where is the average height of obstacles, is the distance of signal transmission path, is the average height of terrain around the communication base station, and are the weight coefficients of obstacle obstruction and terrain height obstruction, respectively, wherein and are greater than 0.

[0089] In wireless signal propagation, obstacles are the main attenuation factor. Tall buildings, trees or other obstacles can significantly obstruct signal propagation, therefore, the average height of obstacles is added to the logarithmic term in the formula to reflect the nonlinear impact of obstacles on signals;

[0090] The average height of terrain around the base station affects the scattering and reflection paths of signals, especially in uneven areas, signal propagation loss usually increases with the increase of distance, therefore is used to represent the square effect of distance.

[0091] Obstacle height has a significant impact on signal attenuation. Larger will increase so that the signal needs more compensation power, the square of the distance of the signal transmission path is used in the formula to emphasize the significant attenuation of long distance on signal strength. The overall characteristics of the terrain are considered to affect signal propagation, especially in areas with complex terrain. Since the impact of obstacle obstruction on signals is large and direct, it is set that and and are greater than 0.

[0092] Step 2: Collect equipment data in the communication base station to be monitored, generate a state space of the communication base station based on the real-time power demand of the communication base station, equipment data in the communication base station, and the average temperature of the environment where the communication base station is located, wherein the equipment data includes the remaining amount of energy reserve of the backup power supply and the health status of the power generation equipment in the communication base station.

[0093] The real-time power demand of the communication base station is calculated by combining the signal receiving power, the transmitting power and the power of the cooling system in the communication base station, wherein the formula for the real-time power demand of the communication base station is:

[0094]

[0095] In the formula, is the real-time power demand of the communication base station at time t, ​Let be the power of the cooling system inside the communication base station at time t. The power requirements of auxiliary equipment within the communication base station at time t, specifically including the power required by routers, switches, and monitoring equipment, are obtained from the equipment specifications provided by the manufacturers. These documents typically list the typical power consumption range of the equipment. By summing the nominal power consumption of all equipment, the total power requirements of the base station's auxiliary equipment under normal circumstances are obtained.

[0096] The power of the cooling system inside the communication base station at time t The formula used for the calculation is:

[0097]

[0098] In the formula, Let t be the cooling power requirement at the reference temperature. For real-time ambient temperature, As the reference temperature, This is the temperature sensitivity constant.

[0099] Among them, reference temperature Generally 25 , The cooling power requirement at time t at a reference temperature is obtained from the technical specifications of the air conditioning equipment or refrigeration system, which specifies its rated power at the reference temperature. Manufacturers may provide performance data for the equipment under different environmental conditions, including power requirements under standard conditions. The typical range is from several kilowatts (kW) to tens of kilowatts.

[0100] Data from equipment within the communication base station to be monitored is collected. Based on the real-time power demand of the communication base station, the data from the equipment within the base station, and the average temperature of the environment in which the base station is located, a state space for the communication base station is generated. This state space includes four parameters: the real-time power demand of the communication base station, the remaining energy reserve of the backup power supply, the health status of the power generation equipment within the communication base station, and the average temperature of the environment in which the communication base station is located. Specifically, these are expressed as follows:

[0101]

[0102]

[0103] In the formula, This represents the state vector at time t. Let t be the real-time power demand of the communication base station. Let t be the remaining electrical energy reserve of the backup power supply. and Let be the health status of the i-th power generation device within the communication base station at time t, and let be the average temperature of the environment surrounding the communication base station at time t, where i is the index of the power generation device. n represents the number of power generation devices started at time t.

[0104] The health status of the power generation device in the communication base station at time t characterized by the actual output power and actual power generation efficiency of the power generation device, and the specific formula is:

[0105]

[0106] In the formula, is the total usage time of the i-th power generation device before time t, is the power generation power index of the i-th power generation device at time t, is the power generation efficiency index of the i-th power generation device at time t; record the time of each start and stop of the device, and periodically accumulate these times to obtain the total usage time, record the running time during device maintenance or inspection, and the cumulative time can usually be found in the maintenance log.

[0107] wherein the formula for calculating the power generation power index of the i-th power generation device is:

[0108]

[0109] In the formula, is the actual output power of the i-th power generation device at time t, is the rated power of the i-th power generation device, and the technical specifications of the device, including the rated power, are usually attached to the nameplate of the power generation device. This is the most direct and authoritative source, in addition to the nameplate, there may be other identification tags or plates on the device, which list the rated power.

[0110] wherein the actual output power of the power generation device is calculated according to the formula:

[0111]

[0112] In the formula, and are the voltage and current at the output end of the i-th power generation device at time t, is the power factor of the output end, and modern power generation devices are usually equipped with built-in power factor tables or displays, which can directly read the power factor.

[0113] wherein the formula for calculating the power generation efficiency index of the i-th power generation device is:

[0114]

[0115] In the formula, is the input power of the i-th power generation device at time t, where the formula for calculating the input power is:

[0116]

[0117] where, is the fuel heat value, is the fuel consumption rate of the i-th power generation device at time t, specifically the amount of fuel consumed per unit of time, is the conversion efficiency. The fuel heat value will be affected by the environmental humidity, resulting in a decrease in the fuel heat value, where the specific correction formula is:

[0118]

[0119] where, is the initial fuel heat value, and the fuel supplier will usually provide the technical specifications of the fuel, including its heat value. These information can be found in the procurement contract or product specification, is the real-time environmental humidity, is the reference environmental humidity, generally 50%.

[0120] Step 3: Define the action space of the communication base station, which includes starting and shutting down specific power generation devices, switching to backup power, performing equipment maintenance or repair, and adjusting the input power of the power generation devices, while designing a reward function based on the power generation cost of the power generation device, the power supply balance index, and the equipment wear and tear.

[0121] Design a reward function based on the power generation cost of the power generation device, the power supply balance index, and the equipment wear and tear, where the formula for calculating the reward function is:

[0122]

[0123] where, is the reward function, is the power generation cost of the i-th power generation device at time t, is the power supply balance index, is the equipment wear and tear, where , and are the weight coefficients of the power generation cost of the power generation device, the power supply balance index, and the equipment wear and tear, respectively, where and , and are all greater than 0, where the power generation cost is calculated by the fuel consumption rate , and the specific formula is:

[0124]

[0125] wherein, is the unit price of fuel.

[0126] equipment wear The formula used for the calculation is:

[0127]

[0128] wherein, is the number of power generation equipment started;

[0129] wherein the power supply balance index The specific formula used is:

[0130]

[0131] wherein, is the real-time power demand of the communication base station at time t.

[0132] The impact of the power generation cost is measured by an exponential function , which means that a lower cost will significantly increase the reward value. The use of the exponential function indicates that the negative impact of cost increase on the reward is nonlinear, and a small increase in cost will lead to a significant decrease in reward. The reward function emphasizes the importance of reducing costs. Cost is a key factor in the operation of the power system, and reducing costs can improve economic benefits, so the power generation cost is inversely proportional to the reward function .

[0133] The power supply balance index highlights the stability and reliability of power supply. The balance of power supply directly affects the safety of the power grid, and the use of a logarithmic function represents the gain of power supply balance. The logarithmic function is suitable for describing the effect of diminishing returns, and initial small improvements have a greater impact on the increase in rewards, while the incremental returns of subsequent improvements gradually decrease, so the power supply balance index is directly proportional to the reward function .

[0134] Wear is an unavoidable factor in the operation of all equipment, and the square term represents that the negative impact of wear on the reward is accelerating. That is, the weakening effect of wear increase on the reward is intensified, and the reward function is more sensitive to the increase in wear, encouraging the reduction of equipment wear to extend the service life and improve efficiency, so the equipment wear is inversely proportional to the reward function .

[0135] Step 4: Establish a reinforcement learning model based on the state space of the communication base station, the action space of the communication base station and the reward function, and train the reinforcement learning model by using the Q-learning algorithm through finite element analysis to obtain the communication base station emergency power generation operation scheduling model.

[0136] The Q-learning algorithm is used to train the reinforcement learning model through finite element analysis, and the specific formula is:

[0137]

[0138] In the formula, represents the state After performing the action , the cumulative weighted reward obtained is represents the action performed in the current state, is the state vector at time t, is the state after performing the action , represents the immediate reward after performing the action , which is calculated by the reward function, is the learning rate for controlling convergence, is the discount factor, where Both are actions contained in the action space of the communication base station.

[0139] The specific training steps include: creating a Q table, the dimension of the table is the state space multiplied by the action space, the Q values in the initial table are all 0, setting the learning parameters: determining the learning rate , the discount factor , where the learning rate has a value range of , the discount factor has a range of , represents the immediate reward, tends to 1 indicating future rewards. Use the Q-learning algorithm for training: obtain the initial state from the environment, based on the above formula, constantly update the Q values after performing different actions in different states based on the reward function, repeat the above process until convergence, complete the training of the reinforcement learning model, obtain the communication base station emergency power generation operation scheduling model. Since the Q-learning algorithm is a mature technical means, it is not described here.

[0140] Step 5: Based on the obtained communication base station emergency power generation operation scheduling model, input the real-time state space data of the communication base station to be monitored, and the model outputs the optimal action in the action space under the current state space for execution, wherein the optimal action in the action space is a specific action with the maximum expected return in the action space under the current state space.

[0141] Real-time acquisition of state information of the communication base station. Including power demand, battery capacity, weather conditions, current power generation equipment state, etc., ensure that the data is accurate and timely, format the acquired real-time data into a state compatible with the model, in the trained Q table, according to the current state, find the Q value of all possible actions, select the action with the highest Q value, and execute the selected optimal action in the communication base station.

[0142] Please refer to Figure 2 The application also provides a communication base station emergency power generation operation scheduling system based on reinforcement learning, which is used to execute the above-mentioned communication base station emergency power generation operation scheduling method based on reinforcement learning, and comprises:

[0143] A power demand analysis module is used to collect transmission flow data of the communication base station to be monitored under emergency conditions, combine the transmission flow data of the communication base station with the terrain feature data, calculate the signal receiving power and the transmitting power of the communication base station, and combine the signal receiving power, the transmitting power with the power of the cooling system in the communication base station to calculate the real-time power demand of the communication base station.

[0144] A state space definition module is used to collect equipment data in the communication base station to be monitored, generate a state space of the communication base station based on the real-time power demand of the communication base station, the equipment data in the communication base station and the average temperature of the environment where the communication base station is located, wherein the equipment data includes the remaining amount of power reserve of the backup power supply and the health status of the power generation equipment in the communication base station.

[0145] A reward and action space definition module is used to define the action space of the communication base station, wherein the action space of the communication base station includes starting and stopping specific power generation equipment, switching to a backup power supply, performing equipment maintenance or repair, and adjusting the input power of the power generation equipment, and a reward function is designed based on the power generation cost of the power generation equipment, the power supply balance index and the equipment loss.

[0146] A scheduling model training module is used to establish a reinforcement learning model based on the state space of the communication base station, the action space of the communication base station and the reward function, and the reinforcement learning model is trained by using a Q-learning algorithm through finite element analysis to obtain a communication base station emergency power generation operation scheduling model.

[0147] The power generation operation scheduling module is configured to input real-time state space data of the communication base station to be monitored into the obtained communication base station emergency power generation operation scheduling model, and output an optimal action in the action space under the current state space for execution, wherein the optimal action in the action space is a specific action with the maximum expected return in the action space under the current state space.

[0148] The above formulas are dimensionless values calculated, and the formulas are obtained by collecting a large amount of data to simulate a formula of the most recent real situation, and the preset parameters in the formula are set by a person skilled in the art according to the actual situation.

[0149] The above embodiments can be realized wholly or partially by software, hardware, firmware or any other combination. When realized by software, the above embodiments can be realized wholly or partially in the form of a computer program product. Those skilled in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized by hardware or software methods depends on the specific application and design constraints of the technical solutions.

[0150] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, which can be located in one place or distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0151] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for scheduling emergency power generation operations at communication base stations based on reinforcement learning, characterized in that, The specific steps include: In emergency situations, the transmission traffic data of the communication base station to be monitored is collected. Combined with the transmission traffic data of the communication base station and the terrain feature data, the signal receiving power and transmission power of the communication base station are calculated. Combined with the signal receiving power, transmission power and the power of the cooling system inside the communication base station, the real-time power demand of the communication base station is calculated. Data on equipment within the communication base station to be monitored is collected. Based on the real-time power demand of the communication base station, the data on equipment within the communication base station, and the average temperature of the environment in which the communication base station is located, a state space of the communication base station is generated. The equipment data includes the remaining power reserve of the backup power supply and the health status of the power generation equipment within the communication base station. Define the action space of the communication base station, which includes starting and stopping specific power generation equipment, switching to backup power, performing equipment maintenance or repair and adjusting the input power of the power generation equipment, and designing a reward function based on the power generation cost of the power generation equipment, the power supply balance index and equipment loss. A reinforcement learning model is established based on the state space, action space, and reward function of the communication base station. The reinforcement learning model is trained by the Q-learning algorithm through finite element analysis to obtain the emergency power generation operation scheduling model of the communication base station. Based on the obtained emergency power generation operation scheduling model of the communication base station, the real-time state space data of the communication base station to be monitored is input, and the model outputs the optimal action in the action space under the current state space for execution. The optimal action in the action space is the specific action with the highest expected return in the action space under the current state space.

2. The method for scheduling emergency power generation operations of communication base stations based on reinforcement learning according to claim 1, characterized in that: In emergency situations, transmission traffic data of the communication base station to be monitored is collected. Combined with the transmission traffic data and terrain feature data, the signal receiving power and transmitting power of the communication base station are calculated. The specific formula used to calculate the receiving power is as follows: In the formula, Let be the received power of the communication base station at time t. The reference power represents the power of the communication base station when it is not transmitting data. The power factor required per unit flow rate. Let be the downlink data flow at time t. Indicates the terrain correction value; The formula used to calculate the transmit power is: In the formula, Let t be the transmission power of the communication base station. Let be the uplink data flow at time t; Terrain correction value Calculations are performed based on terrain feature data, which includes the average elevation of the terrain surrounding the communication base station, the average elevation of surrounding obstacles, and the distance of the signal transmission path, along with terrain correction values. The specific formula used for the calculation is as follows: In the formula, The average height of the obstacle. The distance of the signal transmission path. This represents the average elevation of the terrain surrounding the communication base station. and These are the weighting coefficients for obstacle obstruction and terrain height obstruction, respectively. and and All are greater than 0.

3. The method for scheduling emergency power generation operations of communication base stations based on reinforcement learning according to claim 2, characterized in that: By combining the signal receiving power, transmitting power, and the power of the cooling system within the communication base station, the real-time power demand of the communication base station is calculated. The formula used to determine the real-time power demand of the communication base station is as follows: In the formula, Let t be the real-time power demand of the communication base station. Let be the power of the cooling system inside the communication base station at time t. Let t be the power requirement of auxiliary equipment within the communication base station, specifically including the power required by routers, switches, and monitoring equipment; The power of the cooling system inside the communication base station at time t The formula used for the calculation is: In the formula, Let t be the cooling power requirement at the reference temperature. For real-time ambient temperature, As the reference temperature, This is the temperature sensitivity constant.

4. The method for scheduling emergency power generation operations of communication base stations based on reinforcement learning according to claim 1, characterized in that: Data from equipment within the communication base station to be monitored is collected. A state space for the communication base station is generated based on its real-time power demand, the data from the equipment within the base station, and the average temperature of the environment in which the base station is located. This state space includes four parameters: the real-time power demand of the communication base station, the remaining energy reserve of the backup power supply, the health status of the power generation equipment within the communication base station, and the average temperature of the environment in which the communication base station is located. Specifically, these are expressed as follows: In the formula, This represents the state vector at time t. Let t be the real-time power demand of the communication base station. Let t be the remaining electrical energy reserve of the backup power supply. and Let i represent the health status of the i-th power generation device within the communication base station at time t, and let i represent the average temperature of the environment surrounding the communication base station at time t. , where n represents the number of power generation devices started at time t.

5. The method for scheduling emergency power generation operations of communication base stations based on reinforcement learning according to claim 4, characterized in that: The health status of the power generation equipment in the communication base station at time t It is characterized by the actual output power and actual power generation efficiency of the power generation equipment, and the specific formula used is as follows: In the formula, Let be the total usage time of the i-th generating device before time t. Let be the power generation index of the i-th generating unit at time t. Let be the power generation efficiency index of the i-th power generation device at time t; where the formula for calculating the power generation index of the i-th power generation device at time t is: In the formula, Let be the actual output power of the i-th generating unit at time t. Let be the rated power of the i-th power generation device, where is the actual output power of the power generation device. The formula used for the calculation is: In the formula, and Let be the voltage and current at the output terminal of the i-th generating device at time t, respectively. The power factor at the output terminal; The formula used to calculate the power generation efficiency index of the i-th power generation equipment is as follows: In the formula, Let be the input power of the i-th generating device at time t, where the formula for calculating the input power is: In the formula, The calorific value of the fuel. Let be the fuel consumption rate of the i-th power generation unit at time t, specifically the amount of fuel consumed per unit time. For conversion efficiency.

6. The method for scheduling emergency power generation operations of communication base stations based on reinforcement learning according to claim 5, characterized in that: The reward function is designed based on the power generation cost of the power generation equipment, the power supply equilibrium index, and equipment losses. The formula used to calculate the reward function is as follows: In the formula, For the reward function, Let i be the power generation cost of the i-th power generation device at time t. The electricity supply balance index, For equipment wear and tear, of which , and These are the weighting coefficients for power generation equipment cost, power supply balance index, and equipment loss, respectively. and , and All are greater than 0; Equipment loss The formula used for the calculation is: Among them, the power supply balance index The specific formula used is as follows: In the formula, Let t be the real-time power demand of the communication base station.

7. The method for scheduling emergency power generation operations of communication base stations based on reinforcement learning according to claim 1, characterized in that: The reinforcement learning model is trained using the Q-learning algorithm through finite element analysis. The specific formula used is as follows: In the formula, Representing state Next action The cumulative weighted reward obtained afterwards This indicates the action to be performed in the current state. Let be the state vector at time t. To perform the action The state after that, Indicates the execution of an action The immediate reward is calculated by the reward function. To control the learning rate during convergence, , where All of these are actions contained within the action space of the communication base station; Based on the above formula, the Q-values ​​after performing different actions under different states are continuously updated based on the reward function. The above process is repeated until convergence, thus completing the training of the reinforcement learning model and obtaining the emergency power generation operation scheduling model for communication base stations.

8. A communication base station emergency power generation operation scheduling system based on reinforcement learning, characterized in that: The reinforcement learning-based emergency power generation operation scheduling system for communication base stations is used to execute the reinforcement learning-based emergency power generation operation scheduling method for communication base stations as described in any one of claims 1-7, including: The power demand analysis module is used to collect transmission traffic data of the communication base station to be monitored in emergency situations. Combining the transmission traffic data of the communication base station with terrain feature data, it calculates the signal receiving power and transmission power of the communication base station. Combining the signal receiving power, transmission power and the power of the cooling system inside the communication base station, it calculates the real-time power demand of the communication base station. The state space definition module is used to collect equipment data within the communication base station to be monitored. Based on the real-time power demand of the communication base station, the equipment data within the communication base station, and the average temperature of the environment where the communication base station is located, the state space of the communication base station is generated. The equipment data includes the remaining power reserve of the backup power supply and the health status of the power generation equipment within the communication base station. The reward and action space definition module is used to define the action space of the communication base station. The action space of the communication base station includes starting and stopping specific power generation equipment, switching to backup power, performing equipment maintenance or repair and adjusting the input power of the power generation equipment. At the same time, a reward function is designed based on the power generation cost of the power generation equipment, the power supply balance index and equipment loss. The scheduling model training module is used to establish a reinforcement learning model based on the state space, action space and reward function of the communication base station. The reinforcement learning model is trained by Q-learning algorithm through finite element analysis to obtain the emergency power generation operation scheduling model of the communication base station. The power generation operation scheduling module is used to input the real-time state space data of the communication base station to be monitored based on the obtained emergency power generation operation scheduling model of the communication base station. The model outputs the optimal action in the action space under the current state space for execution, wherein the optimal action in the action space is the specific action with the highest expected return in the action space under the current state space.

Citation Information

Patent Citations

  • Communication base station energy management system based on machine learning

    CN118839991A

  • DQN-based 5G fusion intelligent power distribution network energy management method

    CN113988356A

  • New energy power system elastic optimization method based on deep reinforcement learning

    CN114330113A