A power distribution method and device based on communication power supply system

By building a multi-dimensional target reward function and reinforcement learning training model, the power supply redundancy problem caused by the lack of dynamic adjustment of traditional communication power supply power supply systems is solved, and efficient adaptation and reliable power supply to complex network environments are achieved.

CN120320330BActive Publication Date: 2025-08-19FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510811601.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-08-19
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Traditional communication power supply systems rely on fixed rules to allocate power, lack dynamic adjustment capabilities, resulting in redundancy in power supply.

Method used

A multi-dimensional target reward function based on equipment energy consumption cost, residual power variance and battery life decay is constructed, and the optimal power supply allocation scheme is generated through reinforcement learning training.

Benefits of technology

It significantly improves the adaptability of the communication power supply system to complex and variable communication network environments, avoids power supply redundancy, and ensures power supply reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120320330B_ABST
    Figure CN120320330B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of communication power supply energy distribution, and discloses a power supply distribution method and device based on a communication power supply system. The present invention accurately coordinates multiple objectives and overcomes the limitations of a single objective by constructing a multi-dimensional target reward function based on equipment energy consumption cost, remaining power balance, and battery life decay. The present invention also uses historical operation data sets to conduct reinforcement learning training on the initial power supply distribution model, so that the model gradually learns the optimal power supply strategy under different scenarios, and finally generates a target power supply distribution model that can autonomously optimize decisions. The present invention significantly improves the adaptability of the communication power supply system to complex and changeable communication network environments, and while ensuring power supply reliability, it effectively avoids the power supply redundancy problem caused by static allocation strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication power supply energy distribution, and in particular to a power supply distribution method and device based on a communication power supply system. Background Art

[0002] Telecommunications power supply systems are a core component of modern communications networks, providing stable and reliable power for various communications equipment and building loads within communication stations. With the rapid development of technologies like 5G, the Internet of Things, and cloud computing, the scale of communication networks continues to expand. Therefore, the introduction of intelligent technologies to optimize and transform telecommunications power supply systems has become an inevitable trend in the communications industry.

[0003] At present, traditional communication power supply systems often rely on fixed rules to distribute electricity, such as load balancing and priority presets, and lack dynamic adjustment capabilities, resulting in power supply redundancy. Summary of the Invention

[0004] The present invention provides a power distribution method and device based on a communication power supply system, which solves the technical problem that traditional communication power supply systems often rely on fixed rules to distribute electric energy, lack dynamic adjustment capabilities, and lead to power supply redundancy.

[0005] A first aspect of the present invention provides a power distribution method based on a communication power supply system, comprising:

[0006] Acquire a historical operation data set of the communication power supply system, wherein the historical operation data set includes a plurality of historical operation data within a preset time period;

[0007] Using device energy consumption cost data, device remaining power variance data, and battery life attenuation data as optimization targets, we construct the target reward function of the initial power distribution model.

[0008] Using each of the historical operation data to input the initial power supply distribution model for training to obtain a plurality of predicted historical distribution data;

[0009] Based on the target reward function, using each of the historical operating data and the associated predicted historical distribution data to perform reward calculation, and determining a target power supply distribution model according to the reward calculation result;

[0010] The target power distribution model is used to solve the real-time operation data to generate an optimal power distribution plan.

[0011] Optionally, the step of performing reward calculation based on the target reward function using each of the historical operation data and the associated predicted historical distribution data, and determining a target power supply distribution model according to the reward calculation result, includes:

[0012] Performing reward calculation on each of the historical operation data and the associated predicted historical distribution data using the target reward function to obtain a plurality of reward values;

[0013] Calculating a cumulative reward using the plurality of reward values and a preset discount factor to obtain a cumulative reward value;

[0014] When the cumulative reward value is greater than a preset cumulative threshold, the initial power supply distribution model is used as a target power supply distribution model.

[0015] Optionally, when the cumulative reward value is less than or equal to the preset cumulative threshold, each of the predicted historical distribution data is used to calculate the loss value to obtain the loss value;

[0016] With the goal of minimizing the loss value, adjusting the network parameters of the initial power distribution model and performing sample priority updates on the plurality of historical operation data;

[0017] Based on the historical operation data set after the sample priority is updated, the process jumps to the step of using the historical operation data set to input the initial power supply distribution model for training to obtain predicted historical distribution data.

[0018] Optionally, updating the sample priority of the plurality of historical operation data includes:

[0019] Determining a target state action value corresponding to each of the historical operation data using the network parameters, the historical operation data, and the associated predicted historical allocation data;

[0020] An error calculation is performed using the instant reward value, the target state action value, the preset discount factor, and the preset standard state action value to obtain a target error value corresponding to each of the historical operation data;

[0021] The priority value corresponding to each of the historical operation data is obtained by performing a priority calculation using the absolute value of the target error value and a preset parameter factor;

[0022] Performing a sum operation using all the priority values to obtain a priority value sum;

[0023] Performing a ratio operation on each priority value and the priority value sum to obtain a sampling probability corresponding to each historical operation data;

[0024] Sorting the historical operation data from large to small according to the sampling probability;

[0025] Calculate the sampling probability interval corresponding to each of the sorted historical operation data;

[0026] Randomly generate multiple random numbers, and use each of the random numbers to perform interval retrieval, matching the sampling probability interval of each of the random numbers as the target sampling probability interval;

[0027] The historical operation data associated with the target sampling probability interval is selected to form an updated historical operation data set.

[0028] Optionally, the error calculation is performed using the immediate reward value, the target state action value, the preset discount factor, and the preset standard state action value to obtain the target error value corresponding to each of the historical operation data, including:

[0029] Performing a multiplication operation using the preset discount factor and the preset standard state action value to obtain a target multiplication value corresponding to each of the historical operation data;

[0030] Performing a sum operation using the instant reward value and the target multiplier value to obtain a target sum value corresponding to each of the historical operation data;

[0031] The target sum value and the target state action value are used to perform a difference operation to obtain a target error value corresponding to each of the historical operation data.

[0032] Optionally, solving the real-time operating data using the target power distribution model to generate an optimal power distribution solution includes:

[0033] Solving the real-time operation data by using the target power distribution model to generate the optimal power distribution plan that meets the preset power distribution constraints;

[0034] The preset power supply distribution constraint conditions include total power distribution constraint, single device power distribution constraint, power distribution change constraint and operation data constraint.

[0035] A second aspect of the present invention provides a power distribution device based on a communication power supply system, comprising:

[0036] A data acquisition module, configured to acquire a historical operation data set of the communication power supply system, wherein the historical operation data set includes a plurality of historical operation data within a preset time period;

[0037] The reward construction module is used to construct the target reward function of the initial power distribution model based on the device energy consumption cost data, the device remaining power variance data, and the battery life decay data as optimization targets;

[0038] A model training module, configured to input each of the historical operation data into the initial power supply distribution model for training to obtain a plurality of predicted historical distribution data;

[0039] a model output module, configured to perform reward calculation based on the target reward function using each of the historical operating data and the associated predicted historical distribution data, and determine a target power supply distribution model according to the reward calculation result;

[0040] The electric energy distribution module is used to solve the real-time operation data through the target power distribution model to generate an optimal power distribution plan.

[0041] The third aspect of the present invention provides an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the power distribution method based on the communication power supply system as described in any one of the above items.

[0042] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the power distribution method based on the communication power supply system as described in any one of the above items.

[0043] A fifth aspect of the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the power distribution method based on the communication power supply system as described in any one of the above items.

[0044] It can be seen from the above technical solutions that the present invention has the following advantages:

[0045] This invention constructs a multi-dimensional objective reward function based on equipment energy consumption costs, remaining power balance, and battery life decay, precisely coordinating multiple objectives and overcoming the limitations of a single objective. It also uses historical operating data sets to conduct reinforcement learning training on the initial power distribution model, allowing the model to gradually learn the optimal power supply strategy under different scenarios, ultimately generating a target power distribution model capable of autonomously optimizing decisions. This invention significantly improves the adaptability of the communication power supply system to complex and changing communication network environments, while ensuring power supply reliability and effectively avoiding the power supply redundancy problem caused by static allocation strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0047] Figure 1A flowchart of a power distribution method based on a communication power supply system provided in accordance with the first embodiment of the present invention;

[0048] Figure 2 A flowchart of a power distribution method based on a communication power supply system provided in the second embodiment of the present invention;

[0049] Figure 3 This is a structural diagram of the communication power supply system;

[0050] Figure 4 This is a schematic diagram of reward value convergence comparison;

[0051] Figure 5 This is a structural block diagram of a power distribution device based on a communication power supply system provided in the third embodiment of the present invention;

[0052] Figure 6 This is a structural block diagram of a computer device provided in Example 4 of the present invention. DETAILED DESCRIPTION

[0053] The embodiments of the present invention provide a power distribution method and device based on a communication power supply system, which is used to solve the technical problem that traditional communication power supply systems often rely on fixed rules to distribute electric energy, lack dynamic adjustment capabilities, and lead to power supply redundancy.

[0054] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0055] Affected by multiple factors such as load fluctuations and ambient temperature variations, the energy allocation process of telecommunication power supply systems exhibits significant uncertainty and nonlinear characteristics. Therefore, this paper proposes an intelligent energy allocation method for telecommunication power supply systems based on an improved Deep Q-Network (DQN). This method comprehensively considers the remaining power of the device, grid electricity prices, and environmental factors, defines the state space as a multidimensional vector, and models the allocation problem. Constraints are designed by comprehensively considering factors such as the total power supply of the device, the device power demand, and the battery status. Prioritized experience replay is used to improve the Deep Q-Network, and a loss function is used to update the current network parameters. The trained network is used to learn the energy allocation strategy, and energy allocation is achieved after solving the allocation problem. Experimental test results show that this method achieves higher system power stability and ideal allocation results when implementing intelligent energy allocation.

[0056] See also Figure 1 , Figure 1 This is a flowchart of the steps of a power distribution method based on a communication power supply system provided in Example 1 of the present invention.

[0057] The present invention provides a power distribution method based on a communication power supply system, comprising:

[0058] Step 101: Acquire a historical operation data set of a communication power supply system.

[0059] The communication power supply system refers to a complete system that provides power supply for key facilities such as communication base stations and data centers. The communication power supply system serves as a physical carrier for data collection.

[0060] The historical operation dataset refers to the multi-dimensional time series data related to power supply decisions recorded during the past operation of the communication power supply system. Specifically, it includes the normalized SOC of each energy storage unit, real-time electricity prices, and environmental parameters (such as temperature, humidity, and wind speed). The historical operation dataset serves as a training set for reinforcement learning, covering a sufficiently diverse range of scenarios to train the initial power distribution model, enabling the model to learn the optimal power supply strategy under different scenarios (such as high temperature and high load, low electricity price period, etc.).

[0061] In an embodiment of the present invention, a historical operation data set of a communication power supply system is obtained.

[0062] Step 102: Taking the device energy consumption cost data, the device remaining power variance data, and the battery life attenuation data as optimization targets, a target reward function of the initial power supply allocation model is constructed.

[0063] In an embodiment of the present invention, the target reward function of the initial power distribution model is constructed with the device energy consumption cost data, the device remaining power variance data and the battery life attenuation data as optimization targets.

[0064] Step 103: Use each historical operation data to input the initial power supply distribution model for training to obtain a plurality of predicted historical distribution data.

[0065] In an embodiment of the present invention, various historical operation data are input into an initial power supply distribution model for reinforcement learning training to obtain a plurality of predicted historical distribution data.

[0066] Step 104 : Based on the target reward function, each historical operation data and the associated predicted historical distribution data are used to perform reward calculation, and a target power supply distribution model is determined according to the reward calculation result.

[0067] The initial power distribution model refers to the original decision model that has not been optimized at the beginning of the reinforcement learning training process. Specifically, it refers to a deep Q network (DQN) with a fixed structure but untrained parameters.

[0068] The target power supply allocation model refers to a target power supply allocation model that has the ability to optimize decision-making after the initial model is trained through reinforcement learning using historical operation data sets.

[0069] In an embodiment of the present invention, each historical operation data and the associated predicted historical distribution data are input into a target reward function to perform reward calculation, and a target power supply distribution model is determined according to the reward calculation result.

[0070] Step 105: Solve the real-time operation data using the target power distribution model to generate an optimal power distribution plan.

[0071] In an embodiment of the present invention, first, the real-time operating data is solved by the target power distribution model to obtain a set of initial power distribution schemes, and then the constrained optimization problem is solved based on a preset solver to obtain the optimal power distribution scheme that meets the preset power distribution constraints.

[0072] This invention constructs a multi-dimensional objective reward function based on equipment energy consumption costs, remaining power balance, and battery life decay, precisely coordinating multiple objectives and overcoming the limitations of a single objective. It also uses historical operating data sets to conduct reinforcement learning training on the initial power distribution model, allowing the model to gradually learn the optimal power supply strategy under different scenarios, ultimately generating a target power distribution model capable of autonomously optimizing decisions. This invention significantly improves the adaptability of the communication power supply system to complex and changing communication network environments, while ensuring power supply reliability and effectively avoiding the power supply redundancy problem caused by static allocation strategies.

[0073] See also Figure 2 , Figure 2 This is a flowchart of the steps of a power distribution method based on a communication power supply system provided in the second embodiment of the present invention.

[0074] The present invention provides a power distribution method based on a communication power supply system, comprising:

[0075] Step 201: Acquire a historical operation data set of a communication power supply system, wherein the historical operation data set includes a plurality of historical operation data within a preset time period.

[0076] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a communication power supply system, which includes a perception layer, a decision layer, and an execution layer that are sequentially connected in communication;

[0077] The perception layer includes smart meters, environmental parameter collection sensors, battery management units, and load prediction modules;

[0078] Smart meters are used to collect real-time electricity prices and provide an economic basis for subsequent power distribution decisions;

[0079] Environmental parameter collection sensors include temperature sensors, humidity sensors, and wind speed sensors, which are used to collect temperature, humidity, and wind speed, respectively. Ambient temperature has a significant impact on device operational stability and battery performance. For example, high temperatures can affect device heat dissipation, leading to performance degradation or even failure. Temperature also significantly affects battery charge and discharge efficiency and lifespan. By collecting temperature data, the system can adjust power distribution strategies accordingly to ensure safe and stable device operation and optimize battery management.

[0080] The Battery Management System (BMS) manages battery-related status and parameters, with its core function being to monitor the state of charge (SOC) of each energy storage unit. By accurately understanding the battery's SOC, the system can rationally schedule battery charging and discharging, avoiding overcharging and overdischarging, extending battery life, and ensuring the battery's stable and reliable energy storage function within the power supply system. It is a key management component for the stable operation of energy storage systems.

[0081] The load forecasting module analyzes historical operational data sets to predict future load demand. Accurate load demand forecasts help plan power distribution plans in advance and avoid power shortages or oversupply.

[0082] The decision-making layer includes edge computing devices, which are used to implement the improved target power distribution model. These devices receive various data from the perception layer, analyze and process the data using the deep Q network algorithm corresponding to the target power distribution model, and make power distribution decisions based on the current system status and goals (such as lowest cost, highest power supply stability, etc.).

[0083] The execution layer includes the Programmable Logic Controller (PLC) and power distribution circuits;

[0084] Programmable logic controllers, which are used to precisely control the distribution and transmission of power based on the decisions made by edge computing devices;

[0085] The power distribution circuit is the physical execution channel for power distribution. It is responsible for rationally distributing electrical energy to various electrical devices or energy storage units according to the control instructions of the PLC. By adjusting circuit parameters and other methods, it can achieve power distribution of different power levels. It is the key hardware link that converts power supply decisions into actual power supply operations.

[0086] In the embodiment of the present invention, before the historical operation data set is used to perform reinforcement learning training on the initial power distribution model, it is necessary to define the state space of the historical operation data set. It is necessary to include all factors that may affect the energy allocation decision. In this invention, the remaining battery capacity, real-time electricity price and environmental parameters are comprehensively considered, and the state space is defined as a multidimensional vector, specifically:

[0087]

[0088] Where, represents the state space, Indicates the Devices in The power demand during the period, where the equipment refers to the power distribution equipment connected to the communication power supply system. Indicates the Devices in The remaining battery power during the time period, , Indicates the real-time electricity price, Represents environmental parameters, Indicates the total number of devices.

[0089] Step 202: Taking the device energy consumption cost data, the device remaining power variance data, and the battery life attenuation data as optimization targets, a target reward function of the initial power distribution model is constructed.

[0090] It should be noted that the target reward function The optimization goal needs to be reflected. The present invention sets the power distribution goal of the communication power supply system to improve system stability and extend equipment life while minimizing energy consumption costs. The following target reward function is constructed:

[0091]

[0092] Where, represents the target reward function, Represents the equipment energy consumption cost data, specifically representing the equipment energy consumption cost calculated based on the action and electricity price, where: represents the current time step, Indicates time The forecast historical distribution data under represents the electricity price parameter under time, Represents the variance data of the remaining power of the device, where: represents the next time step, Indicates the battery of the 1st to nth device at time The remaining power set of the device specifically represents the variance of the remaining power of the device, which is used to measure the stability of the system. Represents battery life attenuation data, where Indicates the The battery of the device, in time The remaining power when Indicates calculation of Device battery life When the remaining power is reduced, Represents the sum of the lifespan decay of all n devices, specifically a linear function between battery life and remaining power, used to measure the impact of remaining battery power on device lifespan. 、 、 Represents the weight coefficient, which is used to balance the importance of different optimization objectives.

[0093] The initial power distribution model includes an input layer, a hidden layer, and an output layer. The input layer includes 12 neurons (corresponding to 4 devices × 3 operating parameters), the hidden layer includes 3 fully connected layers (256-128-64 nodes) and ReLU activation functions, and the output layer includes 4 action values Q (s t , a j ), s t represents the state space, a j represents the action space.

[0094] Step 203: Use each historical operation data to input the initial power supply distribution model for training to obtain a plurality of predicted historical distribution data.

[0095] In the embodiment of the present invention, the historical operation data set is used to input the initial power supply distribution model to obtain the corresponding predicted historical distribution data, that is, the action space corresponding to the state space needs to contain all possible energy distribution strategies, and the action space is a discrete space point, each action space is a vector representing each device at time The amount of power allocated within the The details are as follows:

[0096]

[0097] Where, Indicates the Devices in The amount of electricity allocated during the time period.

[0098] Step 204 : Based on the target reward function, each historical operation data and the associated predicted historical distribution data are used to perform reward calculation, and a target power supply distribution model is determined according to the reward calculation result.

[0099] Furthermore, step 204 may include the following sub-steps:

[0100] S11. Calculate rewards for each historical operation data and the associated predicted historical distribution data using a target reward function to obtain multiple reward values.

[0101] In an embodiment of the present invention, the target reward function constructed in step 202 is used to calculate the reward value by inputting each historical operation data and the associated predicted historical distribution data into the target reward function to obtain the reward value corresponding to each historical operation data.

[0102] S12. Calculate the cumulative reward using multiple reward values and a preset discount factor to obtain a cumulative reward value.

[0103] In the embodiment of the present invention, the goal of intelligent energy distribution is to find the optimal power distribution solution so that starting from any initial state, The cumulative discounted rewards can be maximized, so the cumulative reward value expression is as follows:

[0104]

[0105] Where, represents the cumulative reward value, Indicates the expected value corresponding to the preset power distribution scheme, Indicates time range The preset discount factor within Represents the predicted historical distribution data, specifically characterized in time Communication power supply system status The action selected by the power distribution strategy under the state Extracted from historical operation data sets (battery remaining power, real-time electricity prices and environmental parameters), Indicates the time range (i.e. the preset period).

[0106] S13. When the cumulative reward value is greater than a preset cumulative threshold, the initial power supply distribution model is used as the target power supply distribution model.

[0107] In an embodiment of the present invention, when the cumulative reward value is greater than the preset cumulative threshold, that is, when convergence is reached, it means that the initial power supply distribution model has fully learned the power supply distribution strategy that can effectively cope with various operating conditions during the reinforcement learning training process, and the reward obtained has met the expected goal. Therefore, the current initial power supply distribution model is used as the target power supply distribution model.

[0108] Furthermore, step 204 may also include the following sub-steps:

[0109] S14. When the cumulative reward value is less than or equal to the preset cumulative threshold, each predicted historical distribution data is used to calculate the loss value to obtain the loss value.

[0110] In this embodiment of the present invention, when the cumulative reward value is less than or equal to the preset cumulative threshold, it means that the performance of the current initial power distribution model has not yet reached the expected standard. In this case, the loss value is calculated using the predicted historical distribution data to obtain the loss value. The calculation of the loss value can quantify the gap between the current strategy of the model and the optimal strategy. The specific expression of the loss value is:

[0111]

[0112] Where, represents the loss value, Indicates the total number of historical running data in the historical running data set, that is, the number of batch samples; Indicates the The standard state action value of the historical operation data is specifically the target state action value under the optimal power supply allocation strategy. For example, the theoretical optimal value of the execution action under the optimal power supply allocation strategy is generated through expert experience as the standard state action value. Indicates that the current model Forecast historical allocation data The value prediction of the current initial power distribution model is Next action ( is one of the four candidate actions) (of the four action values in the output layer, only the actual action is used to calculate the loss), Represents the network parameters of the current initial power distribution model.

[0113] It is worth mentioning that the output layer contains 4 action values In the case of loss calculation, the loss calculation focuses on the actual execution actions in the historical operation data. The corresponding predicted value, In the historical operation data, if the execution action is (belongs to one of the four actions), then only the predicted value corresponding to this action is calculated and standard state action value The predicted values of the other three unexecuted actions do not participate in the loss calculation of the current historical running data (because the loss needs to align with the value deviation of the actual executed action). Therefore, the loss function It is actually the average error between the predicted value and the standard value of the actual execution action of each sample in the batch sample (historical operation data set).

[0114] S15. With the goal of minimizing the loss value, the network parameters of the initial power supply distribution model are adjusted, and a priority experience replay strategy is used to update the sample priority of multiple historical operation data.

[0115] The priority experience replay strategy specifically adjusts the sampling probability of multiple historical operation data input into the initial power supply distribution model for training, so that the historical operation data with higher priority has a higher probability of being selected for training.

[0116] The goal is to minimize the loss value, which means using the gradient descent method to minimize the loss function and iteratively updating the network parameters of the initial power supply distribution model through back propagation. When the loss value converges to the preset threshold, the optimization process is terminated. At this time, the model's prediction deviation of the power supply strategy has met the actual application requirements.

[0117] It should be noted that in actual training, when the loss value drops to the preset threshold, optimization is stopped. This is not the pursuit of the mathematical "global minimum" (because data and computing power are limited in actual scenarios, the global optimum is difficult to achieve).

[0118] It's important to note that the prioritized experience replay strategy optimizes training efficiency by assigning different sampling priorities to experiences. Its core idea is to assign a higher sampling probability to experiences with higher learning value (such as high TD errors or key decision points), thereby improving model learning efficiency and accelerating convergence.

[0119] Furthermore, S15 may include the following sub-steps:

[0120] S151. Adjust the network parameters of the initial power supply distribution model with the goal of minimizing the loss value.

[0121] In the embodiment of the present invention, with the goal of minimizing the loss value, after calculation using the above loss value expression, the back propagation algorithm is used to adjust the network parameters of the initial power distribution model.

[0122] S152: Using network parameters, historical operation data, and associated predicted historical allocation data, determine the target state action value corresponding to each historical operation data.

[0123] In the embodiment of the present invention, based on the updated network parameters, the state-action sequence in the historical operation data is used as input, and the actual distribution results after the execution of the action (such as energy consumption, power change, etc.) reflected in the predicted historical distribution data are used as the constraint conditions for the model to calculate the long-term value, and the target state action corresponding to the updated initial power distribution model is obtained. Here, the target state action is also used. The purpose of predicting historical distribution data is to enable the model to learn the actual impact of actions and avoid disconnection between value assessment and real scenarios. In other words, the predicted historical distribution data is the output of the model, which is used to constrain / calibrate the value calculation and ensure the target state action value. It not only reflects the model prediction, but also fits the actual operating rules of the communication power supply system.

[0124] S153: Calculate the error using the immediate reward value, the target state action value, the preset discount factor, and the preset standard state action value to obtain the target error value corresponding to each historical operation data.

[0125] Furthermore, S153 may include the following sub-steps:

[0126] S1531. Perform a multiplication operation on the preset discount factor and the preset standard state action value to obtain a target multiplication value corresponding to each historical operation data.

[0127] S1532: Perform a sum operation using the instant reward value and the target multiplier value to obtain the target sum value corresponding to each historical operation data.

[0128] S1533. Perform a difference operation on the target sum value and the target state action value to obtain the target error value corresponding to each historical operation data.

[0129] In the specific implementation, the above process is encapsulated into the form of a formula, as follows:

[0130]

[0131] Where, represents the target error value, Indicates the immediate reward value, specifically in the state Next action After that, the immediate reward obtained from the environment is obtained through the target reward function associated with the above-mentioned target power distribution model (that is, the updated initial power distribution model). represents the target state action value, Indicates the preset standard state action value. It should be noted that the preset standard state action value is the preset standard power supply distribution network in the state Next action Multiple initial state action values are calculated, and then the maximum value is selected as the preset standard state action value.

[0132] S154. Calculate the priority using the absolute value of the target error value and a preset parameter factor to obtain a priority value corresponding to each historical operation data.

[0133] In the specific implementation, the above process is encapsulated into the form of a formula, as follows:

[0134]

[0135] Where, Indicates the priority value, Indicates the preset parameter factor, select a very small positive number to avoid the priority value being 0;

[0136] S155. Perform a sum operation using all priority values to obtain a priority value sum.

[0137] In the embodiment of the present invention, all priority values are used to perform sum operation to obtain the priority value and value. .

[0138] S156. Perform a ratio operation on each priority value and the priority value sum to obtain a sampling probability corresponding to each historical operation data.

[0139] In the specific implementation, the above process is encapsulated into the form of a formula, as follows:

[0140]

[0141] Where, Indicates the The sampling probability corresponding to the historical running data, Indicates the The priority value corresponding to each historical operation data.

[0142] S157. Sort the historical operation data from large to small according to the sampling probability.

[0143] In an embodiment of the present invention, historical operation data are sorted from large to small according to sampling probability and from large to small according to priority value. High-priority samples (large errors and more critical to model optimization) are sampled first and input at high frequency. For example, a sample with a high priority value may be selected 8 times in 10 training rounds. A model that originally required 100 rounds of training to converge may only need 70 rounds to meet the standard, thereby shortening the training cycle, while improving the prediction accuracy of key scenarios (such as power supply anomalies and high battery load) and reducing strategy errors.

[0144] S158. Calculate the sampling probability interval corresponding to each historical operation data after sorting.

[0145] In the embodiment of the present invention, the sampling probability interval corresponding to each sorted historical operation data is calculated.

[0146] In the specific implementation, the above process is encapsulated into the form of a formula, as follows:

[0147]

[0148] Where, Indicates the starting point of the interval, Indicates the upper limit of the sampling probability interval, Represents the total number of sampled probabilities.

[0149] For ease of understanding, let's take three historical operating data as an example:

[0150] The sampling probability of the historical running data ranked first is 5%;

[0151] The sampling probability of the historical running data ranked second is 4%;

[0152] The sampling probability of the historical running data ranked third is 3%;

[0153] Then, the sampling probability interval of the historical running data ranked first is (0, 0.05);

[0154] The sampling probability interval of the second ranked historical running data is [0.05, 0.09);

[0155] The sampling probability interval of the historical running data ranked third is [0.09, 0.12).

[0156] S159 , randomly generate multiple random numbers, and use each random number to perform interval search, matching the sampling probability interval where each random number is located as the target sampling probability interval.

[0157] In the embodiment of the present invention, a plurality of random numbers between 0 and 1 are randomly generated, and then interval retrieval is performed using the random numbers to determine which sampling probability interval each random number falls within, and then use it as the target sampling probability interval.

[0158] S1510 , selecting historical operation data associated with the target sampling probability interval to form an updated historical operation data set.

[0159] In the embodiment of the present invention, the historical operation data associated with these target sampling probability intervals are selected to form a new historical operation data set, and the sample priority update is completed. This new historical operation data set is the updated historical operation data set.

[0160] It is worth mentioning that the experience replay mechanism allows the model to reuse historical running data sets multiple times, breaking the data usage restrictions of a single training session, allowing historical running data to be repeatedly called during multiple rounds of model optimization, fully releasing the policy learning value of the data, and avoiding one-time consumption of data value. Then, based on dynamic selection of sample priorities, it can focus on mining samples that are more valuable for model optimization (such as samples with large prediction deviations and those that can reduce losses), extracting more effective information from limited historical running data, and improving data utilization.

[0161] S16: Based on the historical operation data set after the sample priority is updated, jump to the step of using the historical operation data set to input the initial power supply distribution model for training to obtain predicted historical distribution data.

[0162] In the embodiment of the present invention, based on the historical operation data set after the sample priority is updated, the process jumps to the step of using the historical operation data set to input the initial power supply distribution model to obtain predicted historical distribution data.

[0163] It should be noted that when sampling, weighted random sampling is performed according to priority, and "important" samples (that is, historical running data sets) are selected for training with a higher probability, which can improve the learning speed of the deep Q network. Then it is trained in a loop. For each time step, the current state is first observed and the action is selected according to the current strategy. And execute, observe the new state and instant rewards . Then from the experience replay buffer A batch of samples is sampled according to the priority.

[0164] This invention constructs a multi-dimensional objective reward function based on equipment energy consumption costs, remaining power balance, and battery life decay, precisely coordinating multiple objectives and overcoming the limitations of a single objective. It also uses historical operating data sets to conduct reinforcement learning training on the initial power distribution model, allowing the model to gradually learn the optimal power supply strategy under different scenarios, ultimately generating a target power distribution model capable of autonomously optimizing decisions. This invention significantly improves the adaptability of the communication power supply system to complex and changing communication network environments, while ensuring power supply reliability and effectively avoiding the power supply redundancy problem caused by static allocation strategies.

[0165] Step 205: Solve the real-time operation data using the target power distribution model to generate an optimal power distribution plan that meets the preset power distribution constraints.

[0166] The preset power distribution constraints include total power distribution constraints, single device power distribution constraints, power distribution change constraints and operation data constraints.

[0167] It should be noted that it is necessary to ensure that the total amount of power allocated to all devices at any time point does not exceed the available power. The total power allocation constraints are as follows:

[0168]

[0169] Where, represents the total available power in period t, Indicates the preset available power.

[0170] Considering that each device has its own corresponding power demand range, its corresponding allocated power must also meet this range. Therefore, the power allocation constraints for a single device are specifically as follows:

[0171]

[0172] Where, Indicates the Devices in Minimum power demand during the time period, Indicates the Devices in The maximum power demand during the time period.

[0173] Operational data constraints include battery remaining power safety constraints and real-time electricity price constraints;

[0174] The remaining battery capacity of the communication power supply system should be kept within a safe range to avoid overcharging or over-discharging. Therefore, the specific safety constraints for the remaining battery capacity are as follows:

[0175]

[0176] Where, Indicates the Devices in The remaining battery power during the time period, Indicates the minimum safe power of the battery. Indicates the maximum safe capacity of the battery.

[0177] Set a static threshold for real-time electricity prices , so the real-time electricity price constraint is:

[0178]

[0179] Where, Indicates the maximum total electricity consumption allowed when electricity prices are high. represents the static threshold, Indicates the set of devices to be powered.

[0180] The power distribution change between adjacent time steps is set not to exceed a certain threshold, so the power distribution change constraint is specifically:

[0181]

[0182] Where, Indicates the Devices in The amount of electricity allocated during the period, Indicates the maximum allowed value of power distribution change.

[0183] In an embodiment of the present invention, first, the real-time operating data is solved by using a target power distribution model to obtain a set of initial power distribution plans. Then, based on a model predictive control rolling optimization solver, such as a quadratic programming solver (QP) and a nonlinear programming solver (NLP), the constrained optimization problem is solved within a short time window to obtain an optimal power distribution plan that meets the preset power distribution constraints.

[0184] The following is an experimental example:

[0185] In this experiment, power system simulation software constructed a medium-sized communication base station environment. This environment contained 10 communication devices, including five high-power base station communication devices (such as 4G / 5G base stations) and five low-power IoT sensor devices. The base station communication devices consumed an average of approximately 2000W, while the IoT sensor devices consumed an average of approximately 5W. Two high-capacity battery packs, each with a capacity of 500kWh, served as energy storage units.

[0186] By simulating three different application scenarios, adjusting the parameters of the communication equipment load and the initial battery charge, the allocation effect of different allocation strategies in actual application scenarios was tested. The specific application scenario parameter configuration is shown in Table 1:

[0187] Table 1. Parameter configuration comparison table for application scenarios

[0188]

[0189] Under the power distribution method of the present invention, the comparison results of the algorithm reward values before and after the improvement of the deep Q network are as follows: Figure 4 shown.

[0190] It can be seen from the above experimental results that after the deep Q network is improved by adopting priority experience replay and target reward function, the algorithm has higher convergence than the original algorithm. This is mainly because the priority experience replay mechanism enables the model to be more frequently exposed to experiences that are critical to improving performance, thereby accelerating the learning process. In order to improve the reliability of the experimental results, this experiment recorded the system power supply quality under different energy allocation strategies. Among them, Scheme A is a fixed-cycle equal distribution strategy, specifically, it evenly distributes the power supply according to a 24-hour cycle, does not distinguish between load fluctuations, and only guarantees basic power supply. Scheme B is a single-objective Q learning strategy, specifically based on the original DQN framework, and only uses the lowest device energy consumption cost as the reward target for training. Scheme C is a multi-objective deep Q network strategy optimized by priority experience replay in the present invention, specifically using device energy consumption, power variance, and battery life attenuation as reward targets. Learning is accelerated by priority experience replay. After energy allocation for the same scenario, the higher the system power supply quality, the better the allocation effect of the representative strategy. The specific experimental results are shown in Table 2:

[0191] Table 2. Comparison table of experimental results

[0192]

[0193] This paper uses two indicators, the power supply stability index (SSI) and energy efficiency, to measure power supply quality. The experimental results clearly show that the proposed intelligent energy allocation strategy significantly improves power supply stability and energy efficiency by adjusting and optimizing energy allocation in real time, offering significant advantages over traditional fixed allocation strategies.

[0194] A communication base station scenario with 10 devices was constructed in the simulation platform. The test results are shown in Table 3 below:

[0195] Table 3. Test results comparison table

[0196]

[0197] In a test of 10 communication base station scenarios built on a simulation platform, the performance of various indicators in different scenarios highlighted advantages. In high-load scenarios, power supply stability reached 98.7%, cost reduction rate reached 21.3%, and response delay was only 0.48 seconds. In extreme electricity price scenarios, the cost reduction rate was particularly outstanding, reaching 34.7%, power supply stability remained at 96.5%, and response delay was 0.51 seconds. In low-load scenarios, power supply stability reached 99.1%, cost reduction rate was 12.5%, and response delay was as low as 0.45 seconds. Overall, the test results show that in different complex scenarios, the system has better performance in power supply stability, cost control, and response speed.

[0198] The present invention improves the efficiency of power supply management through a variety of innovative technologies: adopting a dynamic priority experience replay mechanism, adjusting sample priority based on TD error to improve training efficiency; constructing a multi-objective composite reward function to collaboratively optimize electricity price cost, power balance and battery life loss; using real-time electricity price threshold response technology to automatically switch power supply modes according to electricity price fluctuations to reduce operating costs; and performing multi-dimensional state fusion modeling to integrate four-dimensional parameters of equipment, environment, load and power grid to enhance state characterization capabilities.

[0199] By introducing a reinforcement learning framework, this invention continuously evaluates the comprehensive performance of power distribution strategies under multiple objectives, including energy cost, power balance, and battery life. It also autonomously optimizes the decision-making model based on a cumulative reward mechanism, enabling intelligent dynamic adjustment of power distribution strategies. Compared to traditional fixed-rule distribution methods, this invention significantly reduces system operating costs and extends battery life while ensuring power supply reliability. It also improves the balance and rationality of power distribution, avoids the power supply redundancy issues caused by static distribution strategies, and provides a more intelligent and efficient power management solution for communication networks.

[0200] See also Figure 5 , Figure 5 This is a structural block diagram of a power distribution device based on a communication power supply system provided in Example 3 of the present invention.

[0201] The present invention provides a power distribution device based on a communication power supply system, comprising:

[0202] The data acquisition module 301 is used to acquire a historical operation data set of the communication power supply system, wherein the historical operation data set includes a plurality of historical operation data within a preset time period;

[0203] The reward construction module 302 is used to construct a target reward function of the initial power distribution model using the device energy consumption cost data, the device remaining power variance data, and the battery life decay data as optimization targets;

[0204] The model training module 303 is used to use various historical operation data to input the initial power distribution model for training to obtain multiple predicted historical distribution data;

[0205] The model output module 304 is configured to calculate rewards based on the target reward function using the historical operating data and the associated predicted historical distribution data, and determine the target power distribution model according to the reward calculation results;

[0206] The power distribution module 305 is used to solve the real-time operation data through the target power distribution model to generate an optimal power distribution plan.

[0207] Furthermore, the model output module 304 includes:

[0208] A reward value submodule is used to calculate rewards for each historical operation data and the associated predicted historical distribution data using a target reward function to obtain multiple reward values;

[0209] A cumulative reward value submodule is used to calculate the cumulative reward using multiple reward values and a preset discount factor to obtain a cumulative reward value;

[0210] The target power supply allocation model submodule is used to use the initial power supply allocation model as the target power supply allocation model when the cumulative reward value is greater than a preset cumulative threshold.

[0211] Furthermore, the model output module 304 further includes:

[0212] The loss value submodule is used to calculate the loss value using each predicted historical distribution data when the cumulative reward value is less than or equal to the preset cumulative threshold to obtain the loss value;

[0213] The sample priority update submodule is used to adjust the network parameters of the initial power distribution model with the goal of minimizing the loss value, and to update the sample priority of multiple historical operation data;

[0214] The jump submodule is used to jump to the step of using the historical operation data set to input the initial power supply distribution model for training based on the historical operation data set after the sample priority is updated, so as to obtain the predicted historical distribution data.

[0215] Furthermore, the sample priority update submodule includes:

[0216] A target state action value unit is used to determine a target state action value corresponding to each historical operation data by using network parameters, historical operation data and associated predicted historical distribution data;

[0217] A target error value unit is used to calculate the error using the immediate reward value, the target state action value, the preset discount factor, and the preset standard state action value to obtain the target error value corresponding to each historical operation data;

[0218] The priority value unit is used to calculate the priority using the absolute value of the target error value and the preset parameter factor to obtain the priority value corresponding to each historical operation data;

[0219] A priority value sum value unit is used to perform a sum operation using all priority values to obtain a priority value sum value;

[0220] The sampling probability unit is used to perform a ratio operation between each priority value and the priority value and value to obtain a sampling probability corresponding to each historical operation data;

[0221] A sorting unit is used to sort the historical operation data from large to small according to the sampling probability;

[0222] The sampling probability interval unit is used to calculate the sampling probability interval corresponding to each historical operation data after sorting;

[0223] The target sampling probability interval unit is used to randomly generate multiple random numbers, and use each random number to perform interval retrieval, matching the sampling probability interval of each random number as the target sampling probability interval;

[0224] The updating unit is used to select historical operation data associated with the target sampling probability interval to form an updated historical operation data set.

[0225] Furthermore, the target error value unit includes:

[0226] A target multiplication subunit is used to perform a multiplication operation using a preset discount factor and a preset standard state action value to obtain a target multiplication value;

[0227] A target sum subunit, configured to perform a sum operation using the instant reward value and the target multiplier value to obtain a target sum value;

[0228] The target error value operation subunit is used to perform a difference operation between the target sum value and the target state action value to obtain a target error value.

[0229] Furthermore, the power distribution module 305 includes:

[0230] The constraint solving submodule is used to solve the real-time operation data through the target power distribution model to generate the optimal power distribution plan that meets the preset power distribution constraints;

[0231] The preset power distribution constraints include total power distribution constraints, single device power distribution constraints, power distribution change constraints and operation data constraints.

[0232] This invention constructs a multi-dimensional objective reward function based on equipment energy consumption costs, remaining power balance, and battery life decay, precisely coordinating multiple objectives and overcoming the limitations of a single objective. It also uses historical operating data sets to conduct reinforcement learning training on the initial power distribution model, allowing the model to gradually learn the optimal power supply strategy under different scenarios, ultimately generating a target power distribution model capable of autonomously optimizing decisions. This invention significantly improves the adaptability of the communication power supply system to complex and changing communication network environments, while ensuring power supply reliability and effectively avoiding the power supply redundancy problem caused by static allocation strategies.

[0233] See also Figure 6 , Figure 6 This is a structural block diagram of a computer device provided in Example 4 of the present invention.

[0234] An electronic device according to an embodiment of the present invention includes: a memory 401 and a processor 402, wherein the memory 401 stores a computer program; when the computer program is executed by the processor 402, the processor 402 executes a power distribution method based on a communication power supply system according to any of the above embodiments.

[0235] Memory 401 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Memory 401 has storage space 403 for program code 413 for executing any of the method steps described above. For example, storage space 403 for program code may include individual program codes 413 for implementing various steps in the method described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. The program codes may be compressed, for example, in a suitable format. When executed by a processing device, these codes cause the processing device to execute the various steps in the method described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. The program codes may be compressed, for example, in a suitable format. When these codes are executed by a computing and processing device, the computing and processing device is caused to execute the various steps of the above-described power distribution method based on the communication power supply system.

[0236] The fifth embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the power distribution method based on the communication power supply system according to any of the above embodiments is implemented.

[0237] Embodiment 6 of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the power distribution method based on the communication power supply system as any of the above embodiments.

[0238] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0239] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0240] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0241] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0242] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0243] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A power distribution method based on a communication power supply system, characterized in that: include: Acquire a historical operation data set of the communication power supply system, wherein the historical operation data set includes a plurality of historical operation data within a preset time period; Using device energy consumption cost data, device remaining power variance data, and battery life attenuation data as optimization targets, we construct the target reward function of the initial power distribution model. Using each of the historical operation data to input the initial power supply distribution model for training to obtain a plurality of predicted historical distribution data; Based on the target reward function, using each of the historical operating data and the associated predicted historical distribution data to perform reward calculation, and determining a target power supply distribution model according to the reward calculation result; Solving the real-time operation data through the target power distribution model to generate an optimal power distribution plan; The reward calculation is performed based on the target reward function using each of the historical operation data and the associated predicted historical distribution data, and a target power supply distribution model is determined according to the reward calculation result, including: Performing reward calculation on each of the historical operation data and the associated predicted historical distribution data using the target reward function to obtain a plurality of reward values; Calculating a cumulative reward using the plurality of reward values and a preset discount factor to obtain a cumulative reward value; When the cumulative reward value is greater than a preset cumulative threshold, the initial power supply distribution model is used as a target power supply distribution model; Also includes: When the cumulative reward value is less than or equal to the preset cumulative threshold, the loss value is calculated using each of the predicted historical distribution data to obtain a loss value; With the goal of minimizing the loss value, adjusting the network parameters of the initial power distribution model and performing sample priority updates on the plurality of historical operation data; Based on the historical operation data set after the sample priority is updated, jumping to the step of using each of the historical operation data to input the initial power supply distribution model for training to obtain multiple predicted historical distribution data; Solving the real-time operation data using the target power distribution model to generate an optimal power distribution plan includes: Solving the real-time operation data by using the target power distribution model to generate the optimal power distribution plan that meets the preset power distribution constraints; The preset power supply distribution constraint conditions include total power distribution constraint, single device power distribution constraint, power distribution change constraint and operation data constraint.

2. The power distribution method based on the communication power supply system according to claim 1, characterized in that: The updating of sample priorities of the plurality of historical operation data includes: Determining a target state action value corresponding to each of the historical operation data using the network parameters, the historical operation data, and the associated predicted historical allocation data; An error calculation is performed using the instant reward value, the target state action value, the preset discount factor, and the preset standard state action value to obtain a target error value corresponding to each of the historical operation data; The priority value corresponding to each of the historical operation data is obtained by performing a priority calculation using the absolute value of the target error value and a preset parameter factor; Performing a sum operation using all the priority values to obtain a priority value sum; Performing a ratio operation on each priority value and the priority value sum to obtain a sampling probability corresponding to each historical operation data; Sorting the historical operation data from large to small according to the sampling probability; Calculate the sampling probability interval corresponding to each of the sorted historical operation data; Randomly generate multiple random numbers, and use each of the random numbers to perform interval retrieval, matching the sampling probability interval of each of the random numbers as the target sampling probability interval; The historical operation data associated with the target sampling probability interval is selected to form an updated historical operation data set.

3. The power distribution method based on the communication power supply system according to claim 2, characterized in that: The instant reward value, the target state action value, the preset discount factor and the preset standard state action value are used to perform error calculation to obtain the target error value corresponding to each of the historical operation data, including: Performing a multiplication operation using the preset discount factor and the preset standard state action value to obtain a target multiplication value corresponding to each of the historical operation data; Performing a sum operation using the instant reward value and the target multiplier value to obtain a target sum value corresponding to each of the historical operation data; The target sum value and the target state action value are used to perform a difference operation to obtain a target error value corresponding to each of the historical operation data.

4. A power distribution device based on a communication power supply system, characterized in that: include: A data acquisition module, configured to acquire a historical operation data set of the communication power supply system, wherein the historical operation data set includes a plurality of historical operation data within a preset time period; The reward construction module is used to construct the target reward function of the initial power distribution model based on the device energy consumption cost data, the device remaining power variance data, and the battery life decay data as optimization targets; A model training module, configured to input each of the historical operation data into the initial power supply distribution model for training to obtain a plurality of predicted historical distribution data; a model output module, configured to perform reward calculation based on the target reward function using each of the historical operating data and the associated predicted historical distribution data, and determine a target power supply distribution model according to the reward calculation result; An electric energy distribution module is used to solve the real-time operation data through the target power distribution model to generate an optimal power distribution plan; The model output module includes: a reward value submodule, configured to perform reward calculation on each of the historical operation data and the associated predicted historical distribution data using the target reward function to obtain a plurality of reward values; A cumulative reward value submodule, configured to calculate a cumulative reward using the plurality of reward values and a preset discount factor to obtain a cumulative reward value; a target power supply allocation model submodule, configured to use the initial power supply allocation model as a target power supply allocation model when the cumulative reward value is greater than a preset cumulative threshold; The model output module also includes: a loss value submodule, configured to calculate a loss value using each of the predicted historical distribution data to obtain a loss value when the cumulative reward value is less than or equal to the preset cumulative threshold; a sample priority updating submodule, configured to adjust the network parameters of the initial power distribution model with the goal of minimizing the loss value, and to update the sample priorities of the plurality of historical operation data; a jump submodule, configured to jump to the step of inputting each of the historical operation data into the initial power distribution model for training to obtain a plurality of predicted historical distribution data based on the historical operation data set after the sample priority is updated; The power distribution module includes: A constraint solving submodule, configured to solve the real-time operation data using the target power distribution model to generate the optimal power distribution solution that satisfies the preset power distribution constraint conditions; The preset power supply distribution constraint conditions include total power distribution constraint, single device power distribution constraint, power distribution change constraint and operation data constraint.

5. An electronic device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the power distribution method based on the communication power supply system according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the power distribution method based on the communication power supply system according to any one of claims 1 to 3 is implemented.

7. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer is caused to execute the power distribution method based on the communication power supply system according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Power equipment distribution method and device, electronic equipment and computer readable medium

    CN115759444A

  • Power distribution method and device, computer equipment and readable storage medium

    CN118627822A