5G base station energy storage scheduling method and device based on mathematical optimization and reinforcement learning

By combining the two-stage scheduling method of mathematical optimization and reinforcement learning, the Time2Vec-Decision Transformer model is used to perform 5G base station energy storage scheduling, solving the problems of idle energy storage and uncertain scheduling of 5G base station energy storage, and achieving the stability and economic improvement of power grid operation.

CN120338451BActive Publication Date: 2025-08-22ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510823990.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-08-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

5G base station energy storage batteries are idle for a long time in a high-reliability grid environment and have not fully utilized their value. Traditional scheduling methods rely on accurate predictions to be easily affected by environmental fluctuations, and large-scale base station scheduling increases system uncertainty, requiring higher environmental adaptability and anti-interference capabilities.

Method used

Combining mathematical optimization and reinforcement learning, the Time2Vec-Decision Transformer model is used for two-stage scheduling. Recently, planning and intraday correction are combined, and action correction is used to use real-time information of the power grid to improve the environmental adaptability and anti-interference ability of decision-making.

Benefits of technology

It realizes precise scheduling in the case of fluctuations in the power grid environment, improves the economy and safety of power grid operation, reduces operating costs, stabilizes the power fluctuations in the power grid, and fully utilizes the energy storage adjustment capabilities of 5G base stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338451B_ABST
    Figure CN120338451B_ABST
Patent Text Reader

Abstract

The present invention discloses a 5G base station energy storage scheduling method and device based on mathematical optimization and reinforcement learning, which belongs to the field of power grid scheduling. The objective function of the day-ahead optimization model is constructed to obtain the day-ahead scheduling action vector; based on the intraday real-time information of the power grid environment, the real-time state vector and reward value are obtained; based on the model, the corrected final action vector is obtained and the power grid system is adjusted to obtain the adjusted real-time state vector; the decision sequence sample sets at multiple moments are added to the sample pool, and the Time2Vec‑Decision Transformer model is optimized and trained; after optimization, the power grid is made in real time to complete the scheduling of the power grid system. The day-ahead-intraday scheduling method is adopted to realize the process of the power grid making decisions using real-time state information, improve the algorithm's ability to resist random fluctuations, and enable the algorithm to give more accurate results than the predicted information scheduling in a fluctuating environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power grid dispatching, and in particular relates to a 5G base station energy storage dispatching method and device based on mathematical optimization and reinforcement learning. Background Art

[0002] 5G base stations are equipped with energy storage batteries during construction to ensure power supply stability and communication service quality. They also possess a certain degree of power consumption regulation capability, offering potential advantages in participating in grid dispatch, balancing power loads, and promoting the absorption of new energy. However, in the current highly reliable power grid environment, the energy storage batteries in 5G base stations are often idle and float-charged, serving only as a backup power source. This results in the failure to fully utilize the energy storage value, necessitating the exploration of a reasonable dispatch mechanism to enable more effective participation in grid regulation. Furthermore, because the power load of 5G base stations is closely related to communication service demand, large-scale base station dispatching increases system uncertainty, placing higher demands on the environmental adaptability and anti-interference capabilities of the dispatching algorithm.

[0003] Currently, extensive research is underway both domestically and internationally on strategies for integrating 5G base station energy storage into grid dispatch. Traditional dispatch methods primarily rely on mathematical optimization algorithms. These methods offer advantages in their robust theoretical frameworks and interpretable results. However, these methods rely heavily on precise forecast data and are susceptible to fluctuations in the external environment, which can lead to deviations from the optimal solution. A rational intraday correction method is urgently needed that can implement real-time action correction and rapid adjustments.

[0004] With the rapid development of artificial intelligence (AI) technology, reinforcement learning (RL) has garnered widespread attention for its exceptional performance in sequential decision-making. RL can learn scheduling patterns from environmental data, with decision-making capabilities comparable to expert experience. It can provide efficient adjustment plans in real time and exhibits strong robustness against uncertainty. Applying RL to the intraday real-time optimization of 5G base station energy storage can provide more accurate results than day-ahead scheduling in the face of fluctuating grid operating conditions, thereby improving the economic efficiency and safety of grid operations. Summary of the Invention

[0005] The purpose of the present invention is to address the deficiencies of the existing technology and provide a 5G base station energy storage scheduling method and device based on mathematical optimization and reinforcement learning.

[0006] The object of the present invention is to achieve the following technical solution: a 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning, comprising the following steps:

[0007] (1) Construct the objective function of the day-ahead optimization model and minimize it to obtain the time The day-ahead scheduling action vector;

[0008] (2) Based on the real-time information of the power grid environment during the day, obtain the time of the power grid environment The real-time state vector is calculated and the time is obtained The reward value;

[0009] (3) Interception time Before The reward value at the moment, the real-time state vector and the intraday corrected action vector are combined with the moment The real-time state vector and reward value of the time are obtained based on the Time2Vec-Decision Transformer model. The intraday correction action vector and the time The day-ahead scheduling action vector is added to get the time The final corrected motion vector of

[0010] (4) According to the time The corrected final action vector is used to adjust the power grid system and obtain the time The adjusted real-time state vector;

[0011] (5) Repeat steps (1) to (4) to obtain a decision sequence sample set consisting of reward values, real-time state vectors, corrected final action vectors, and adjusted real-time state vectors at multiple moments and add them to the sample pool;

[0012] (6) Randomly select B decision sequence sample sets from the sample pool to train the Time2Vec-DecisionTransformer model to obtain the optimized Time2Vec-DecisionTransformer model;

[0013] (7) The optimized Time2Vec-Decision Transformer model is used for real-time decision-making of the power grid to complete the scheduling of the power grid system.

[0014] Furthermore, the step (1) specifically includes the following sub-steps:

[0015] (1.1) Construct the objective function of the day-ahead optimization model based on the forecast vector of user load power, the purchase price vector, and the sales price vector at time T;

[0016] (1.2) Minimizing the objective function based on the charge and discharge power constraints of the 5G base station energy storage system to obtain a set of scheduling actions for the day-ahead phase at T time points. The set of scheduling actions for the day-ahead phase at T time points includes the predicted values ​​of the day-ahead injection power of all 5G base stations at T time points.

[0017] (1.3) Then, the time is selected from the set of scheduling actions in the day-ahead phase of T time The day-ahead scheduling action vector.

[0018] Furthermore, the objective function of the day-ahead optimization model is for ,in, represents the first optimization objective function, represents the second optimization objective function, represents the normalization function, represents the first weighted weight, represents the second weighted weight;

[0019] The first optimization objective function It is constructed based on the prediction vector of user load power at T moments;

[0020] The second optimization objective function It is constructed based on the prediction vector of user load power at T moments, the electricity purchase price vector and the electricity sales price vector.

[0021] Furthermore, the moment of the power grid environment The real-time state vector contains the time ,time The actual value of the grid injection power, electricity purchase price, electricity sales price, user load power and the information collection of all 5G base stations at the time The information set of all 5G base stations includes the time The actual values ​​of the power status and communication load rate of all 5G base stations at that time.

[0022] Furthermore, the step (3) specifically includes the following sub-steps:

[0023] (3.1) From the historical sequence of decisions, extract the moment Before The reward value at the moment, the real-time state vector and the intra-day correction action vector are then filled with zeros, and the position of the zero filling is recorded with the position mask vector; then and the moment The reward value is combined with the real-time state vector to form The reward value sequence and real-time state sequence at each moment and The intraday correction action sequence of each moment; and construct the A time index sequence of moments;

[0024] (3.2) The Time2Vec-Decision Transformer model includes a Time2Vec model, a linear layer model, and a Decision Transformer model;

[0025] The time index sequence is then encoded using the Time2Vec model to obtain an encoded time index sequence; the reward value sequence, real-time state sequence, and intra-day correction action sequence are then encoded using the linear layer model to obtain an encoded reward value sequence, real-time state sequence, and intra-day correction action sequence.

[0026] (3.3) The encoded time index sequence is then added to the encoded reward value sequence, real-time state sequence, and intra-day corrected action sequence respectively to obtain the time-embedded reward value sequence, real-time state sequence, and intra-day corrected action sequence, and then concatenated into the total sequence after time processing;

[0027] (3.4) Then the total sequence and position mask vector after time processing are input into the Decision Transformer model, and the output is The intraday correction action sequence of the moment and obtain the moment The intraday correction action vector of

[0028] (3.5) Change the time The day-ahead scheduling action vector and time Add the intraday correction action vectors to get the moment The corrected final motion vector.

[0029] The present invention also includes a 5G base station energy storage scheduling device based on mathematical optimization and reinforcement learning, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used for the above-mentioned 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning.

[0030] The present invention also includes a computer-readable storage medium on which a program is stored. When the program is executed by a processor, it implements the above-mentioned 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning.

[0031] The beneficial effects of the present invention are:

[0032] 1) This paper establishes a scheduling framework that integrates mathematical optimization day-ahead planning and reinforcement learning intraday correction for the power grid dispatching environment of 5G base station energy storage. Through two-stage scheduling, day-ahead and intraday, it can adjust the optimal strategy in real time according to the power grid environment. It fully utilizes the interpretability of power grid physical model theory and the environmental adaptability of reinforcement learning data-driven models to achieve the complementary advantages of physical and data models and improve the performance of decision-making actions.

[0033] 2) Based on the Decision Transformer reinforcement learning framework, this paper introduces Time2Vec temporal embedding coding and proposes a Time2Vec-Decision Transformer model, which improves the agent's ability to extract temporal feature correlations and facilitates the delivery of optimal actions. Compared with the original algorithm, the proposed algorithm has better results and performance, not only achieving better convergence results but also faster convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a flowchart of a 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning;

[0035] Figure 2 Action error graphs for the Time2Vec-Decision Transformer model and the Decision Transformer model;

[0036] Figure 3 The power injection fluctuation curve of the power grid without the Time2Vec-Decision Transformer model is shown;

[0037] Figure 4 Inject power fluctuation curve graph for the power grid using the Time2Vec-Decision Transformer model;

[0038] Figure 5 This is a structural diagram of a 5G base station energy storage scheduling device based on mathematical optimization and reinforcement learning. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to illustrate the present invention, rather than to represent all embodiments. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.

[0040] Example 1: Figure 1As shown, the present invention provides a 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning, comprising the following steps:

[0041] (1) Constructing the objective function of the day-ahead optimization model Minimize and get the time Day-ahead scheduling action vector .

[0042] The step (1) specifically includes the following sub-steps:

[0043] (1.1) Based on the prediction vector of user load power at time T , electricity purchase price vector And the electricity price vector Constructing the objective function of the day-ahead optimization model .

[0044] The prediction vector of user load power at the T time for ,in, Indicates time The predicted value of user load power.

[0045] The electricity purchase price vector at the T moments for ,in, Indicates time The electricity purchase price.

[0046] The electricity price vector at the T moments for ,in, Indicates time The electricity purchase price.

[0047] The objective function of the day-ahead optimization model for ,in, represents the first optimization objective function, represents the second optimization objective function, represents the normalization function, represents the first weighted weight, Represents the second weighted weight.

[0048] The first optimization objective function The prediction vector of user load power at time T is constructed, specifically: ,in, Indicates time The grid injection power, represents the average value of the grid injection power at T moments, .

[0049] The moment Grid injection power The calculation formula is ,in, Indicates time Time The predicted value of the charging power of 5G base stations, Indicates time Time The predicted value of the discharge power of a 5G base station, Indicates the number of 5G base stations in the power grid system, .

[0050] The second optimization objective function According to the prediction vector of user load power at time T , electricity purchase price vector and electricity price vector Constructed, specifically .

[0051] (1.2) Objective function based on the charging and discharging power constraints of 5G base station energy storage Minimize and get the scheduling action set of the day-ahead phase at T moments , specifically:

[0052] Based on the constraints, the objective function Minimize and get the predicted value of charging power of each 5G base station at time T and the predicted value of discharge power , and then get the time Time Predicted value of the injection power of a 5G base station in the day-ahead phase : , and obtain the scheduling action set of the day-ahead phase at T moments : ,in, Indicates time The day-ahead scheduling action vector.

[0053] The specific charging and discharging power constraints of the 5G base station energy storage are:

[0054] time Time Predicted value of charging power of 5G base stations Cannot be greater than the upper limit of charging power ;time Time Predicted value of discharge power of 5G base stations Cannot be greater than the upper limit of discharge power ;time Time The power consumption of a 5G base station Cannot be less than the time Time The minimum safe backup power required for a 5G base station And cannot be greater than the time Time The maximum power limit of a 5G base station , that is, satisfy .

[0055] The moment Time The minimum safe backup power required for a 5G base station Through the moment Time Power load power of a 5G base station and minimum backup time The calculation formula is The time Time Power load power of a 5G base station The calculation formula is ,in, Indicates the The fixed load component of each 5G base station, Indicates the The coefficient of variation of the time-varying load of a 5G base station, Indicates time Time The predicted value of the communication load rate of each 5G base station.

[0056] (1.3) Then, the set of scheduling actions from the day-ahead phase at T time points is Select the moment Day-ahead scheduling action vector , the day-ahead scheduling action vector for .

[0057] (2) Based on the real-time information of the power grid environment during the day, obtain the time of the power grid environment The real-time state vector , the grid environment at the moment The real-time state vector Including time ,time The actual value of the grid injection power, electricity purchase price, electricity sales price, user load power and the information collection of all 5G base stations at the time The information set of all 5G base stations includes the time The actual values ​​of the power status and communication load rate of all 5G base stations at that time.

[0058] The moment The real-time state vector for ,in, Indicates time The grid injection power, Indicates time The actual value of the user load power, Represents the information set of all 5G base stations.

[0059] The moment Information collection of all 5G base stations for ,in, Indicates time Time The power status of each 5G base station, Indicates time Time The actual value of the communication load rate of each 5G base station.

[0060] And use the time The real-time state vector Calculate the time Reward value .

[0061] The moment Reward value The calculation formula is:

[0062] ;

[0063] ;

[0064] ;

[0065] in, The historical average value of the power injected into the grid.

[0066] (3) Interception time Before The reward value at the moment, the real-time state vector and the intraday corrected action vector are combined with the moment The real-time state vector and reward value , based on the Time2Vec-Decision Transformer model, we get the time The intraday corrective action vector And with the moment The day-ahead scheduling action vector is added to obtain the time The corrected final motion vector.

[0067] The step (3) specifically includes the following sub-steps:

[0068] (3.1) From the historical sequence of decisions, extract the moment Before The reward value at the moment, the real-time state vector and the intraday correction action vector are then filled with zeros, and the position of the zero filling is recorded with the position mask vector Mask; then and the moment The reward value is combined with the real-time state vector to form The reward value sequence and real-time state sequence at each moment and The intraday correction action sequence of each moment; and construct the A time index sequence of moments.

[0069] The moment Before The reward value at this moment is 、…、 、…、 ,in, Indicates time The reward value of any moment before ; the moment Before The real-time state vector at the moment is 、…、 、…、 , Indicates time The real-time state vector of any previous moment; Before The intraday correction action vector at the moment is 、…、 、…、 , Indicates time The intraday correction action vector at any previous moment.

[0070] described The reward value sequence at each moment is ; The real-time state sequence at each moment is ; The intraday correction action sequence at each moment is ; The time index sequence of each moment is .

[0071] (3.2) The Time2Vec-Decision Transformer model includes a Time2Vec model, a linear layer model, and a Decision Transformer model.

[0072] Then, the time index sequence is encoded by the Time2Vec model to obtain the encoded time index sequence; and the reward value sequence, real-time state sequence and intra-day correction action sequence are encoded by the linear layer model respectively to obtain the encoded reward value sequence, real-time state sequence and intra-day correction action sequence, specifically:

[0073] Time index series At any moment Input into the Time2Vec model to get the time The corresponding encoded time vector The specific process is: first, the time Perform linear transformation using Calculate the first element of the time code vector , and serves as the non-periodic feature of the time-encoding vector; and Indicates that the Time2Vec model calculates the Elements The model weight parameters at the time of . Then the periodic change method is adopted The other elements of the time coding vector are calculated as the periodic characteristics of the time coding vector, where Represents a periodic activation function; finally, the Time2Vec model is calculated dimensional time encoding vector Output. The encoding function of the Time2Vec model is to convert time Encoding becomes a representation of other implicit information dimensional time encoding vector : , extracting non-periodic features and periodic features through the Time2Vec model can help the model mine time The implicit information is contained in the model, which facilitates model decision-making.

[0074] Time index series Repeat the above steps at each moment to obtain the encoded time index sequence .

[0075] Then the reward value sequence Input into the linear layer model to obtain the encoded reward value sequence ,in, Indicates time The encoded reward value vector.

[0076] The real-time status sequence Input into the linear layer model to obtain the corresponding encoded real-time state sequence ,in, Indicates time The encoded real-time state vector.

[0077] Sequence the intraday correction action Input into the linear layer model to obtain the corresponding encoded intraday correction action sequence ,in, Indicates time The encoded intra-day corrected motion vector.

[0078] (3.3) The encoded time index sequence is then added to the encoded reward value sequence, real-time state sequence, and intra-day corrected action sequence respectively to obtain the time-embedded reward value sequence, real-time state sequence, and intra-day corrected action sequence.

[0079] The reward value sequence with time embedding is ,in, Indicates time The reward value vector with time embedding, .

[0080] The real-time state sequence with time embedding is ,in, Indicates time The real-time state vector with time embedding, .

[0081] The intraday correction action sequence with time embedding is ,in, Indicates time The intraday corrected action vector with time embedding, .

[0082] Then, the reward value sequence with time embedding, the real-time state sequence and the intra-day correction action sequence are reordered at the same time, and then spliced ​​to obtain the total sequence after time processing : , each element in the total sequence after time processing is dimensional vector.

[0083] (3.4) Then the total sequence after time processing The position mask vector Mask is input into the Decision Transformer model together with the position mask vector Mask, and the position mask vector Mask is used to shield the position filled by the zero-filling operation so that the zero-filled elements do not affect the decision process of the Decision Transformer model; the Decision Transformer model extracts the total sequence after time processing through the attention mechanism The correlation between the elements in the equation is used to infer the correction value of the injection power of each 5G base station at each moment. After calculation by the Decision Transformer model, the output is Intraday correction action sequence at each moment .

[0084] described Intraday correction action sequence at each moment for ,in, For the moment The intraday correction action vector of Intraday correction action sequence at each moment Get the moment The intraday corrective action vector , the moment The intraday corrective action vector for ,in, Indicates time Time The correction value of the injection power of a 5G base station, that is, the moment The intraday corrective action vector Including time The correction value of the injection power of all 5G base stations at this time.

[0085] (3.5) Change the time Day-ahead scheduling action vector and time The intraday corrective action vector Add up and get the time The final corrected motion vector , the final corrected motion vector for ,in, Indicates time Time The actual value of the injection power of a 5G base station, .

[0086] (4) According to the time The final corrected motion vector Adjust the power grid system to get the time The adjusted real-time state vector .

[0087] The step (4) is specifically as follows: the power grid environment is corrected with the final action vector For dispatching instructions, control the charging and discharging power adjustment of each 5G base station in the power grid, and obtain the real-time status information vector after the power grid system is adjusted The specific process is as follows: the power grid environment is corrected with the final action vector To provide dispatch instructions, control the charging and discharging power adjustment of each 5G base station in the power grid, and then use the corrected final action vector According to the power balance equation , calculate the time Grid injection power Then according to the time The power status of each 5G base station and the corresponding actual value of injection power Physical laws , calculate the time The power status of all 5G base stations 、…、 、…、 Then, from the real-time information of the adjusted power grid environment, the time The actual value of the purchase price, sales price, user load power and time The actual value of the communication load rate of all 5G base stations at that time.

[0088] Then the time ,time Grid injection power , electricity purchase price, electricity sales price, actual value and time of user load power The actual values ​​of the power status and communication load rate of all 5G base stations at this time are used to obtain the real-time status information vector after the power grid system is adjusted. .

[0089] (5) Repeat steps (1) to (4) to obtain the reward values ​​at multiple moments , real-time state vector ,time The final corrected motion vector and the adjusted real-time state vector The decision sequence sample set composed of is added to the sample pool. The decision sequence sample set consists of four tuples composition.

[0090] (6) Randomly select B decision sequence sample sets from the sample pool to train the Time2Vec-DecisionTransformer model to obtain the optimized Time2Vec-DecisionTransformer model, specifically:

[0091] (6.1) Randomly select B decision sequence sample sets from the sample pool. Any decision sequence sample set is represented by a 4-tuple , and then divide the B decision sequence sample sets into groups, the number of decision sequence sample sets in each group is .

[0092] (6.2) Then, the real-time state vectors of all randomly selected decision sequence sample sets are input into the Time2Vec-Decision Transformer model to obtain the corresponding adjusted action set.

[0093] Then, the total loss function is constructed based on the randomly selected B decision sequence sample sets and the adjusted action set. , the total loss function The calculation formula is:

[0094] ;

[0095] ;

[0096] in, Indicates the No. The corrected final action vector in the set of decision sequence samples, Indicates the No. The adjusted action vector corresponding to the decision sequence sample set, , .

[0097] The said No. The adjusted action vector corresponding to the decision sequence sample set is given by No. The real-time state vector in the decision sequence sample set Input into the Time2Vec-Decision Transformer model to obtain.

[0098] And by minimizing the total loss function Train the Time2Vec-Decision Transformer model.

[0099] (6.3) Repeat steps (6.1)-step (6.2) until the total loss function It converges to obtain the optimized Time2Vec-Decision Transformer model.

[0100] (7) The optimized Time2Vec-Decision Transformer model is used for real-time decision-making of the power grid to complete the scheduling of the power grid system, specifically:

[0101] Get the current time Day-ahead scheduling action vector and real-time state vector , and calculate the current time Reward value ; Then intercept the current moment Before The reward value, real-time state vector and intraday correction action vector at the moment are combined with the current moment The real-time state vector and reward value , based on the optimized Time2Vec-DecisionTransformer model, we get the current moment The intraday corrective action vector ; Then the current moment Day-ahead scheduling action vector and the current moment The intraday corrective action vector Add together to get the current moment The final corrected motion vector Finally, the current moment The final corrected motion vector As the current moment The control instructions of the power grid environment are issued to each 5G base station in the power grid environment to control its charging and discharging power, playing a role in coordinated scheduling. It can effectively reduce the variance of the injected power of the power grid, reduce operating costs, and give full play to the "peak shaving and valley filling" advantage of 5G base station energy storage, which is conducive to the safe and economical operation of the power grid.

[0102] In order to verify the effectiveness of the two-stage scheduling method for 5G base station energy storage based on mathematical optimization and improved Decision Transformer proposed in this invention, a test was conducted in a power grid operation scenario. Figure 2 、 Figure 3 and Figure 4 shown.

[0103] Figure 2The action error graphs for the Time2Vec-Decision Transformer model and the Decision Transformer model show the convergence curves before and after the model improvement. Figure 2 The convergence error decrease trend shows that the Time2Vec-Decision Transformer model has a faster decrease trend and convergence effect than the Decision Transformer model alone. The Time2Vec-Decision Transformer model converges better, not only achieving better convergence results, but also faster convergence speed. This shows that the combination of the Time2Vec model and the Decision Transformer model adopted by the present invention has the effect of significantly improving the decision-making ability of the model.

[0104] Figure 3 The power injection fluctuation curve of the power grid without the Time2Vec-Decision Transformer model is shown; Figure 4 The power injection curve of the power grid using the Time2Vec-Decision Transformer model is shown in the figure. Figure 3 and Figure 4 It can be clearly observed that the fluctuation of the power injected into the grid is smaller and the curve is smoother when the Time2Vec-Decision Transformer model proposed in the present invention is applied than when the Time2Vec-Decision Transformer model is not applied. This shows that the Time2Vec-Decision Transformer model proposed in the present invention has a significant ability to smooth out the fluctuation of grid power, which is conducive to the safe and stable operation of the grid.

[0105] In summary, the present invention can comprehensively consider the operating characteristics of 5G energy storage, according to the safety constraints of its operation and the uncertainties it faces, through a two-stage 5G base station energy storage collaborative scheduling framework based on mathematical optimization and reinforcement learning, and use the Time2Vec model structure to enhance the time feature extraction capability of the Decision Transformer algorithm, improve the scheduling decision-making effect, give full play to the characteristics of the flexible resources of base station energy storage, achieve the "low storage and high release" operation effect, and improve the safety and economy of power grid operation.

[0106] Example 2: This embodiment relates to a 5G base station energy storage scheduling device based on mathematical optimization and reinforcement learning, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used for a 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning in the above-mentioned Example 1; the device embodiment can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer.

[0107] like Figure 5 As shown in the figure, at the hardware level, the model includes processors, internal buses, network interfaces, memory, and non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0108] Improvements to a technology can be clearly categorized as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with technological advancements, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can integrate a digital system onto a PLD through their own programming, eliminating the need for a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using "logic compiler" software. This is similar to the software compiler used during program development. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0109] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.

[0110] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0111] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0112] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0113] The present invention may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0114] Example 3: An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning of the above-mentioned Example 1 is implemented.

[0115] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning, characterized in that: The following steps are involved: (1) Construct the objective function of the day-ahead optimization model and minimize it to obtain the time The day-ahead scheduling action vector; (2) Based on the real-time information of the power grid environment during the day, obtain the time of the power grid environment The real-time state vector is calculated and the time is obtained The reward value; (3) Interception time Before The reward value at the moment, the real-time state vector and the intraday corrected action vector are combined with the moment The real-time state vector and reward value of the time are obtained based on the Time2Vec-Decision Transformer model. The intraday correction action vector and the time The day-ahead scheduling action vector is added to obtain the time The corrected final motion vector; (4) According to the time The corrected final action vector is used to adjust the power grid system and obtain the time The adjusted real-time state vector; (5) Repeat steps (1) to (4) to obtain a decision sequence sample set consisting of reward values, real-time state vectors, corrected final action vectors, and adjusted real-time state vectors at multiple moments and add them to the sample pool; (6) Randomly select B decision sequence sample sets from the sample pool to train the Time2Vec-Decision Transformer model to obtain the optimized Time2Vec-Decision Transformer model; (7) The optimized Time2Vec-Decision Transformer model is used for real-time decision-making of the power grid to complete the scheduling of the power grid system.

2. A 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning according to claim 1, characterized in that: The step (1) specifically includes the following sub-steps: (1.1) Construct the objective function of the day-ahead optimization model based on the forecast vector of user load power, the purchase price vector, and the sales price vector at time T; (1.2) Minimizing the objective function based on the charge and discharge power constraints of the 5G base station energy storage system to obtain a set of scheduling actions for the day-ahead phase at T time points. The set of scheduling actions for the day-ahead phase at T time points includes the predicted values ​​of the day-ahead injection power of all 5G base stations at T time points. (1.3) Then, the time is selected from the set of scheduling actions in the day-ahead phase of T time The day-ahead scheduling action vector.

3. A 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning according to claim 2, characterized in that: The objective function of the day-ahead optimization model for ,in, represents the first optimization objective function, represents the second optimization objective function, represents the normalization function, represents the first weighted weight, represents the second weighted weight; The first optimization objective function It is constructed based on the prediction vector of user load power at T moments; The second optimization objective function It is constructed based on the prediction vector of user load power at T moments, the electricity purchase price vector and the electricity sales price vector.

4. A 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning according to claim 1, characterized in that: The grid environment at the moment The real-time state vector contains the time ,time The actual value of the grid injection power, electricity purchase price, electricity sales price, user load power and the information collection of all 5G base stations at the time The information set of all 5G base stations includes the time The actual values ​​of the power status and communication load rate of all 5G base stations at that time.

5. A 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning according to claim 1, characterized in that: The step (3) specifically includes the following sub-steps: (3.1) From the historical sequence of decisions, extract the moment Before The reward value at the moment, the real-time state vector and the intra-day correction action vector are then filled with zeros, and the position of the zero filling is recorded with the position mask vector; then and the moment The reward value is combined with the real-time state vector to form The reward value sequence and real-time state sequence at each moment and The intraday correction action sequence of each moment; and construct the A time index sequence of moments; (3.2) The Time2Vec-Decision Transformer model includes a Time2Vec model, a linear layer model, and a Decision Transformer model; Then the time index sequence is encoded by the Time2Vec model to obtain the encoded time index sequence; The reward value sequence, real-time state sequence and intra-day correction action sequence are respectively encoded by the linear layer model to obtain the encoded reward value sequence, real-time state sequence and intra-day correction action sequence; (3.3) The encoded time index sequence is then added to the encoded reward value sequence, real-time state sequence, and intra-day corrected action sequence respectively to obtain the time-embedded reward value sequence, real-time state sequence, and intra-day corrected action sequence, and then concatenated into the total sequence after time processing; (3.4) Then the total sequence and position mask vector after time processing are input into the Decision Transformer model, and the output is The intraday correction action sequence of the moment and obtain the moment The intraday correction action vector of (3.5) Change the time The day-ahead scheduling action vector and time Add the intraday correction action vectors to get the moment The corrected final motion vector.

6. A 5G base station energy storage scheduling device based on mathematical optimization and reinforcement learning, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning as described in any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that A program is stored thereon, and when the program is executed by the processor, it implements a 5G base station energy storage scheduling method based on mathematical optimization and reinforcement learning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Generative decision model-based power grid operation adjustment method

    CN117154845A

  • Graph neural network-based power grid dispatching decision-making method and large model

    CN119294872A