5G base station group intra-day optimization operation method based on reinforcement learning

By building an intraday optimization operation model of 5G base station group and using deep Q network reinforcement learning algorithm, the problem of high operating costs of 5G base stations is solved, and the optimization operation and cost reduction of base station groups are achieved.

CN120302323APending Publication Date: 2025-07-11TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510716925.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing 5G base station intraday scheduling methods are limited by the accuracy of communication traffic prediction, making it difficult to achieve optimal economic benefits, and traditional optimization algorithms are difficult to meet real-time requirements in large-scale problems, resulting in high operating costs.

Method used

A 5G base station group intraday optimization operation model based on communication load migration is built, and it is converted into a Markov decision-making process model. The agent is trained using deep Q network reinforcement learning algorithm to optimize the base station operation cost through reward functions.

Benefits of technology

The optimized operation of 5G base station clusters has been realized, the operating costs have been reduced, the rapid solution needs for intraday optimization have been met, and the stability and economic benefits of the system have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302323A_ABST
    Figure CN120302323A_ABST
Patent Text Reader

Abstract

The invention provides a 5G base station group intra-day optimization operation method based on reinforcement learning, and relates to the technical field of 5G base station optimization scheduling, and the method is characterized in that the method comprises the steps: constructing a 5G base station group intra-day optimization operation model considering communication load migration, and carrying out the optimization of the 5G base station group intra-day optimization operation model under a scene that the electricity price of each base station is different; the demand response of the power distribution network is participated through a base station communication load migration mechanism, so that the daily electricity purchase cost of a base station operator is minimum; according to the rapid solving requirement of the 5G base station group intra-day optimization scheduling problem, the problem is converted into a Markov decision process model, a deep Q network reinforcement learning algorithm is adopted to train an intelligent agent, and rapid solving of the problem is achieved. Through the method provided by the invention, the optimization of the intra-day operation cost of the 5G base station group can be effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of 5G base station optimization scheduling, and particularly relates to an intra-day optimized operation method for a group of 5G base stations based on reinforcement learning. Background Art

[0002] With the rapid popularization of the new generation of mobile communication technology (5G), the energy consumption problem of 5G base stations has become increasingly prominent. Since the power consumption of a single 5G base station is relatively large while its coverage range is limited, exploring effective energy-saving strategies to reduce its operating cost has become a key issue. 5G base stations have various energy consumption management means, including communication load migration and energy storage scheduling, etc. These characteristics enable them to participate in the demand response plan of the distribution network as adjustable resources, thereby effectively reducing the operating cost. Existing research mainly focuses on the day-ahead scheduling level, but limited by the accuracy of communication traffic prediction, relying solely on the day-ahead scheduling scheme often fails to achieve the optimal economic benefit. Especially when the prediction deviation is large, it may lead to the inability of the base station to meet the communication service demand during actual operation, affecting the system stability. Therefore, it is of great value to carry out research on the optimized operation of 5G base stations at the intra-day time scale.

[0003] In the existing research on the intra-day scheduling of 5G base stations, some literature aims to minimize the operating cost and focuses on the real-time optimization of the energy storage system. However, in addition to energy storage scheduling, 5G base stations can also adjust their energy consumption through the dynamic migration strategy of communication loads, that is, transferring the communication load from the base stations in the areas with higher electricity prices to the base stations in the areas with lower electricity prices. This load migration mechanism based on electricity price differences provides a new idea for reducing the energy consumption cost of base stations and is worthy of further research in the intra-day scenario.

[0004] When studying the intra-day optimization problem, it is necessary to consider the rapid solution of the intra-day operation plan. When solving the intra-day optimization problem, the computational efficiency of the solution algorithm is crucial. Currently, three types of methods are mainly used for the intra-day optimization problems in the field of distribution networks: model predictive control, Lyapunov optimization algorithm, and reinforcement learning algorithm. Among them, model predictive control realizes dynamic decision-making by rolling the time window and combining the latest prediction data; Lyapunov optimization uses the energy function theory to decompose the long-term optimization goal into a series of short-term sub-problems. These two traditional methods both require building an accurate mathematical model and rely on a numerical solver for iterative calculation, and often fail to meet the real-time requirements when facing large-scale optimization problems. In contrast, reinforcement learning, as a data-driven optimization method, does not require complex iterative operations in the decision-making process, and its characteristics are more in line with the strict requirements of the intra-day operation problem for timeliness.

[0005] In summary, based on the deficiencies in the current research field of 5G base station daily scheduling and the characteristics of daily optimization methods, the present invention proposes a method for optimizing the daily operation of a group of 5G base stations based on reinforcement learning. First, a daily operation scenario considering the communication load migration of 5G base stations is constructed, and a daily optimization operation model for a group of 5G base stations is established with the goal of minimizing the operation cost of 5G base station operators. Second, to meet the rapid solution requirement for the daily optimization scheduling problem of a group of 5G base stations, the problem is transformed into a Markov decision process model, and a deep Q-network (DQN) reinforcement learning algorithm is used to train the intelligent agent. Finally, the effectiveness of the proposed method is verified through numerical examples. Summary of the Invention

[0006] The object of the present invention is to provide a method for optimizing the daily operation of a group of 5G base stations based on reinforcement learning, which solves the technical problems of high operation cost and lack of operation optimization of 5G communication base stations, and realizes the technical effect of fully optimizing the operation of 5G base stations and reducing the operation cost.

[0007] To achieve the above object of the invention, the technical solution adopted by the present invention is as follows:

[0008] A method for optimizing the daily operation of a group of 5G base stations based on reinforcement learning, comprising the following steps: Step 1, constructing a daily optimization operation model for a group of 5G base stations based on communication load migration;

[0009] Step 2, transforming the daily optimization operation problem of a group of 5G base stations into a Markov decision process model, and proposing a training process for a reinforcement learning intelligent agent for the daily optimization problem of a group of 5G base stations.

[0010] As an improvement, in the above Step 1, constructing a daily optimization operation model for a group of 5G base stations based on communication load migration includes:

[0011] (1) Objective function of the daily optimization operation model for a group of 5G base stations:

[0012]

[0013] where I is the total number of 5G base stations, c i,t is the electricity price of base station i at time t, is the power consumption of base station i at time t;

[0014] (2) Constraint conditions

[0015] 1. Energy constraint:

[0016]

[0017] where, and respectively represent the static power consumption and dynamic power consumption of base station i at time t, and β represents the energy efficiency coefficient of the base station;

[0018] 2. Constraint on the connection relationship between the base station and the user:

[0019]

[0020] Among them, x i,j,t represents the connection relationship between base station i and user j at time t, 1 represents connection, 0 represents non - connection, and J is the total number of communication load users;

[0021] 3. Constraint on the transmission power of the base station:

[0022]

[0023] Among them, is the transmission power of base station i at time t, is the transmission power corresponding to base station i connecting to user j, N0 is the noise power, A is the fixed path loss value, d i,j is the geographical distance between base station i and user j, B and L j respectively represent the bandwidth of user j and the user traffic demand, g i,j is the channel gain between base station i and user j, d0 is the reference distance, α is the path loss exponent, P i tr,max is the maximum transmission power of base station i;

[0024] 4. Constraint on the bandwidth of the base station:

[0025]

[0026] Among them, B represents the bandwidth demand of the user, and B max represents the total bandwidth that the base station can provide;

[0027] 6. Constraint on the transmission traffic of the base station:

[0028]

[0029] Among them, represents the upper limit of the traffic processing capacity of base station i.

[0030] As an improvement, in step 2, the problem of the day - within optimization operation of the 5G base station group is transformed into a Markov decision process model, specifically including:

[0031] (1) State space: Design the state space from the perspective of communication load users. The transmission power of the base station connecting to communication users is related to the traffic demand of the users and the distance from the users to the base station. The operating cost of the base station is related to the electricity price of the base station and the transmission power. Therefore, to optimize the operating cost of the base station, the environmental state observed by the agent should include the above information, that is, the state space should include the traffic demand of communication users, the distance between communication users and each base station, the electricity price and transmission power of each base station. Therefore, the state of user j can be expressed as:

[0032]

[0033] where S j represents the state space of user j, d BSi is the distance between the i-th base station and user j, c BSi is the electricity price of the i-th base station, is the transmission power of the i-th base station;

[0034] (2) Action space: Similarly, design the action space from the perspective of communication load users. The action space is the index of the base station. The agent selects the base station to which the communication user connects based on the state space. This action matches the decision variable x i,j,t in the intra-day optimal operation problem of the base station. The action space of communication j is:

[0035] A j ={1, 2,..., I}(8)

[0036] where A j is the action space corresponding to user j;

[0037] (3) Reward function: Since the goal of the intra-day optimization problem of 5G base stations is to minimize the total operating cost, the reward function is designed based on the objective function (1). To avoid the bandwidth occupancy and transmission power overlimit of the base station caused by user access, a penalty function is introduced to ensure the transmission power and bandwidth constraints of the base station. The reward function is specifically as follows:

[0038]

[0039] where is the transmission power corresponding to the connection between base station i and user j, P i tr is the transmission power of base station i at time t, p B , p tr are the penalty coefficients corresponding to the transmission power overlimit and bandwidth occupancy overlimit of the base station respectively;

[0040] The training objective of the agent is to maximize the cumulative reward, which is specifically shown in the following formula:

[0041]

[0042] Among them, γ is the discount factor.

[0043] As an improvement, in step 2, the reinforcement learning agent training process for the intra-day optimization problem of 5G base station clusters includes: Step 1, generating communication load distribution data based on the Poisson distribution, generating communication load traffic demand data based on the normal distribution, and the base station electricity price comes from the actual load electricity price. To improve the generalization ability of the trained agent, multiple sets of training data are generated, corresponding to different intra-day scenarios (including peak communication demand periods / low-demand periods / transition periods);

[0044] Step 2, initializing the training network Q(s,a;θ) and the target network Q(s,a;θ - ), where θ and θ - are the parameters of the training network and the target network respectively. Initialize the 5G base station status and the base station-communication load connection relationship, and initialize the experience replay buffer;

[0045] Step 3, in each action, the communication load selects an action a (connect to a certain base station) according to the ε-greedy policy and executes it, obtains the reward R and the next state s′ (the real-time electricity price of each base station, the transmission power of each base station, the traffic demand of the communication load, and the distance to each base station), and stores the experience (s,a,R,s′) in the experience replay buffer;

[0046] Step 4, randomly sample an experience sample (s,a,R,s′) from the experience replay buffer, calculate the target Q value y = R + γmaxQ(s′,a′;θ - ), use the mean square error as the loss function, and use the gradient descent method to update the training network parameter θ;

[0047] Step 5, update the target network, and regularly copy the parameter θ of the training network to the target network parameter θ - ;

[0048] Step 6, continuously repeat the above steps until the training network and the target network converge to obtain the action value function Q(s,a).

[0049] The beneficial effects of the present invention are as follows: By establishing an optimization operation model and cooperating with the Markov decision process model, the operation rules of the base station can be fully reflected and obtained, and an optimization guidance can be given to it. Cooperating with multiple constraint conditions can further plan the optimization direction, thereby improving the overall optimization effect and achieving the effect of fully reducing the base station operation cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is the flowchart of the implementation of an intra-day optimization operation method for 5G base station clusters based on reinforcement learning according to the present invention;

[0051] Figure 2 Schematic diagram of the neural network constructed by the DQN reinforcement learning algorithm for the intra-day optimized operation method of a 5G base station group based on reinforcement learning according to the present invention;

[0052] Figure 3 Schematic diagram of the intra-day operation scenario of a 5G base station group in the embodiment of the intra-day optimized operation method of a 5G base station group based on reinforcement learning according to the present invention;

[0053] Figure 4 Schematic diagram of the electricity price curves of different types of loads in the embodiment of the intra-day optimized operation method of a 5G base station group based on reinforcement learning according to the present invention;

[0054] Figure 5 Schematic diagram of the convergence curve of the training results of the agent in the embodiment of the intra-day optimized operation method of a 5G base station group based on reinforcement learning according to the present invention. Detailed implementation manners

[0055] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0056] It should be noted that the terms "first" and "second" in the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0057] As Figure 1 and Figure 2 shown, the present invention proposes an intra-day optimized operation method for a 5G base station group based on reinforcement learning, and the method includes the following steps:

[0058] Step 1: Construct an intra-day optimized operation model for a 5G base station group based on communication load migration;

[0059] Step 2: Convert the intra-day optimization operation problem of the 5G base station group into a Markov decision process model, and propose a reinforcement learning agent training process for the intra-day optimization problem of the 5G base station group;

[0060] The above-mentioned Step 1 constructs an intra-day optimization operation model of the 5G base station group based on communication load migration, including:

[0061] (1) Objective function

[0062]

[0063] where I is the total number of 5G base stations, c i,t is the electricity price of base station i at time t, is the power consumption of base station i at time t.

[0064] (2) Constraint conditions

[0065] 1) Energy constraint:

[0066]

[0067] where, and respectively represent the static power consumption and dynamic power consumption of base station i at time t, and β represents the energy efficiency coefficient of the base station.

[0068] 2) Base station - user connection relationship constraint:

[0069]

[0070] where x i,j,t represents the connection relationship between base station i and user j at time t, 1 represents connection, 0 represents non - connection, and J is the total number of communication load users.

[0071] 3) Base station transmission power constraint:

[0072]

[0073] where, is the transmission power of base station i at time t, is the transmission power corresponding to base station i connecting to user j, N0 is the noise power, A is the fixed path loss value, d i,j is the geographical distance between base station i and user j, B and L j are the bandwidth and traffic demand of user j respectively, g i,j is the channel gain between base station i and user j, d0 is the reference distance, α is the path loss exponent, P i tr,max is the maximum transmission power of base station i.

[0074] 4) Base station bandwidth constraint:

[0075]

[0076] Among them, B represents the bandwidth requirement of the user, and B max represents the total bandwidth that the base station can provide.

[0077] 5) Base station transmission traffic constraint:

[0078]

[0079] Among them, represents the upper limit of the traffic processing volume of base station i.

[0080] The above step 2 transforms the problem of the daily optimized operation of the 5G base station group into a Markov decision process model, specifically including:

[0081] The Markov decision process model is usually represented by the tuple G = <S, A, P, r, γ>. Among them, S represents the state space of the agent, A represents the action space of the agent, P is the state transition function, r represents the reward function, and γ is the discount factor.

[0082] In an ideal situation, the agent can consider the states of all users and then output a decision plan for all users in one action. However, the communication load migration problem is a combinatorial optimization problem, and the total number of combinations is I J (I is the number of base stations, and J is the number of communication loads), which will cause the agent to face a huge action space and state space. Therefore, it is difficult to train an agent that can simultaneously decide on all communication user load migration plans. To effectively reduce the scale of the action space and state space, during the training process, let the agent make allocation decisions for each user one by one, and then design the state and action spaces from the perspective of communication users.

[0083] Based on the above analysis, the Markov decision process model for the 5G base station optimization problem is specifically modeled as follows.

[0084] State space: Design the state space from the perspective of communication load users. The transmission power required for the base station to connect communication users is related to the traffic demand of the users and the distance from the users to the base station. The operating cost of the base station is related to the electricity price of the base station and the transmission power. Therefore, to optimize the operating cost of the base station, the environmental state observed by the agent should include the above information, that is, the state space should include the traffic demand of communication users, the distance between communication users and each base station, the electricity price and transmission power of each base station. Therefore, the state of user j can be expressed as:

[0085]

[0086] Among them, Sj Denote the state space of user j as d BSi The distance between the i-th base station and user j is c BSi The electricity price of the i-th base station is The transmission power of the i-th base station is

[0087] Action space: The action space is the index of the base stations. The agent selects the base station for the communication user to access based on the state space. This action matches the decision variable x in the problem of the day-ahead optimal operation of the base stations. i,j,t The action space for communicating with j is

[0088] A j ={1, 2, ..., I}(8)

[0089] where A j is the action space corresponding to user j.

[0090] Reward function: Since the goal of the day-ahead optimization problem of 5G base stations is to minimize the total operating cost, the reward function is designed based on the objective function (1). To avoid the bandwidth occupation and transmission power overlimit of the base stations caused by user access, a penalty function is introduced to ensure the transmission power and bandwidth constraints of the base stations. The reward function is as follows:

[0091]

[0092] where is the transmission power corresponding to the connection between base station i and user j, P i tr is the transmission power of base station i at time t, p B , p tr are the penalty coefficients corresponding to the overlimit of the transmission power and bandwidth occupation of the base station respectively.

[0093] The training objective of the agent is to maximize the cumulative reward, which is shown as the following formula.

[0094]

[0095] where γ is the discount factor.

[0096] Step 2 proposes a reinforcement learning agent training process for the day-ahead optimization problem of 5G base station clusters, which specifically includes:

[0097] The Deep Q Network (DQN) is a model-free reinforcement learning algorithm and belongs to the value iteration method. The goal of this method is to use a neural network to approximate the action value function Q(s, a), enabling the agent to select the optimal action in a given state. In the DQN reinforcement learning algorithm, Q(s, a) is updated iteratively to approximate the optimal action value function, and its iterative update formula is as follows.

[0098] Q(s,a) ← Q(s,a) + α[R + γ max Q(s′,a′) - Q(s,a)] (11)

[0099] Among them, Q(s,a) represents the action value function, s represents the current state, a represents the action taken in the current state, α represents the learning rate, R is the immediate reward, s′ is the next state, and a′ is the action taken in the next state.

[0100] Based on using a deep neural network to approximate the action value function, the DQN algorithm also introduces two mechanisms, experience replay and target network, to improve the learning efficiency. The specific method principles are as follows.

[0101] Experience replay: The agent stores the experience (state, action, reward, next state) of each interaction with the environment in an experience pool. During training, a batch of data is randomly sampled from the experience pool for training instead of directly using the latest interaction data. This mechanism breaks the correlation between data and improves the sample utilization rate.

[0102] Target network: The DQN reinforcement learning algorithm contains two neural networks. One is the training network (used to select actions and update the action value function), and the other is the target network (used to estimate the action value function). Among them, the introduction of the target network helps to stabilize the training process and avoid large fluctuations in the estimated value of the action value function.

[0103] In summary, for the problem of the intraday optimal operation of a 5G base station group, the process of training an agent based on the DQN algorithm is as follows.

[0104] Step 1: Generate communication load distribution data based on the Poisson distribution, generate communication load traffic demand data based on the normal distribution, and the base station electricity price comes from the actual load electricity price. To improve the generalization ability of the trained agent, multiple sets of training data are generated, corresponding to different intraday scenarios (including peak / valley / transition periods of communication demand).

[0105] Step 2: Initialize the training network Q(s,a;θ) and the target network Q(s,a;θ - ), where θ and θ - are the parameters of the training network and the target network respectively. Initialize the 5G base station state and the base station - communication load connection relationship, and initialize the experience replay buffer.

[0106] Step 3: In each action, the communication load selects an action a (connect to a certain base station) according to the ε-greedy policy and executes it, obtaining a reward R and the next state s′ (the real-time electricity price of each base station, the transmit power of each base station, the traffic demand of the communication load, and the distance to each base station), and storing the experience (s, a, R, s′) in the experience replay buffer.

[0107] Step 4: Randomly sample an experience sample (s, a, R, s′) from the experience replay buffer, calculate the target Q value y = R + γmaxQ(s′, a′; θ - ), use the mean squared error as the loss function, and update the training network parameter θ using the gradient descent method.

[0108] Step 5: Update the target network, and regularly copy the parameter θ of the training network to the target network parameter θ - .

[0109] Step 6: Continuously repeat the above steps until the training network and the target network converge, obtaining the action value function Q(s, a).

[0110] As Figures 3 to 5 shown, this embodiment takes 24 hours as the overall operating scenario within a day. The scenario settings are as follows: In an 8km×8km area, 100 5G base stations are evenly distributed, with a distance of 800m between each base station, and the coverage radius of the base station is 600m, specifically as Figure 3 shown. The example area is divided into 2000 small areas, and any small area represents a communication user aggregation point. The entire area consists of an industrial load area, a commercial load area, and a residential load area, and the time-of-use electricity prices in different areas are as Figure 4 shown. The communication parameter settings are as follows: The static power of the base station is 2.3kW, and the maximum transmit power is 1kW; the fixed path loss value A = -35dB; the maximum transmission bandwidth of the base station is 100MHz; the communication user bandwidth demand W = 2MHz; the energy efficiency coefficient β = 2.81. The traffic demand L (Mbit / s) of the user is randomly taken within the range of [50, 100]. MATLAB R2022a is used for example verification.

[0111] (1) Training convergence

[0112] The relevant parameter settings of the DQN reinforcement learning agent are as follows: The Q network consists of an input layer, a hidden layer, and an output layer. The hidden layer has two layers, each layer contains 64 neurons, and the activation function is ReLu. The discount factor is 0.99, the learning rate is 0.001, the capacity of the experience replay pool is 10000, and the minimum sampling batch is 32.

[0113] It can be seen from the model training convergence curve that as the number of training increases, the rewards obtained by the agent continue to grow, and the reward returns tend to be stable after 1200 times of training. The final average reward return is about -5091.

[0114] (2) Optimization results

[0115] Set the following two scenarios for example comparison.

[0116] Scenario 1: Do not consider communication load migration.

[0117] Scenario 2: Consider communication load migration participating in demand response

[0118]

[0119] According to the above results, the total operating cost of Scenario 2 is reduced by 412.43 yuan compared with Scenario 1, indicating that considering communication load migration participating in demand response can reduce the operating cost of the base station group.

[0120] The above are only the preferred embodiments of the present invention patent, and are not used to limit the present invention patent. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention patent shall be included in the protection scope of the present invention patent.

Claims

1. A method for intra-day optimized operation of a 5G base station group based on reinforcement learning, characterized in that, It includes the following steps: Step 1: Construct an intraday optimization operation model for a 5G base station group based on communication load migration; Step 2: Transform the intraday optimization operation problem of the 5G base station group into a Markov decision process model, and propose a reinforcement learning agent training process for the intraday optimization problem of the 5G base station group.

2. The method for optimizing the daily operation of a 5G base station group based on reinforcement learning according to claim 1, wherein In the above Step 1, constructing an intraday optimization operation model for a 5G base station group based on communication load migration includes: (1) The objective function of the intraday optimization operation model for the 5G base station group: Among them, I is the total number of 5G base stations, and c i,t is the electricity price of base station i at time t, and is the power consumption of base station i at time t; (2) Constraint conditions 1. Energy constraint: Among them, and respectively represent the static power consumption and dynamic power consumption of base station i at time t, and β represents the energy efficiency coefficient of the base station; 2. Constraint on the connection relationship between the base station and the user: where x i,j,t represents the connection relationship between base station i and user j at time t, 1 means connected, 0 means not connected, and J is the total number of communication load users; 3. Constraint on the transmit power of the base station: Among them, is the transmission power of base station i at time t, is the transmission power corresponding to the user j connected to base station i, N0 is the noise power, A is the fixed path loss value, d i,j is the geographical distance between base station i and user j, B and L j are the bandwidth and user traffic demand of user j respectively, g i,j is the channel gain between base station i and user j, d0 is the reference distance, α is the path loss exponent, is the maximum transmission power of base station i; 4. Constraint on the bandwidth of the base station: Among them, B represents the bandwidth requirement of the user, and B max represents the total bandwidth that the base station can provide; 5. Constraint on the transmission traffic of the base station: Among them, represents the upper limit of the traffic processing capacity of base station i.

3. The method for optimizing the intra-day operation of a 5G base station group based on reinforcement learning according to claim 1, characterized in that In the above Step 2, transforming the intraday optimization operation problem of the 5G base station group into a Markov decision process model specifically includes: (1) State space: Design the state space from the perspective of communication load users. The transmit power of the base station connecting to communication users is related to the traffic demand of the users and the distance from the users to the base station. The operating cost of the base station is related to the electricity price of the base station and the transmit power. Therefore, to optimize the operating cost of the base station, the environmental state observed by the agent should include the above information, that is, the state space should include the traffic demand of communication users, the distance between communication users and each base station, the electricity price and transmit power of each base station. Therefore, the state of user j can be expressed as: Among them, S j represents the state space of user j, d BSi is the distance between the i-th base station and user j, c BSi is the electricity price of the i-th base station, is the transmission power of the i-th base station; (2) Action space: The action space is also designed from the perspective of communication load users. The action space is the index of the base station. The agent selects the base station for communication user access based on the state space. This action matches the decision variable x in the problem of optimizing the daily operation of the base station. The action space for communication j is as follows: i,j,t matches, and the action space for communication j is: A j = {1, 2,..., I}(8) Among them, A j is the action space corresponding to user j; (3) Reward function: Since the goal of the intraday optimization problem of the 5G base station is to minimize the total operating cost, a reward function is designed based on the objective function (1). To avoid the bandwidth occupancy and transmit power over-limit of the base station caused by user access, a penalty function is introduced to ensure the transmit power and bandwidth constraints of the base station. The reward function is specifically as follows: Among them, is the transmission power corresponding to base station i connecting to user j, is the transmission power of base station i at time t, p B and p tr are the penalty coefficients corresponding to the transmission power limit and bandwidth occupancy limit of the base station respectively; The training goal of the agent is to maximize the cumulative reward, specifically as shown in the following formula: where γ is the discount factor.

4. The method for optimizing the intra-day operation of a 5G base station group based on reinforcement learning according to claim 1, wherein In the above Step 2, the proposed reinforcement learning agent training process for the intraday optimization problem of the 5G base station group includes: Step 1: Generate communication load distribution data based on the Poisson distribution, generate communication load traffic demand data based on the normal distribution, and the electricity price of the base station is from the actual load electricity price. To improve the generalization ability of the trained agent, multiple groups of training data are generated, corresponding to different intraday scenarios (including peak / valley / transition periods of communication demand); Step 2: Initialize the training network Q(s,a;θ) and the target network Q(s,a;θ - ), where θ and θ - are the parameters of the training network and the target network respectively, initialize the 5G base station status and the base station-communication load connection relationship, and initialize the experience replay buffer; Step 3: In each action, the communication load selects an action a (connect to a certain base station) according to the ε-greedy strategy and executes it, obtains the reward R and the next state s′ (the real-time electricity price of each base station, the transmit power of each base station, the traffic demand of the communication load and the distance to each base station), and stores the experience (s, a, R, s′) in the experience replay buffer; Step 4: Randomly sample an experience sample (s, a, R, s′) from the experience replay buffer, calculate the target Q-value y = R + γmaxQ(s′, a′; θ - ), use the mean squared error as the loss function, and update the training network parameter θ using the gradient descent method; Step 5, update the target network, and regularly copy the parameters θ of the training network to the target network parameters θ - ; Step 6: Continuously repeat the above steps until the training network and the target network converge to obtain the action value function Q(s, a).