Base station control method and device

By dividing the base station network area into discrete grids and using the action value network model to predict the reward value, the shortcomings of base station energy consumption management in the existing technology are solved, and intelligent control and energy consumption optimization of base stations and cells are realized.

CN120111633APending Publication Date: 2025-06-06TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510187784.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems in the energy consumption management of base stations with limited energy saving effects, relying on manual analysis, difficulty in dealing with large changes in traffic loads, and difficulty in cooperating with each other between base stations.

Method used

By dividing the network area covered by the base station into discrete grids, the cells in each grid are equivalent in pairs, and using the pre-trained action value network model, the reward values ​​of different working states are predicted based on the cell's characteristic parameters and predicted traffic, thereby controlling the working state of the base station and the cell.

Benefits of technology

Intelligent control of base stations and cells is realized, which effectively reduces energy consumption, improves network resource utilization, reduces idle and waste of resources, and reduces the need for human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111633A_ABST
    Figure CN120111633A_ABST
Patent Text Reader

Abstract

The invention provides a base station control method and device, which can be used in the technical field of wireless communication facility control. The method comprises the following steps: dividing a network area covered by a base station into at least two discrete grids according to a co-coverage relationship between affiliated cells of the base station; for each cell, generating an input feature vector according to the feature parameter of the cell in the target time period, the feature parameters of other cells in the grid where the cell is located in the target time period and the feature parameters of other cells which are outside the grid where the cell belongs and belong to the base station where the cell belongs in the target time period; according to the input feature vector and the predicted flow of each grid at a target time after the target time period, utilizing a pre-trained action value network model to predict reward values of the cell in different working states; and controlling the work of the base station and each affiliated cell of the base station at the target time according to the reward values of different working states of each cell predicted by the action value network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of wireless communication facility control, and in particular to a base station control method and device. Background Art

[0002] With the rapid development of communication networks, the energy consumption of base stations has attracted widespread attention.

[0003] Early research work proposed many simple and direct heuristic methods, such as shutting down the central cell according to the coverage relationship between cells, pre-defining some thresholds or triggers, adjusting the control state of the base station according to the traffic load, keeping the minimum number of cells open according to the greedy method, etc. These methods are simple and easy to implement, but the energy saving effect is limited, and they rely on manual analysis of historical traffic data and cannot cope with large changes in traffic load.

[0004] With the development of deep RL, some RL-based energy-saving methods have been proposed. Most of these methods use Q-learning or other RL algorithms to train a single agent to control all working modes of the base station, or use global rewards to train multiple independent agents to control the mode. When faced with large-scale problems, these methods find it difficult to determine which changes in control states are the cause of the overall performance improvement or reduction, making it difficult to optimize the strategy.

[0005] In summary, the existing related work has the following limitations: (1) Most methods use heuristic or meta-heuristic methods for optimization. The control algorithm is relatively simple and cannot be adjusted in detail according to the system status, which limits the energy-saving efficiency; (2) Most of them are only based on the status of a single base station, and it is difficult for base stations to cooperate with each other, which further limits the energy-saving efficiency. Summary of the invention

[0006] In view of the problems in the prior art, the embodiments of the present application provide a base station control method and device, which can at least partially solve the problems in the prior art.

[0007] On the one hand, the present application proposes a base station control method, comprising:

[0008] Divide the network area covered by the base station into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, and the number of the base station is at least one;

[0009] For each of the cells, an input feature vector is generated according to feature parameters of the cell in the target period, feature parameters of other cells in the grid where the cell is located in the target period, and feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target period;

[0010] According to the input feature vector and the predicted flow of each grid at a target time after the target period, a pre-trained action value network model is used to predict the reward value of the cell in different working states, wherein the reward value is related to the working state of the cell, the energy consumption value of the base station to which the cell belongs, and the energy consumption value of the cells in the grid where the cell is located;

[0011] According to the reward value of each of the cells in different working states predicted by the action value network model, the base station and each of the cells attached to the base station are controlled to work at the target time.

[0012] In some embodiments, the network area covered by the base station is divided into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, including:

[0013] Selecting a first cell from the network area covered by the base station to create a first grid, so that the network coverage of the first grid is equal to the network coverage of the first cell;

[0014] Determine other cells equivalent to the first cell according to the co-coverage relationship between the first cell and other cells other than the first cell in the affiliated cells of the base station;

[0015] Adding other cells equivalent to the first cell to the first grid;

[0016] Continue to select a second cell from other subordinate cells of the base station outside the first grid to create a second grid, so that the coverage range of the second grid is equal to the coverage range of the second cell;

[0017] Determine other cells equivalent to the second cell according to the co-coverage relationship between other subordinate cells of the base station outside the first grid and the second cell;

[0018] Adding other cells equivalent to the second cell to the second grid;

[0019] This process is deduced in this way until all cells under the base station are divided into grids.

[0020] In some embodiments, for each of the cells, an input feature vector is generated according to feature parameters of the cell in the target period, feature parameters of other cells in the grid where the cell is located in the target period, and feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target period, including:

[0021] For each of the cells, generating a first input feature vector of the cell according to the feature parameters of the cell in the target time period;

[0022] Generate a second input feature vector for each other cell in the grid where the cell is located according to the feature parameters of each other cell in the target time period;

[0023] A third input feature vector of each of the other cells is generated according to feature parameters of other cells that are outside the grid to which the cell belongs and belong to the base station to which the cell belongs within the target time period.

[0024] In some embodiments, the characteristic parameters of each of the cells in the target time period include at least one of the following: current time, traffic load of the grid to which the cell belongs within the last m time steps, and equipment parameters of the cell, where m is a positive integer.

[0025] In some embodiments, based on the input feature vector and the predicted flow of each grid at a target time after the target period, a pre-trained action value network model is used to predict the reward value of the cell in different working states, including:

[0026] Integrate the first input feature vector and each of the second input feature vectors using a first attention layer to obtain a first integrated embedding vector;

[0027] Integrate the first input feature vector and each of the third input feature vectors using a second attention layer to obtain a second integrated embedding vector;

[0028] Concatenate the first integrated embedding vector, the second integrated embedding vector, and the global feature vector into a single vector;

[0029] The single vector is processed by using a fully connected layer to obtain reward values ​​of different working states of the cell.

[0030] In some embodiments, the global feature vector is generated according to the time step and the flow demand of each of the grids at the target time.

[0031] In some embodiments, before predicting the reward value of the cell in different working states using a pre-trained action value network model based on the input feature vector and the predicted traffic of each grid at a target time after the target time period, the method further includes:

[0032] Acquire each sample data in the sample set, wherein each sample data includes characteristic parameters of an affiliated cell of the base station in a historical period, characteristic parameters of other cells in the grid where the cell is located in the historical period, characteristic parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, a reward value of the cell at a historical time after the historical period, a time step, a traffic demand of each of the grids at the historical time, a first mask, and a second mask;

[0033] Generate a historical input feature vector according to the feature parameters of the affiliated cell of the base station in each sample data in a historical period, the feature parameters of other cells in the grid where the cell is located in the historical period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period; generate a historical global feature vector according to the time step in the sample data and the traffic demand of each grid at the historical time; and generate a mask vector according to the first mask and the second mask in the sample data;

[0034] Taking the historical input feature vector, the historical global feature vector and the mask vector corresponding to each sample data as input, and taking the reward value in the sample data as the target output, the first action value network model is trained until the output of the first action value network model meets the first preset requirement;

[0035] Taking the historical input feature vector and the historical global feature vector corresponding to each sample data as input, and the output of the first action value network model under the same sample data as the target output, the second action value network model is trained until the output of the second action value network model meets the second preset requirements, and the second action value network model whose output meets the second preset requirements is used as the trained action value network model.

[0036] In some embodiments, for a sample data, the first mask in the sample data indicates whether the traffic demand of the grid where the cell is located can be met if the subsidiary cell in the sample data is set to a sleep state; the second mask in the sample data indicates whether the base station to which the cell belongs can be shut down if the subsidiary cell is set to a sleep state.

[0037] In some embodiments, the reward value of each cell is negatively correlated to the sum of the energy consumption of the remote radio units of cells belonging to the same grid as the cell, and the energy consumption of the baseband unit and the cooling equipment of the base station to which the cell belongs.

[0038] In some embodiments, controlling the base station and each of the attached cells of the base station to operate at the target time according to the reward value of each of the cells in different operating states predicted by the action value network model includes:

[0039] For each subordinate cell of the base station, determining a preliminary working state of the cell according to reward values ​​of different working states of the cell predicted by the action value network model, wherein the working state of each cell includes an activated state, a deactivated state and a dormant state;

[0040] For each grid under the base station, if the total capacity of the cells in the grid whose initial working state is activated is less than the predicted total traffic demand, the cells whose initial working state is closed are modified to activated in descending order of cell capacity until the total capacity of the cells whose working state is activated is greater than the total traffic demand;

[0041] The predicted total traffic demand in each grid is divided according to the capacity of the activated cells and then allocated to the corresponding cells;

[0042] If all cells in a base station correspond to the closed state, the base station and each subordinate cell of the base station are controlled to be closed; otherwise, the base station is controlled to be turned on, and the cell is controlled to work according to the corresponding working state and allocated traffic of each subordinate cell of the base station.

[0043] On the other hand, the present application proposes a base station control device, comprising:

[0044] A division module, configured to divide the network area covered by the base station into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, and the number of the base station is at least one;

[0045] A first generating module is used to generate an input feature vector for each cell according to the feature parameters of the cell in the target time period, the feature parameters of other cells in the grid where the cell is located in the target time period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target time period;

[0046] A prediction module, used to predict the reward value of the cell in different working states by using a pre-trained action value network model according to the input feature vector and the predicted flow of each grid at a target time after the target period, wherein the reward value is related to the influence of the working state of the cell, the energy consumption value of the base station to which the cell belongs, and the energy consumption value of the cells in the grid where the cell is located;

[0047] The control module is used to control the operation of the base station and each subordinate cell of the base station at the target time according to the reward value of each cell in different working states predicted by the action value network model.

[0048] In some embodiments, the partitioning module is specifically used to:

[0049] Selecting a first cell from the network area covered by the base station to create a first grid, so that the network coverage of the first grid is equal to the network coverage of the first cell;

[0050] Determine other cells equivalent to the first cell according to the co-coverage relationship between the first cell and other cells other than the first cell in the affiliated cells of the base station;

[0051] Adding other cells equivalent to the first cell to the first grid;

[0052] Continue to select a second cell from other subordinate cells of the base station outside the first grid to create a second grid, so that the coverage range of the second grid is equal to the coverage range of the second cell;

[0053] Determine other cells equivalent to the second cell according to the co-coverage relationship between other subordinate cells of the base station outside the first grid and the second cell;

[0054] Adding other cells equivalent to the second cell to the second grid;

[0055] This process is deduced in this way until all cells under the base station are divided into grids.

[0056] In some embodiments, the first generating module is specifically used for:

[0057] For each of the cells, generating a first input feature vector of the cell according to the feature parameters of the cell in the target time period;

[0058] Generate a second input feature vector for each other cell in the grid where the cell is located according to the feature parameters of each other cell in the target time period;

[0059] A third input feature vector of each of the other cells is generated according to feature parameters of other cells that are outside the grid to which the cell belongs and belong to the base station to which the cell belongs within the target time period.

[0060] In some embodiments, the characteristic parameters of each of the cells in the target time period include at least one of the following: current time, traffic load of the grid to which the cell belongs within the last m time steps, and equipment parameters of the cell, where m is a positive integer.

[0061] In some embodiments, the prediction module is specifically used to:

[0062] Integrate the first input feature vector and each of the second input feature vectors using a first attention layer to obtain a first integrated embedding vector;

[0063] Integrate the first input feature vector and each of the third input feature vectors using a second attention layer to obtain a second integrated embedding vector;

[0064] Concatenate the first integrated embedding vector, the second integrated embedding vector, and the global feature vector into a single vector;

[0065] The single vector is processed by using a fully connected layer to obtain reward values ​​of different working states of the cell.

[0066] In some embodiments, the global feature vector is generated according to the time step and the flow demand of each of the grids at the target time.

[0067] In some embodiments, the apparatus further comprises:

[0068] An acquisition module, used to acquire each sample data in the sample set, wherein each sample data includes characteristic parameters of an affiliated cell of the base station in a historical period, characteristic parameters of other cells in the grid where the cell is located in the historical period, characteristic parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, a reward value of the cell at a historical time after the historical period, a time step, a flow demand of each of the grids at the historical time, a first mask, and a second mask;

[0069] A second generating module is used to generate a historical input feature vector according to the feature parameters of the affiliated cell of the base station in each sample data in a historical period, the feature parameters of other cells in the grid where the cell is located in the historical period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, generate a historical global feature vector according to the time step in the sample data and the traffic demand of each of the grids at the historical time, and generate a mask vector according to the first mask and the second mask in the sample data;

[0070] A first training module is used to take the historical input feature vector, the historical global feature vector and the mask vector corresponding to each sample data as input, and the reward value in the sample data as the target output, to train the first action value network model until the output of the first action value network model meets the first preset requirement;

[0071] The second training module is used to take the historical input feature vector and the historical global feature vector corresponding to each sample data as input, and the output of the first action value network model under the same sample data as the target output, to train the second action value network model until the output of the second action value network model meets the second preset requirements, and use the second action value network model whose output meets the second preset requirements as the trained action value network model.

[0072] In some embodiments, for a sample data, the first mask in the sample data indicates whether the traffic demand of the grid where the cell is located can be met if the subsidiary cell in the sample data is set to a sleep state; the second mask in the sample data indicates whether the base station to which the cell belongs can be shut down if the subsidiary cell is set to a sleep state.

[0073] In some embodiments, the reward value of each cell is negatively correlated to the sum of the energy consumption of the remote radio units of cells belonging to the same grid as the cell, and the energy consumption of the baseband unit and the cooling equipment of the base station to which the cell belongs.

[0074] In some embodiments, the control module is specifically used to:

[0075] For each subordinate cell of the base station, determining a preliminary working state of the cell according to reward values ​​of different working states of the cell predicted by the action value network model, wherein the working state of each cell includes an activated state, a deactivated state and a dormant state;

[0076] For each grid under the base station, if the total capacity of the cells in the grid whose initial working state is activated is less than the predicted total traffic demand, the cells whose initial working state is closed are modified to activated in descending order of cell capacity until the total capacity of the cells whose working state is activated is greater than the total traffic demand;

[0077] The predicted total traffic demand in each grid is divided according to the capacity of the activated cells and then allocated to the corresponding cells;

[0078] If all cells in a base station correspond to the closed state, the base station and each subordinate cell of the base station are controlled to be closed; otherwise, the base station is controlled to be turned on, and the cell is controlled to work according to the corresponding working state and allocated traffic of each subordinate cell of the base station.

[0079] An embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the base station control method described in any of the above embodiments are implemented.

[0080] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the base station control method described in any of the above embodiments are implemented.

[0081] An embodiment of the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided by the above-mentioned method embodiments.

[0082] The base station control method and device provided in the embodiment of the present application divides the network area of ​​the base station into discrete grids, extracts characteristic parameters, and predicts the reward value. By predicting the flow and reward value to select the optimal working state, the intelligent control of the base station and the cell can be realized, which can effectively reduce the energy consumption of the base station and each cell, and can also improve the utilization rate of network resources and reduce idle waste of resources. The reward value is predicted by the pre-trained action value network model, which can realize the automatic decision-making of the working state of the base station and the cell. This automated decision-making process can reduce the need for human intervention and improve the intelligence level of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0084] Figure 1 It is a flowchart of a base station control method provided in one embodiment of the present application.

[0085] Figure 2 It is a partial flow chart of a base station control method provided in one embodiment of the present application.

[0086] Figure 3 It is a partial flow chart of a base station control method provided in one embodiment of the present application.

[0087] Figure 4 It is a partial flow chart of a base station control method provided in one embodiment of the present application.

[0088] Figure 5 It is a partial flow chart of a base station control method provided in one embodiment of the present application.

[0089] Figure 6 It is a structural diagram of an action value network provided in one embodiment of the present application.

[0090] Figure 7 It is a partial flow chart of a base station control method provided in one embodiment of the present application.

[0091] Figure 8 It is a schematic diagram of the optimization process of the policy network provided in one embodiment of the present application.

[0092] Fig. 9 It is a structural diagram of a base station control device provided in one embodiment of the present application.

[0093] Fig.10 It is a schematic diagram of the physical structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0094] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. Here, the illustrative embodiments of the present application and their descriptions are used to explain the present application, but are not intended to limit the present application. It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments can be arbitrarily arranged with each other.

[0095] The terms “first”, “second”, etc. used in this document do not specifically refer to an order or sequence, nor are they used to limit this application. They are only used to distinguish elements or operations described with the same technical terms.

[0096] The words “include,” “including,” “have,” “contain,” etc. used in this document are open-ended terms, meaning including but not limited to.

[0097] As used herein, "and / or" includes any and all permutations of the items described.

[0098] The execution subject of the base station control method provided in the present application includes but is not limited to a computer.

[0099] Figure 1 is a flow chart of a base station control method provided by an embodiment of the present application, such as Figure 1 As shown, the base station control method provided in the embodiment of the present application includes:

[0100] S101. Divide the network area covered by the base station into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, and the number of the base station is at least one;

[0101] In step S101, the base station is first processed, and the network area covered by the base station is divided into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station. The purpose of this is to make the cells in each grid have equivalent characteristics, so as to facilitate subsequent optimization and control.

[0102] S102, for each of the cells, generating an input feature vector according to feature parameters of the cell in the target period, feature parameters of other cells in the grid where the cell is located in the target period, and feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target period;

[0103] In step S102, the state of each cell is described and represented according to the characteristic parameters of each cell, providing input for subsequent prediction and decision-making.

[0104] S103, according to the input feature vector and the predicted flow of each grid at a target time after the target period, using a pre-trained action value network model to predict the reward value of the cell in different working states, wherein the reward value is related to the working state of the cell, the energy consumption value of the base station to which the cell belongs, and the influence of the energy consumption value of the cells in the grid where the cell is located;

[0105] In step S103, the generated input feature vector and the predicted traffic of each grid after the target period are used to predict the reward value of the cell under different working conditions through the pre-trained action value network model. The reward value here is related to the energy consumption value of the cell's working state to the base station and the grid to which it belongs, and the purpose is to guide subsequent decisions through the reward value.

[0106] S104: Control the base station and each subordinate cell of the base station to operate at the target time according to the reward value of each cell in different operating states predicted by the action value network model.

[0107] In step S104, the working state of the base station and each cell is controlled according to the reward value of different working states of each cell of the base station predicted by the action value network model. By selecting the optimal working state according to the predicted reward value, the goals of reducing energy consumption and improving network performance can be achieved.

[0108] The base station control method provided by the present application divides the network area of ​​the base station into discrete grids, extracts characteristic parameters, and predicts the reward value. By predicting the flow and reward value to select the optimal working state, the intelligent control of the base station and the cell can be realized, which can effectively reduce the energy consumption of the base station and each cell, and can also improve the utilization rate of network resources and reduce the idle waste of resources. The reward value is predicted by the pre-trained action value network model, which can realize the automatic decision-making of the working state of the base station and the cell. This automated decision-making process can reduce the need for human intervention and improve the intelligence level of the system.

[0109] like Figure 2 As shown, in some embodiments, dividing the network area covered by the base station into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, may include:

[0110] S1011. Select a first cell from the network area covered by the base station to create a first grid, so that the network coverage of the first grid is equal to the network coverage of the first cell;

[0111] S1012. Determine other cells equivalent to the first cell according to a co-coverage relationship between the first cell and other cells other than the first cell in the affiliated cells of the base station;

[0112] S1013. Add other cells equivalent to the first cell to the first grid;

[0113] S1014. Continue to select a second cell from other subordinate cells of the base station outside the first grid to create a second grid, so that the coverage range of the second grid is equal to the coverage range of the second cell;

[0114] S1015. Determine other cells equivalent to the second cell according to a co-coverage relationship between other subordinate cells of the base station outside the first grid and the second cell;

[0115] S1016. Add other cells equivalent to the second cell to the second grid;

[0116] S1017, and so on, until all cells under the base station are divided into grids.

[0117] Specifically, the network can be divided into multiple discrete grids based on the co-coverage relationship between cells, so that the cells in each grid are equivalent to each other. If the coverage of two cells overlaps so that they can replace each other to communicate with mobile users, then the two cells are equivalent. The co-coverage relationship of cells is modeled using the maximum communication distance, daily communication distance, and inter-cell distance of the cells:

[0118] r i +d ij ≤R j ,r j +d ij ≤R i

[0119] Among them, r i and R i They represent the daily communication distance and maximum communication distance of cell i, respectively, and d ij represents the distance between cell i and cell j. Two cells that satisfy the above formula are considered equivalent cells.

[0120] For example, from the first cell c in the northwest direction of the network 1 (can be replaced by any other starting point) to create a grid M = (c 1 ), find all equivalent cell sets {c 2 ,c 3 ,c 4 ,…}, and select cells from the set one by one to join the grid M until no new cell can be selected that is equivalent to all the cells in M, completing a grid division. Continue to select a cell from the remaining cells and repeat the above process until all cells are divided into the grid.

[0121] like Figure 3 As shown, in some embodiments, for each of the cells, an input feature vector is generated according to the feature parameters of the cell in the target period, the feature parameters of other cells in the grid where the cell is located in the target period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target period, including:

[0122] S1021. For each of the cells, generate a first input feature vector of the cell according to a feature parameter of the cell in a target time period;

[0123] S1022. Generate a second input feature vector for each other cell in the grid where the cell is located according to the feature parameters of each other cell in the target time period;

[0124] S1023: Generate a third input feature vector for each of the other cells that are outside the grid to which the cell belongs and belong to the base station to which the cell belongs within the target time period according to feature parameters of the other cells.

[0125] Specifically, each cell can be taken over by an agent. For the nth cell c n , you can use cell c n The corresponding agent n To observe including with cell c nThe cell IDs under the same base station, the cell IDs in the same virtual grid, and the cell c n And the characteristic parameters of the above-mentioned related cells.

[0126] In the above embodiment, the characteristic parameters of each cell in the target time period include at least one of the following: current time, traffic load of the grid to which the cell belongs within the last m time steps, and equipment parameters of the cell, where m is a positive integer.

[0127] Specifically, each cell c n The characteristic parameters may include the current time t, cell c n Belong to grid g m Traffic load in the last m (e.g. 4) time steps and cell c n The device parameter vector

[0128] like Figure 4 As shown, in some embodiments, according to the input feature vector and the predicted flow of each grid at a target time after the target period, a pre-trained action value network model is used to predict the reward value of the cell in different working states, including:

[0129] S1031, using a first attention layer to integrate the first input feature vector and each of the second input feature vectors to obtain a first integrated embedding vector;

[0130] S1032, using a second attention layer to integrate the first input feature vector and each of the third input feature vectors to obtain a second integrated embedding vector;

[0131] S1033, concatenating the first integrated embedding vector, the second integrated embedding vector, and the global feature vector into a single vector;

[0132] S1034. Use a fully connected layer to process the single vector to obtain reward values ​​of different working states of the cell.

[0133] Specifically, for a cell c n and its agent, first we need to extract nThe capacity and equipment parameters of each cell are related. These parameters are encoded into input feature vectors. These input feature vectors are embedded through two linear layers. In the first linear layer, the original input feature vector is projected into a higher dimensional representation space, and the projected feature vector is further processed in the second linear layer. The embedded representations of other cells from the same grid (or base station) are integrated through a self-attention layer. The attention mechanism allows the model to weight the embeddings of different cells according to relevance, thereby effectively integrating related information. The global feature vector contains global information such as the time step and the traffic demand of each grid at the target time. The integrated embedded representation and global feature vector are concatenated into a single vector. This concatenated single vector contains local cell information and global grid information. The final output represents the predicted reward value of different working states of the cell, guiding the agent to make the best decision in different situations. The above architecture can effectively integrate the local information and global information of related units.

[0134] like Figure 5 As shown, in some embodiments, before predicting the reward value of the cell in different working states using a pre-trained action value network model based on the input feature vector and the predicted traffic of each grid at a target time after the target period, the method further includes:

[0135] S001. Obtain each sample data in the sample set, wherein each sample data includes characteristic parameters of an affiliated cell of the base station in a historical period, characteristic parameters of other cells in the grid where the cell is located in the historical period, characteristic parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, a reward value of the cell at a historical time after the historical period, a time step, a traffic demand of each of the grids at the historical time, a first mask, and a second mask;

[0136] S002. Generate a historical input feature vector according to the feature parameters of the affiliated cell of the base station in each sample data in a historical period, the feature parameters of other cells in the grid where the cell is located in the historical period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period; generate a historical global feature vector according to the time step in the sample data and the traffic demand of each grid at the historical time; and generate a mask vector according to the first mask and the second mask in the sample data;

[0137] S003, taking the historical input feature vector, the historical global feature vector and the mask vector corresponding to each sample data as input, and taking the reward value in the sample data as the target output, training the first action value network model until the output of the first action value network model meets the first preset requirement;

[0138] S004. Taking the historical input feature vector and the historical global feature vector corresponding to each sample data as input, and the output of the first action value network model under the same sample data as the target output, the second action value network model is trained until the output of the second action value network model meets the second preset requirements, and the second action value network model whose output meets the second preset requirements is used as the trained action value network model.

[0139] Specifically, in order to solve the problem of multi-agent cooperation, the policy network is trained to output actions based on the current state. When ordinary multi-agent cooperation algorithms are applied to the problem of base station energy saving, a large number of agents will be a big challenge. On the one hand, the large-scale agents lead to a huge state space, so when an agent makes a decision, it should effectively integrate the states of other agents related to it. To address this problem, two graphs are constructed based on the intra-grid relationship and the intra-base station relationship, and then the graph neural network is used to integrate the information of the associated nodes in the graph. On the other hand, the large-scale agents will also lead to complex interactions between agents, making it difficult to estimate the expected rewards of different actions. Therefore, the idea of ​​the mean field is introduced to integrate the impact of the decisions of other agents on the target agent, and two masks are designed to reduce the difficulty of action value estimation.

[0140] The action value network structure is as follows Figure 6 As shown, for a cell c n and its agent, first we need to extract n The capacity and equipment parameters of each cell are related. These parameters are encoded into input feature vectors. These input feature vectors are embedded through two linear layers. In the first linear layer, the original input feature vector is projected into a higher dimensional representation space, and the projected feature vector is further processed in the second linear layer. The embedding representations of other cells from the same grid (or base station) are integrated through a self-attention layer. The attention mechanism allows the model to weight the embeddings of different cells according to their relevance, thereby effectively integrating related information. The global feature vector contains global information such as time step and traffic demand in the grid. Masks are used to handle the validity of input data to ensure that the model only processes valid data. The integrated embedding representation, global feature vector and mask are concatenated into a single vector. This concatenated vector contains local cell information and global grid information. The final output represents the predicted rewards for different actions, guiding the agent to make optimal decisions in different situations. The above architecture can effectively integrate local and global information of related units.

[0141] In order to train the action-value network, the effects of the actions of all other agents (or neighboring agents) are simplified to two masks: the first mask is used to indicate whether the traffic demand of the grid can be met if the agent sets the corresponding cell to sleep. The second mask is used to indicate whether all other cells under the same base station are set to sleep. In other words, the mask indicates whether the base station can be turned off if the cell is set to sleep mode. These two masks depend on the actions of the relevant agents and directly determine the rewards of the agents. Since these masks come from the actions of other agents, they can only be used as training inputs, so each agent consists of two action-value networks, namely and In the execution phase, the action-value network with unmasked input Predict the rewards associated with different actions and set the corresponding unit activation or deactivation based on the observations. Using the ε-greedy method, the agent determines the state of the unit by sampling according to the predicted rewards. After the unit states determined by all agents are executed, the rewards of all agents are calculated as mentioned above. The observations, actions and rewards of all agents are added to the replay buffer. During the training phase, the action value network with the mask as input The replay buffer is used for training to predict the rewards of different actions based on the observations and masks. The network is optimized with the following loss function:

[0142]

[0143] Among them, the mask m n is obtained based on the actions recorded in the replay buffer. The second action-value network By imitating the first action-value network To optimize,

[0144]

[0145] The mask m n is calculated from the resampled actions, which are the sampling process during the execution phase. The resampling process is equivalent to the sampling process of the mask vector distribution. Therefore, the second action-value network will learn to predict the rewards of various actions, while the policies of the other agents remain unchanged.

[0146] In some embodiments, the reward value of each cell is negatively correlated to the sum of the energy consumption of the remote radio units of cells belonging to the same grid as the cell, and the energy consumption of the baseband unit and the cooling equipment of the base station to which the cell belongs.

[0147] Specifically, the goal of decision-making is to save the total energy consumption of cooling equipment including RRU, BBU and air conditioner, so the design of rewards should encourage agents to cooperate to save energy. When the control status of other cells remains unchanged, the change of the control status of a cell can only affect the energy consumption of its base station and the power grid. Therefore, the reward of each cell can be set to The first term represents the n The energy consumption of RRUs in cells belonging to the same grid, so agents can be encouraged to cooperate with agents in these cells The second term represents cell c n The energy consumption of the BBU and cooling equipment of the base station to which they belong can encourage the agents to cooperate with agents from the same base station.

[0148] like Figure 7 As shown, in some embodiments, controlling the base station and each of the subordinate cells of the base station to work at the target time according to the reward value of each of the cells in different working states predicted by the action value network model includes:

[0149] S1041. For each affiliated cell of the base station, determine the preliminary working state of the cell according to the reward value of different working states of the cell predicted by the action value network model, wherein the working state of each cell includes an activated state, a closed state and a dormant state;

[0150] S1042. For each grid under the base station, if the total capacity of the cells in the grid whose initial working state is activated is less than the predicted total traffic demand, modify the cells whose initial working state is closed to activated in descending order of cell capacity until the total capacity of the cells whose working state is activated is greater than the total traffic demand;

[0151] S1043, dividing the total traffic demand predicted in each grid according to the capacity of cells in the activated state, and then allocating it to corresponding cells;

[0152] S1044. If all cells in a base station correspond to a closed state, the base station and each subordinate cell of the base station are controlled to be closed; otherwise, the base station is controlled to be turned on, and the cell is controlled to work according to the corresponding working state and allocated traffic of each subordinate cell of the base station.

[0153] Specifically, the optional action of the agent is to set its corresponding cell to an activated state, a dormant state, or a deactivated state. According to the decision of the agent, the final decision on the control and flow distribution of the base station and the cell is obtained through the following steps. First, the action output of the agent is adjusted. If the total capacity of the activated cells in the grid is less than the flow demand, the deactivated cells are activated from large to small according to the cell capacity until the total capacity of the activated cells is greater than the demand. Secondly, the predicted flow is distributed on the activated cells. The predicted total flow demand in the grid is divided according to the capacity of the activated cells and then allocated to the corresponding cells. Finally, the control state of the base station and the cell is obtained. If all cells in the base station are in a deactivated state, the base station and its cells will be turned off; otherwise, the control state of the base station is set to open, and if the state of the cell is activated (or deactivated), the state of the cell is set to open (or dormant).

[0154] In order to better understand the present application, the base station control method provided by the present application is described in detail below through a specific embodiment.

[0155] The goal of the base station control method provided in this embodiment is to minimize the energy consumption of the entire system by performing on, off, and sleep control and traffic load distribution on the base station and its affiliated cells, while ensuring that the user's traffic demand can be met. Input: The geographical location of each base station and cell, the BBU energy consumption and cooling equipment energy consumption in each base station, the relationship between the RRU energy consumption corresponding to each cell and the traffic it carries, the affiliation between the cell and the base station, and the traffic carried by each cell. Output: The on and off status of each base station, the on, off, and sleep status of each cell, and how much traffic the on cell carries for which cells.

[0156] Considering that the user's traffic demand is constantly changing, a deep neural network is used to determine the control status of the base station and the cell and the traffic carried by each cell based on the relationship between the base station, the cell and the user and the user's traffic demand. In order to efficiently train the neural network to minimize the energy consumption of the system, mean field reinforcement learning is constructed to handle large-scale base station control problems. The specific process is as follows:

[0157] The workflow of the base station control method provided in this embodiment includes: a. cell equivalence relationship division, constructing the equivalence relationship between cells based on the geographical location and coverage radius of the cells; b. energy consumption model construction, constructing an energy consumption model with the base station and cell control status and cell load conditions as input; c. decision problem construction, converting the optimization problem into a decision problem with equipment parameters and user needs as status, energy saving efficiency as reward, and cell control as action; d. strategy optimization based on mean field reinforcement learning, for the decision problem, constructing a graph neural network as a policy network, and using mean field reinforcement learning to optimize the policy network.

[0158] a. Division of equivalence relationships among communities

[0159] The network is divided into multiple discrete grids based on the co-coverage relationship between cells, so that the cells in each grid are equivalent to each other. If the coverage of two cells overlaps so that they can replace each other to communicate with mobile users, then the two cells are equivalent. The co-coverage relationship of cells is modeled using the maximum communication distance, daily communication distance and inter-cell distance of the cells

[0160] r i +d ij ≤R j ,r j +d ij ≤R i

[0161] where r i and R i They represent the daily communication distance and maximum communication distance of cell i, respectively, and d ij represents the distance between cell i and cell j. Two cells that satisfy the above formula are considered equivalent cells.

[0162] For example, from the first cell c in the northwest direction of the network 1 (can be replaced by any other starting point) to create a grid M = (c 1 ), find all equivalent cell sets {c 2 ,c 3 ,c 4 ,…}, and select cells from the set one by one to join the grid M until no new cell can be selected that is equivalent to all the cells in M, completing a grid division. Continue to select a cell from the remaining cells and repeat the above process until all cells are divided into the grid.

[0163] b. Energy consumption model construction

[0164] A base station usually consists of a communication subsystem and a support subsystem. The remote radio unit (RRU) and baseband unit (BBU) in the communication subsystem are responsible for sending and receiving radio signals and processing baseband signals, respectively. A base station can consist of multiple RRUs and BBUs. Cooling and other auxiliary equipment are part of the support subsystem.

[0165] According to the composition of the base station, a base station b k Total power consumption:

[0166]

[0167] in, represents the power consumption of the communication subsystem, Indicates the power consumed by the cooling equipment to maintain the appropriate operating temperature, roughly the same as The power consumption of the communication subsystem mainly consists of two parts: the power consumption of RRU and BBU.

[0168]

[0169] in, and Represent the power consumption of BBU and RRU respectively.

[0170] BBU is responsible for baseband processing, and its power consumption They all remain relatively constant, are not affected by base station traffic, and are only related to base station hardware.

[0171] When the RRU is turned on, its power consumption is related to the corresponding cell c n Load flow T n There is an approximate linear relationship between

[0172]

[0173] Among them, α n and β n They represent the slope and offset respectively. Due to hardware differences, different base stations have different slopes and offsets. When the RRU is in sleep mode, its power consumption is It is also a fixed value related to hardware facilities.

[0174] The above parameters can be obtained by performing linear regression on real data.

[0175] c. Decision Problem Construction

[0176] The goal of the base station control method provided in this application is to minimize energy consumption by controlling the operating status of base stations and cells and traffic load distribution while ensuring that all user needs have corresponding cells to carry out the load. Among the three subtasks of base station control, cell control and traffic load distribution, cell control is dominant. Once cell control is obtained, base station control and traffic distribution can be obtained through simple rules.

[0177] Cell control is a multi-agent cooperation problem: each cell c n By an agent n Management, these agents cooperate to minimize the energy consumption of the system. In this problem, each agent outputs an action based on its own observation (part of the system state), and then the actions of all agents act together on the system. The system changes its state based on the joint action and feedbacks rewards to each agent.

[0178] Observation: Each agent n The observed values ​​include those with cell c n The cell IDs under the same base station, the cell IDs in the same virtual grid, and the cell cn And the feature vector of the above-mentioned related cells. Each cell c n The characteristic vector includes time t, cell c n Belong to grid g m Traffic load in the last 4 time steps and cell c n The device parameter vector

[0179] Action: The optional action of the agent is to set its corresponding cell to an activated state or a deactivated state. Based on the decision of the agent, the final decision on base station and cell control and traffic allocation is obtained through the following steps. First, adjust the action output of the agent. If the total capacity of the activated cells in the grid is less than the traffic demand, activate the deactivated cells from large to small according to the cell capacity until the total capacity of the activated cells is greater than the demand. Secondly, allocate the predicted traffic on the activated cells. The predicted total traffic demand in the grid is divided according to the capacity of the activated cells and then allocated to the corresponding cells. Finally, obtain the control status of the base station and cells. If all cells in the base station are in a deactivated state, the base station and its cells will be turned off; otherwise, the control state of the base station is set to on, and if the state of the cell is activated (or deactivated), the state of the cell is set to on (or dormant).

[0180] Reward: The goal of decision-making is to save the total energy consumption including RRU, BBU and air conditioner, so the design of reward should encourage agents to cooperate to save energy. When the control state of other cells remains unchanged, the change of the control state of a cell can only affect the energy consumption of its base station and the power grid. Therefore, the reward of each agent is set to The first term represents the n The energy consumption of RRUs in cells belonging to the same grid, so agents can be encouraged to cooperate with agents in these cells The second term represents cell c n The energy consumption of the BBU and cooling equipment of the base station to which they belong can encourage the agents to cooperate with agents from the same base station.

[0181] d. Strategy optimization based on mean field reinforcement learning

[0182] In order to solve the problem of multi-agent cooperation, the policy network is trained to output actions based on the current state. When ordinary multi-agent cooperation algorithms are applied to the problem of base station energy saving, a large number of agents will be a big challenge. On the one hand, the large-scale agents lead to a huge state space, so when an agent makes a decision, it should effectively integrate the states of other agents related to it. To address this problem, two graphs are constructed based on the intra-grid relationship and the intra-base station relationship, and then the graph neural network is used to integrate the information of the associated nodes in the graph. On the other hand, the large-scale agents also lead to complex interactions between agents, making it difficult to estimate the expected rewards of different actions. Therefore, the idea of ​​the mean field is introduced to integrate the impact of the decisions of other agents on the target agent, and two masks are designed to reduce the difficulty of action value estimation.

[0183] The action value network structure is as follows: For a cell c n and its agent, first we need to extract n The capacity and equipment parameters of each cell are related. These parameters are encoded into feature vectors. These feature vectors are embedded through two linear layers. In the first linear layer, the original feature vector is projected into a higher dimensional representation space. The projected feature vector is further processed in the second linear layer. The embedding representations of other cells from the same grid (or base station) are integrated through a self-attention layer. The attention mechanism allows the model to weight the embeddings of different cells according to their relevance, thereby effectively integrating related information. The global feature vector contains global information such as time step and traffic demand in the grid. Masks are used to handle the validity of input data to ensure that the model only processes valid data. The integrated embedding representation, global feature vector and mask are concatenated into a single vector. This concatenated vector contains local cell information and global grid information. The final output represents the predicted rewards for different actions, guiding the agent to make the best decision in different situations. The above architecture can effectively integrate local and global information of related units.

[0184] In order to train the action-value network, the effects of the actions of all other agents (or neighboring agents) are simplified to two masks: the first mask is whether the traffic demand of the grid can be met if the agent sets the corresponding cell to sleep. The second mask is whether all other cells under the same base station are set to sleep. In other words, the mask indicates whether the base station can be turned off if the cell is set to sleep mode. These two masks depend on the actions of the relevant agents and directly determine the rewards of the agents. Since these masks come from the actions of other agents, they can only be used as training inputs, so each agent consists of two action-value networks, namely and In the execution phase, the action-value network with unmasked input Predict the rewards associated with different actions and set the corresponding unit activation or deactivation based on the observations. Using the ε-greedy method, the agent determines the state of the unit by sampling according to the predicted rewards. After the unit states determined by all agents are executed, the rewards of all agents are calculated as mentioned above. The observations, actions and rewards of all agents are added to the replay buffer. During the training phase, the action value network with the mask as input The replay buffer is used for training to predict the rewards of different actions based on the observations and masks. The network is optimized with the following loss function:

[0185]

[0186] Among them, the mask m n is obtained based on the actions recorded in the replay buffer. The second action-value network By imitating the first action-value network To optimize,

[0187]

[0188] The mask m n is calculated from the resampled actions, which are the sampling process during the execution phase. The resampling process is equivalent to the sampling process of the mask vector distribution. Therefore, the second network will learn to predict the rewards for various actions, while the policies of the other agents remain unchanged.

[0189] In order to make the technical solution of the base station control method provided by the above embodiment clearer, a specific implementation method is given below.

[0190] The user wants to optimize the energy consumption of future traffic load conditions based on the hardware parameters of the base station cells in a certain area and the traffic load conditions for a week. This dataset contains the following data:

[0191] Table 1: Dataset information

[0192]

[0193]

[0194] First, according to the method in a (division based on equivalent relationship of cells), the space is divided into multiple virtual grids based on the geographical location and coverage radius of the cells. The cells in each grid are equivalent, and the equivalent relationship between the cells is recorded.

[0195] Then, according to the method in b (energy consumption model construction), an energy consumption model is constructed based on the cooling facility energy consumption, BBU energy consumption and RRU energy consumption coefficient.

[0196] Then, according to the method in c (decision problem construction), the energy consumption optimization problem is transformed into a multi-agent cooperation problem: each cell is managed by an agent, and these agents cooperate to minimize the system energy consumption. The observations of each agent include the cell ID under the same base station as the cell, the cell ID in the same virtual grid, and the feature vectors of the cell and the above-mentioned related cells. The feature vector of each cell includes time, the traffic load of the grid to which the cell belongs in the last 4 time steps, and the energy consumption coefficient of the RRU corresponding to the cell. The action is to set the corresponding cell to an activated state or a deactivated state. The reward is the negative of the sum of the energy consumption of the RRU of the cell belonging to the same grid as the cell and the energy consumption of the BBU and refrigeration equipment of the base station to which the cell belongs.

[0197] Then, according to the method in d (Policy optimization based on mean field reinforcement learning), the policy is optimized based on mean field reinforcement learning. Figure 8 , for this decision-making problem, a graph neural network is constructed as a policy network, and mean field reinforcement learning is used to optimize the policy network. The structure of the graph neural network is: the capacity and equipment parameters of the cell and its related cells are respectively embedded into a 16-dimensional vector through two linear layers. The embeddings of cells from the same grid (or base station) are integrated using the attention layer to obtain an integrated embedding. The two integrated embeddings, global values ​​(such as time steps and traffic requirements in the grid), and masks are concatenated into a single vector, and then passed through two fully connected layers to obtain the final output, that is, the predicted rewards for different operations. For each cell, the effects of all other intelligent actions are simplified to two masks, and the reward prediction ability of the policy network is trained based on these masks.

[0198] Finally, the strategic network is deployed, and the base station and cell control status and the cell traffic load are adjusted according to the actual traffic demand.

[0199] Based on the same inventive concept, the embodiment of the present application also provides a base station control device, such as Fig. 9 As shown, the base station control device provided in the embodiment of the present application includes:

[0200] A division module 21 is used to divide the network area covered by the base station into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, and the number of the base station is at least one;

[0201] A first generating module 22 is used to generate an input feature vector for each cell according to the feature parameters of the cell in the target period, the feature parameters of other cells in the grid where the cell is located in the target period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target period;

[0202] A prediction module 23 is used to predict the reward value of the cell in different working states by using a pre-trained action value network model according to the input feature vector and the predicted flow of each grid at a target time after the target period, wherein the reward value is related to the influence of the working state of the cell, the energy consumption value of the base station to which the cell belongs, and the energy consumption value of the cells in the grid where the cell is located;

[0203] The control module 24 is used to control the base station and each subordinate cell of the base station to work at the target time according to the reward value of each cell in different working states predicted by the action value network model.

[0204] The base station control device provided by the present application divides the network area of ​​the base station into discrete grids, extracts characteristic parameters, and predicts the reward value. By predicting the flow and reward value to select the optimal working state, the intelligent control of the base station and the cell can be realized, which can effectively reduce the energy consumption of the base station and each cell, and can also improve the utilization rate of network resources and reduce idle waste of resources. The reward value is predicted by the pre-trained action value network model, which can realize the automatic decision-making of the working state of the base station and the cell. This automated decision-making process can reduce the need for human intervention and improve the intelligence level of the system.

[0205] In some embodiments, the partitioning module is specifically used to:

[0206] Selecting a first cell from the network area covered by the base station to create a first grid, so that the network coverage of the first grid is equal to the network coverage of the first cell;

[0207] Determine other cells equivalent to the first cell according to the co-coverage relationship between the first cell and other cells other than the first cell in the affiliated cells of the base station;

[0208] Adding other cells equivalent to the first cell to the first grid;

[0209] Continue to select a second cell from other subordinate cells of the base station outside the first grid to create a second grid, so that the coverage range of the second grid is equal to the coverage range of the second cell;

[0210] Determine other cells equivalent to the second cell according to the co-coverage relationship between other subordinate cells of the base station outside the first grid and the second cell;

[0211] Adding other cells equivalent to the second cell to the second grid;

[0212] This process is deduced in this way until all cells under the base station are divided into grids.

[0213] In some embodiments, the first generating module is specifically used for:

[0214] For each of the cells, generating a first input feature vector of the cell according to the feature parameters of the cell in the target time period;

[0215] Generate a second input feature vector for each other cell in the grid where the cell is located according to the feature parameters of each other cell in the target time period;

[0216] A third input feature vector of each of the other cells is generated according to feature parameters of other cells that are outside the grid to which the cell belongs and belong to the base station to which the cell belongs within the target time period.

[0217] In some embodiments, the characteristic parameters of each of the cells in the target time period include at least one of the following: current time, traffic load of the grid to which the cell belongs within the last m time steps, and equipment parameters of the cell, where m is a positive integer.

[0218] In some embodiments, the prediction module is specifically used to:

[0219] Integrate the first input feature vector and each of the second input feature vectors using a first attention layer to obtain a first integrated embedding vector;

[0220] Integrate the first input feature vector and each of the third input feature vectors using a second attention layer to obtain a second integrated embedding vector;

[0221] Concatenate the first integrated embedding vector, the second integrated embedding vector, and the global feature vector into a single vector;

[0222] The single vector is processed by using a fully connected layer to obtain reward values ​​of different working states of the cell.

[0223] In some embodiments, the global feature vector is generated according to the time step and the flow demand of each of the grids at the target time.

[0224] In some embodiments, the apparatus further comprises:

[0225] An acquisition module, used to acquire each sample data in the sample set, wherein each sample data includes characteristic parameters of an affiliated cell of the base station in a historical period, characteristic parameters of other cells in the grid where the cell is located in the historical period, characteristic parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, a reward value of the cell at a historical time after the historical period, a time step, a flow demand of each of the grids at the historical time, a first mask, and a second mask;

[0226] A second generating module is used to generate a historical input feature vector according to the feature parameters of the affiliated cell of the base station in each sample data in a historical period, the feature parameters of other cells in the grid where the cell is located in the historical period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, generate a historical global feature vector according to the time step in the sample data and the traffic demand of each of the grids at the historical time, and generate a mask vector according to the first mask and the second mask in the sample data;

[0227] A first training module is used to take the historical input feature vector, the historical global feature vector and the mask vector corresponding to each sample data as input, and the reward value in the sample data as the target output, to train the first action value network model until the output of the first action value network model meets the first preset requirement;

[0228] The second training module is used to take the historical input feature vector and the historical global feature vector corresponding to each sample data as input, and the output of the first action value network model under the same sample data as the target output, to train the second action value network model until the output of the second action value network model meets the second preset requirements, and use the second action value network model whose output meets the second preset requirements as the trained action value network model.

[0229] In some embodiments, for a sample data, the first mask in the sample data indicates whether the traffic demand of the grid where the cell is located can be met if the subsidiary cell in the sample data is set to a sleep state; the second mask in the sample data indicates whether the base station to which the cell belongs can be shut down if the subsidiary cell is set to a sleep state.

[0230] In some embodiments, the reward value of each cell is negatively correlated to the sum of the energy consumption of the remote radio units of cells belonging to the same grid as the cell, and the energy consumption of the baseband unit and the cooling equipment of the base station to which the cell belongs.

[0231] In some embodiments, the control module is specifically used to:

[0232] For each subordinate cell of the base station, determining a preliminary working state of the cell according to reward values ​​of different working states of the cell predicted by the action value network model, wherein the working state of each cell includes an activated state, a deactivated state and a dormant state;

[0233] For each grid under the base station, if the total capacity of the cells in the grid whose initial working state is activated is less than the predicted total traffic demand, the cells whose initial working state is closed are modified to activated in descending order of cell capacity until the total capacity of the cells whose working state is activated is greater than the total traffic demand;

[0234] The predicted total traffic demand in each grid is divided according to the capacity of the activated cells and then allocated to the corresponding cells;

[0235] If all cells in a base station correspond to the closed state, the base station and each subordinate cell of the base station are controlled to be closed; otherwise, the base station is controlled to be turned on, and the cell is controlled to work according to the corresponding working state and allocated traffic of each subordinate cell of the base station.

[0236] The embodiments of the device provided in the embodiments of the present application can be specifically used to execute the processing flow applied to the above-mentioned various method embodiments. Its functions are not repeated here, and reference can be made to the detailed description of the above-mentioned method embodiments.

[0237] Fig.10 A schematic diagram of the physical structure of an electronic device provided in an embodiment of the present application is shown in FIG. Fig.10 As shown, the electronic device may include: a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302 and the memory 303 communicate with each other through the communication bus 304. The processor 301 may call the logic instructions in the memory 303 to execute the method described in any of the above embodiments.

[0238] In addition, the logic instructions in the above-mentioned memory 303 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.

[0239] This embodiment discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided by the above-mentioned method embodiments.

[0240] This embodiment provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program enables the computer to execute the methods provided by the above-mentioned method embodiments.

[0241] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0242] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that the processes and / or blocks in the flowcharts and / or block diagrams, as well as the combination of the processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0243] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0244] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0245] In the description of this specification, the description with reference to the terms "one embodiment", "a specific embodiment", "some embodiments", "for example", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0246] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A base station control method, characterized in that: include: Divide the network area covered by the base station into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, and the number of the base station is at least one; For each of the cells, an input feature vector is generated according to feature parameters of the cell in the target period, feature parameters of other cells in the grid where the cell is located in the target period, and feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target period; According to the input feature vector and the predicted flow of each grid at a target time after the target period, a pre-trained action value network model is used to predict the reward value of the cell in different working states, wherein the reward value is related to the influence of the working state of the cell on the energy consumption value of the base station to which the cell belongs and the energy consumption value of the cells in the grid where the cell is located; According to the reward value of each of the cells in different working states predicted by the action value network model, the base station and each of the cells attached to the base station are controlled to work at the target time.

2. The method according to claim 1, characterized in that The base station divides the network area covered by the base station into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, including: Selecting a first cell in the network area covered by the base station to create a first grid, so that the network coverage of the first grid is equal to the network coverage of the first cell; Determine other cells equivalent to the first cell according to the co-coverage relationship between the first cell and other cells other than the first cell in the affiliated cells of the base station; Adding other cells equivalent to the first cell to the first grid; Continue to select a second cell from other subordinate cells of the base station outside the first grid to create a second grid, so that the coverage range of the second grid is equal to the coverage range of the second cell; Determine other cells equivalent to the second cell according to the co-coverage relationship between other subordinate cells of the base station outside the first grid and the second cell; Adding other cells equivalent to the second cell to the second grid; This process is deduced in this way until all cells under the base station are divided into grids.

3. The method according to claim 2, characterized in that For each of the cells, generating an input feature vector according to feature parameters of the cell in a target period, feature parameters of other cells in the grid where the cell is located in the target period, and feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target period, including: For each of the cells, generating a first input feature vector of the cell according to the feature parameters of the cell in the target time period; Generate a second input feature vector for each other cell in the grid where the cell is located according to the feature parameters of each other cell in the target time period; A third input feature vector of each of the other cells is generated according to feature parameters of other cells that are outside the grid to which the cell belongs and belong to the base station to which the cell belongs within the target time period.

4. The method according to claim 3, characterized in that The characteristic parameters of each cell in the target time period include at least one of the following: current time, traffic load of the grid to which the cell belongs within the latest m time steps, and equipment parameters of the cell, where m is a positive integer.

5. The method according to claim 4, characterized in that According to the input feature vector and the predicted flow of each grid at a target time after the target period, a pre-trained action value network model is used to predict the reward value of the cell in different working states, including: Integrate the first input feature vector and each of the second input feature vectors using a first attention layer to obtain a first integrated embedding vector; Integrate the first input feature vector and each of the third input feature vectors using a second attention layer to obtain a second integrated embedding vector; Concatenate the first integrated embedding vector, the second integrated embedding vector, and the global feature vector into a single vector; The single vector is processed by using a fully connected layer to obtain reward values ​​of different working states of the cell.

6. The method according to claim 5, characterized in that The global feature vector is generated according to the time step and the flow demand of each of the grids at the target time.

7. The method according to claim 6, characterized in that Before predicting the reward value of the cell in different working states using a pre-trained action value network model based on the input feature vector and the predicted flow of each grid at a target time after the target period, the method further includes: Acquire each sample data in the sample set, wherein each sample data includes characteristic parameters of an affiliated cell of the base station in a historical period, characteristic parameters of other cells in the grid where the cell is located in the historical period, characteristic parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, a reward value of the cell at a historical time after the historical period, a time step, a traffic demand of each grid of the base station at the historical time, a first mask, and a second mask; Generate a historical input feature vector according to the feature parameters of the affiliated cell of the base station in each sample data in a historical period, the feature parameters of other cells in the grid where the cell is located in the historical period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period; generate a historical global feature vector according to the time step in the sample data and the traffic demand of each grid of the base station at the historical time; and generate a mask vector according to the first mask and the second mask in the sample data; Taking the historical input feature vector, the historical global feature vector and the mask vector corresponding to each sample data as input, and taking the reward value in the sample data as the target output, the first action value network model is trained until the output of the first action value network model meets the first preset requirement; Taking the historical input feature vector and the historical global feature vector corresponding to each sample data as input, and the output of the first action value network model under the same sample data as the target output, the second action value network model is trained until the output of the second action value network model meets the second preset requirements, and the second action value network model whose output meets the second preset requirements is used as the trained action value network model.

8. The method according to claim 7, characterized in that For a sample data, the first mask in the sample data indicates whether the traffic demand of the grid where the cell is located can be met if the subsidiary cell in the sample data is set to a sleep state; the second mask in the sample data indicates whether the base station to which the cell belongs can be shut down if the subsidiary cell is set to a sleep state.

9. The method according to claim 7, characterized in that: The reward value of each cell is negatively correlated to the sum of the energy consumption of the remote radio units of the cells belonging to the same grid as the cell, and the energy consumption of the baseband unit and the cooling equipment of the base station to which the cell belongs.

10. The method according to any one of claims 1 to 9, characterized in that: The step of controlling the base station and each of the cells attached to the base station to work at the target time according to the reward value of each cell in different working states predicted by the action value network model includes: For each subordinate cell of the base station, determining a preliminary working state of the cell according to reward values ​​of different working states of the cell predicted by the action value network model, wherein the working state of each cell includes an activated state, a deactivated state and a dormant state; For each grid under the base station, if the total capacity of the cells in the grid whose initial working state is activated is less than the predicted total traffic demand, the cells whose initial working state is closed are modified to activated in descending order of cell capacity until the total capacity of the cells whose working state is activated is greater than the total traffic demand; The predicted total traffic demand in each grid is divided according to the capacity of the activated cells and then allocated to the corresponding cells; If all cells in a base station correspond to the closed state, the base station and each subordinate cell of the base station are controlled to be closed; otherwise, the base station is controlled to be turned on, and the cell is controlled to work according to the corresponding working state and allocated traffic of each subordinate cell of the base station.

11. A base station control device, characterized in that: include: A division module, configured to divide the network area covered by the base station into at least two discrete grids according to the co-coverage relationship between the affiliated cells of the base station, so that the cells in each grid are equivalent to each other, and the number of the base station is at least one; A first generating module is used to generate an input feature vector for each cell according to the feature parameters of the cell in the target time period, the feature parameters of other cells in the grid where the cell is located in the target time period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the target time period; A prediction module, used to predict the reward value of the cell in different working states by using a pre-trained action value network model according to the input feature vector and the predicted flow of each grid at a target time after the target period, wherein the reward value is related to the influence of the working state of the cell, the energy consumption value of the base station to which the cell belongs, and the energy consumption value of the cells in the grid where the cell is located; The control module is used to control the operation of the base station and each subordinate cell of the base station at the target time according to the reward value of each cell in different working states predicted by the action value network model.

12. The device according to claim 11, characterized in that The division module is specifically used for: Selecting a first cell from the network area covered by the base station to create a first grid, so that the network coverage of the first grid is equal to the network coverage of the first cell; Determine other cells equivalent to the first cell according to the co-coverage relationship between the first cell and other cells other than the first cell in the affiliated cells of the base station; Adding other cells equivalent to the first cell to the first grid; Continue to select a second cell from other subordinate cells of the base station outside the first grid to create a second grid, so that the coverage range of the second grid is equal to the coverage range of the second cell; Determine other cells equivalent to the second cell according to the co-coverage relationship between other subordinate cells of the base station outside the first grid and the second cell; Adding other cells equivalent to the second cell to the second grid; This process is deduced in this way until all cells under the base station are divided into grids.

13. The device according to claim 12, characterized in that The first generating module is specifically used for: For each of the cells, generating a first input feature vector of the cell according to the feature parameters of the cell in the target time period; Generate a second input feature vector for each other cell in the grid where the cell is located according to the feature parameters of each other cell in the target time period; A third input feature vector of each of the other cells is generated according to feature parameters of other cells that are outside the grid to which the cell belongs and belong to the base station to which the cell belongs within the target time period.

14. The device according to claim 13, characterized in that The characteristic parameters of each cell in the target time period include at least one of the following: current time, traffic load of the grid to which the cell belongs within the latest m time steps, and equipment parameters of the cell, where m is a positive integer.

15. The device according to claim 14, characterized in that The prediction module is specifically used for: Integrate the first input feature vector and each of the second input feature vectors using a first attention layer to obtain a first integrated embedding vector; Integrate the first input feature vector and each of the third input feature vectors using a second attention layer to obtain a second integrated embedding vector; Concatenate the first integrated embedding vector, the second integrated embedding vector, and the global feature vector into a single vector; The single vector is processed by using a fully connected layer to obtain reward values ​​of different working states of the cell.

16. The device according to claim 15, characterized in that The global feature vector is generated according to the time step and the flow demand of each of the grids at the target time.

17. The device according to claim 16, characterized in that The device also includes: An acquisition module, used to acquire each sample data in the sample set, wherein each sample data includes characteristic parameters of an affiliated cell of the base station in a historical period, characteristic parameters of other cells in the grid where the cell is located in the historical period, characteristic parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, a reward value of the cell at a historical time after the historical period, a time step, a flow demand of each of the grids at the historical time, a first mask, and a second mask; A second generating module is used to generate a historical input feature vector according to the feature parameters of the affiliated cell of the base station in each sample data in a historical period, the feature parameters of other cells in the grid where the cell is located in the historical period, and the feature parameters of other cells outside the grid to which the cell belongs and belonging to the base station to which the cell belongs in the historical period, generate a historical global feature vector according to the time step in the sample data and the traffic demand of each of the grids at the historical time, and generate a mask vector according to the first mask and the second mask in the sample data; A first training module is used to take the historical input feature vector, the historical global feature vector and the mask vector corresponding to each sample data as input, and the reward value in the sample data as the target output, to train the first action value network model until the output of the first action value network model meets the first preset requirement; The second training module is used to take the historical input feature vector and the historical global feature vector corresponding to each sample data as input, and the output of the first action value network model under the same sample data as the target output, to train the second action value network model until the output of the second action value network model meets the second preset requirements, and use the second action value network model whose output meets the second preset requirements as the trained action value network model.

18. The device according to claim 17, characterized in that For a sample data, the first mask in the sample data indicates whether the traffic demand of the grid where the cell is located can be met if the subsidiary cell in the sample data is set to a sleep state; the second mask in the sample data indicates whether the base station to which the cell belongs can be shut down if the subsidiary cell is set to a sleep state.

19. The device according to claim 17, characterized in that The reward value of each cell is negatively correlated to the sum of the energy consumption of the remote radio units of the cells belonging to the same grid as the cell, and the energy consumption of the baseband unit and the cooling equipment of the base station to which the cell belongs.

20. The device according to any one of claims 11 to 19, characterized in that The control module is specifically used for: For each subordinate cell of the base station, determining a preliminary working state of the cell according to reward values ​​of different working states of the cell predicted by the action value network model, wherein the working state of each cell includes an activated state, a deactivated state and a dormant state; For each grid under the base station, if the total capacity of the cells in the grid whose initial working state is activated is less than the predicted total traffic demand, the cells whose initial working state is closed are modified to activated in descending order of cell capacity until the total capacity of the cells whose working state is activated is greater than the total traffic demand; The predicted total traffic demand in each grid is divided according to the capacity of the activated cells and then allocated to the corresponding cells; If all cells in a base station correspond to the closed state, the base station and each subordinate cell of the base station are controlled to be closed; otherwise, the base station is controlled to be turned on, and the cell is controlled to work according to the corresponding working state and allocated traffic of each subordinate cell of the base station.

21. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

23. A computer program product, the computer program product comprising a computer program stored on a non-transitory computer readable storage medium, the computer program comprising program instructions, characterized in that: When the program instructions are executed by a computer, the computer can perform the steps of the method according to any one of claims 1 to 10.