Cell Cooperative Sleep Strategy Generation Model Training Method and Cell Cooperative Sleep Method

By training the neural network based on historical traffic data and expert demonstration action data, combined with deep reinforcement learning and transfer learning, the cell collaborative sleep strategy is optimized, and the problem of poor training effect and limited decision-making of the cell collaborative sleep strategy generation model in the dynamic environment is solved, and adaptive cell switching decisions are realized.

CN119136282BActive Publication Date: 2025-07-22BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411101922.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-12
Publication Date
2025-07-22
Estimated Expiration
2044-08-12

AI Technical Summary

Technical Problem

When facing the dynamics and complexity of the communication environment, the existing cell collaborative sleeping strategy generation model has poor results in traditional deep reinforcement learning, limited decision making in imitation learning, unable to adapt to a highly dynamic environment, and the reward function fluctuates greatly, resulting in poor learning results.

Method used

By training the neural network based on historical traffic data and expert demonstration action data, the current and target neural network of the reinforcement learning agent is initialized, the deep reinforcement learning method is used for joint iterative training, combined with expert demonstration network parameters, the cell collaborative sleep strategy is optimized, and the transfer learning idea is introduced to adapt to the dynamics and complexity of the environment.

Benefits of technology

The training effectiveness and reliability of the cell collaborative sleep strategy generation model is improved, and the problem of poor reinforcement learning training results caused by environmental dynamics and complexity is solved, the decision-making ability in imitation learning is enhanced, and dynamic and adaptive cell switching decisions are realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119136282B_ABST
    Figure CN119136282B_ABST
Patent Text Reader

Abstract

The present application provides a method for training a cell collaborative sleep strategy generation model and a cell collaborative sleep method. The training method includes: training an expert demonstration network based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the collaborative sleep of each cell; migrating the expert demonstration network to the current neural network and the target neural network of the reinforcement learning agent, so as to continue to learn the expert demonstration results based on the deep reinforcement learning method on the basis of this network, continuously update and optimize the strategy, and train a cell collaborative sleep strategy generation model. The present application can improve the training effectiveness and reliability of the cell collaborative sleep strategy generation model, and can solve the problem of limited decision-making in imitation learning. Furthermore, it can improve the application effectiveness and reliability of the cell collaborative sleep strategy generated based on the cell collaborative sleep strategy generation model, so as to achieve dynamic and adaptive cell switching decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of wireless communication networks, and particularly to a method for training a cell collaborative sleep strategy generation model and a cell collaborative sleep method. Background Art

[0002] With the rapid development of modern communication technologies, the communication load has increased explosively. The emergence of 5G communication technology has led to the large-scale deployment of 5G cells. It is statistically shown that cells account for 60% - 80% of the power consumption of communication technologies. The large amount of power consumed by cells not only increases the CO2 emissions, but also additionally increases the operating cost (Operating Expense, OPEX) of operators. In recent years, with the wide application of Artificial Intelligence (AI), more and more research has applied AI algorithms to cell energy saving.

[0003] Traditional single-cell decision-making has problems such as low resource utilization rate, signal interference, and inability to achieve seamless handover, and thus cannot adapt to the current complex communication environment. The introduction of multi-cell collaborative decision-making can improve the performance, coverage, capacity, and user experience of mobile communication networks, and also helps to reduce network operating costs and energy consumption. Multi-cell collaborative decision-making means that different cells cooperate with each other to make on / off decisions to optimize network performance and user experience. Formulating a cell energy-saving plan for collaboration between multiple types of cells can optimize the energy consumption and network performance of the entire network.

[0004] Regarding the problem of multi-cell collaborative energy saving, there has been a lot of research on establishing an optimization problem of cell sleep and resource allocation and using heuristic algorithms for iterative solution. For example: some researchers have proposed a centralized sleep scheme (CSS) to consider the performance of BS state stability, which is defined as the number of times of cell on / off state transitions. Further, in order to minimize system energy consumption and improve system state stability, a bi-objective optimization problem is established, and then a fast exhaustive search algorithm (CSSE) and a low-complexity improved particle swarm optimization algorithm (CSS-PSO) for solving the bi-objective optimization problem are proposed. Secondly, some researchers have developed a multi-objective optimization framework to reduce the overall energy consumption of the network through the cell sleep mode (SLM) in an OFDMA-based cellular system. For finding the optimal active cell set for SLM operation, a genetic algorithm is used to obtain a solution to the considered multi-objective optimization problem with a fast convergence rate.

[0005] However, when conducting large-scale co-cell modeling, the above heuristic co-cell energy-saving scheme may require a large amount of computing resources and time to find a feasible solution, and it cannot guarantee finding the optimal solution. In this regard, some researchers have applied deep reinforcement learning to the problem of cell energy saving. This is because on the one hand, reinforcement learning can dynamically adjust strategies and obtain effective strategies by continuously learning and training a large amount of data. On the other hand, reinforcement learning considers long-term rewards and can take into account the temporal correlation of the cell energy-saving problem over a period of time, and comprehensively consider the impact of cell dormancy on the network. However, the centralized co-cell dormancy algorithm based on reinforcement learning faces the non-stationarity problem in the co-cell dormancy scenario. In the communication environment, traffic shows significant fluctuations at different time periods of a day, and this phenomenon will lead to the instability of the dynamic environment of cell switching. Due to the interference of these factors, the calculation results of the reward function are also prone to large fluctuations. Such fluctuations may cause the agent to be unable to identify whether they are caused by actions or the environment, thereby further affecting the final learning effect.

[0006] To solve this problem, some researchers have proposed a method of imitation learning based on Behavior Cloning (BC). This method learns a stochastic or deterministic policy network by imitating the behavior of experts, thus avoiding the non-stationarity problem faced by the reinforcement learning algorithm in the co-cell dormancy scenario. However, this method only updates the policy network by learning the behavior of experts, does not need to learn rewards, and can only learn the optimal policy under the demonstration of experts, and cannot break through the policy demonstrated by experts.

[0007] That is to say, first, in the training process of the existing co-cell dormancy policy generation model, most use traditional deep reinforcement learning algorithms to establish a dynamic cell switching model. Due to the dynamic nature of the communication environment, the traditional deep reinforcement learning model may result in poor model training effect due to its inability to adapt to the highly dynamic environment; second, existing research introduces imitation learning to solve the non-stationarity problem in deep reinforcement learning. However, by only imitating the expert demonstration actions, it can only learn the policy under the expert demonstration and cannot exceed the policy demonstrated by experts.

[0008] Based on this, there is an urgent need to design a new training method for the co-cell dormancy policy generation model to simultaneously solve the problem of poor training effect of reinforcement learning caused by the dynamic and complex environment and the problem of limited decision-making in the imitation learning model. Summary of the Invention

[0009] In view of this, the embodiments of the present application provide a method for training a co-cell dormancy policy generation model and a co-cell dormancy method to eliminate or improve one or more defects existing in the prior art.

[0010] One aspect of the present application provides a method for training a cell cooperative sleep strategy generation model, including:

[0011] Based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the cooperative sleep of each cell, train a preset neural network to train the neural network into an expert demonstration network for predicting the expert demonstration action data corresponding to each cell in the target area;

[0012] Use the network parameters of the expert demonstration network to initialize the current neural network and the target neural network in the reinforcement learning agent, so that the reinforcement learning agent performs joint iterative training on the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the deep reinforcement learning method, to train the current neural network into a cell cooperative sleep strategy generation model for predicting the cell cooperative sleep strategy corresponding to each cell in the target area.

[0013] In some embodiments of the present application, the training of the preset neural network based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the cooperative sleep of each cell to train the neural network into an expert demonstration network for predicting the expert demonstration action data corresponding to each cell in the target area includes:

[0014] Obtain the target expert demonstration action vectors for controlling the cooperative sleep of each cell corresponding to each unit time;

[0015] Convert the target expert demonstration action vectors corresponding to each unit time into their respective target expert demonstration action indications;

[0016] Use the target expert demonstration action indications corresponding to each unit time as the labels of the historical traffic data of each cell corresponding to each unit time to be pre-obtained, to obtain the samples corresponding to each unit time and the expert demonstration action data set composed of each sample;

[0017] Divide the expert demonstration action data set into a training set and a validation set;

[0018] Based on the training set, train a preset neural network with the stochastic gradient descent algorithm, and use the validation set to verify the trained neural network, and determine the neural network that passes the verification as the current expert demonstration network for predicting the expert demonstration action data in the next unit time after the current time in the target area, and store the expert demonstration network parameters corresponding to the expert demonstration network.

[0019] In some embodiments of the present application, obtaining the target expert demonstration action vectors corresponding to each unit time for controlling the collaborative sleep of each cell includes:

[0020] Calculating the first expert demonstration action vectors corresponding to each unit time for controlling the collaborative sleep of each cell by using an expert demonstration algorithm based on heuristic search;

[0021] And calculating the second expert demonstration action vectors corresponding to each unit time for controlling the collaborative sleep of each cell by using an expert demonstration algorithm based on cell value degree ranking;

[0022] Selecting one of the power consumption values of the first expert demonstration action vectors and the second expert demonstration action vectors corresponding to each unit time as the target expert demonstration action vector corresponding to this unit time according to the power consumption values of the first expert demonstration action vectors and the second expert demonstration action vectors corresponding to each unit time.

[0023] In some embodiments of the present application, initializing the current neural network and the target neural network in the reinforcement learning agent by using the network parameters of the expert demonstration network, so that the reinforcement learning agent jointly iteratively trains the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the deep reinforcement learning method, and training the current neural network into a cell collaborative sleep policy generation model for predicting the cell collaborative sleep policies corresponding to each cell in the target area includes:

[0024] Initializing the current neural network and the target neural network in the reinforcement learning agent by using the expert demonstration network parameters;

[0025] Initializing the experience replay buffer and the maximum number of iterations corresponding to the reinforcement learning agent, where the experience replay buffer is used to store tuple samples, and the tuple samples include: current state, action, reward, and next state;

[0026] Taking the working state of each cell in a unit time as the action corresponding to this unit time, taking the predicted traffic of each cell in this unit time and the sleep state switch action in the previous unit time as the current state corresponding to this unit time, and setting the reward function corresponding to the reward; where the working state includes: the active state represented by 1 and the sleep state represented by 0;

[0027] Initializing the current neural network and the target neural network of the reinforcement learning agent according to the expert demonstration network parameters;

[0028] Select the action according to the current neural network and execute the action, and calculate the corresponding reward based on the reward function. At this time, the environmental state becomes the next state to obtain the corresponding tuple sample, and store the tuple sample in the experience replay buffer;

[0029] If the data of the tuple samples in the experience replay buffer meet a preset quantity threshold, randomly select multiple tuple samples from the experience replay buffer, so that in the current iteration round, the reinforcement learning agent, in a deep reinforcement learning manner, jointly trains the current neural network and the target neural network in the reinforcement learning agent based on each randomly selected tuple sample. Stop training until the number of iterations reaches the maximum number of iterations, and use the network parameters of the current neural network as the model parameters corresponding to the cell collaborative sleep strategy generation model for generating the cell collaborative sleep strategy prediction result within one unit time after the current time for the target area.

[0030] In some embodiments of the present application, the reward function is composed of a power consumption saving benefit function, a user service quality satisfaction function, and an action error function for imitating an expert strategy;

[0031] Among them, the value of the power consumption saving benefit function is calculated based on the total power consumption corresponding to all cells in the target area being in an active state within one unit time obtained in advance, the working state of each cell in this unit time, and a power consumption model;

[0032] The value of the user service quality satisfaction function is calculated based on the user available service rate, user demand traffic, and switching variable obtained in advance;

[0033] The value of the action error function for imitating an expert strategy is calculated based on the target expert demonstration action vector obtained in advance;

[0034] Among them, the user available service rate is calculated based on the bandwidth resources allocated to the associated users by the active cells within one unit time and the signal-to-noise ratio of the active cells to the associated users.

[0035] In some embodiments of the present application, before training a preset neural network based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the collaborative sleep of each cell, it further includes:

[0036] Collect the historical traffic data, total number of users, and total number of cells corresponding to each cell in the target area, and preprocess the historical traffic data to obtain the traffic data of each cell in the target area at each unit time;

[0037] Input the traffic data of each cell at each unit time into a preset traffic prediction model, so that the traffic prediction model outputs the traffic prediction results of each cell at one unit time after each unit time;

[0038] Sum up the traffic prediction results of each cell at one unit time after each unit time to obtain the predicted total traffic of the target area at one unit time after each unit time, and allocate the expected traffic and required traffic of each user in the target area at one unit time after each unit time according to the predicted total traffic and the total number of users.

[0039] In some embodiments of the present application, before training a preset neural network based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the collaborative sleep of each cell, it further includes:

[0040] Set a power consumption model, where the power consumption model corresponding to each cell in one unit time is composed of the fixed power consumption of the cell in this unit time and the transmission power consumption of the cell in this unit time;

[0041] Set the working state of the cell, where the working state of the cell includes: the cell is in the active state with a value of 1 in one unit time, and the cell is in the sleep state with a value of 0 in one unit time.

[0042] In some embodiments of the present application, the first expert demonstration action vector for controlling the collaborative sleep of each cell corresponding to each unit time calculated by using the expert demonstration algorithm based on heuristic search includes:

[0043] Define the working state of each cell in the target area as the position of the particle, define the total power consumption of all cells in the target area as the fitness of the particle, define the particle swarm according to various switch combinations of all cells in the target area, and define the update process of the particle.

[0044] Based on the positions of the particles, the fitness of the particles, the particle swarm, and the update process of the particles, an expert demonstration algorithm based on heuristic search is adopted to search for the optimal solutions to the corresponding cell cooperative dormancy problems for each cell in the target area, so as to obtain the first expert demonstration action vectors corresponding to each unit time for controlling the cooperative dormancy of each cell.

[0045] In some embodiments of the present application, calculating the second expert demonstration action vectors corresponding to each unit time for controlling the cooperative dormancy of each cell by using an expert demonstration algorithm based on cell value degree ranking includes:

[0046] Based on the traffic loads of each cell in each unit time obtained in advance, determine the total traffic load of the target area in each unit time;

[0047] According to the traffic loads of each cell in each unit time and the total traffic load, determine the cell value degrees of each cell in each unit time respectively;

[0048] Input the cell value degrees of each cell in each historical unit time before a unit time into a preset time series prediction model, so that the time series prediction model outputs the cell value degree prediction results of each cell in each unit time;

[0049] Solve the working states corresponding to each cell in the target area according to the cell value degree prediction results of each cell in each unit time, so as to obtain the second expert demonstration action vectors corresponding to each unit time for controlling the cooperative dormancy of each cell.

[0050] Another aspect of the present application provides a cell cooperative dormancy method, including:

[0051] Obtain the target traffic data of each cell in the target area in a historical unit time;

[0052] Input the target traffic data into a preset cell cooperative dormancy policy generation model, so that the cell cooperative dormancy policy generation model outputs the cell cooperative dormancy policy prediction result data within a unit time after the current time in the target area, where the cell cooperative dormancy policy generation model is pre-trained based on the cell cooperative dormancy policy generation model training method provided in the first aspect of the present application;

[0053] According to the cell cooperative dormancy policy prediction result data, perform corresponding dormancy state switch actions for each cell in the target area.

[0054] The third aspect of the present application provides a training device for a cell collaborative sleep strategy generation model, including:

[0055] A model training module based on expert demonstrations, configured to train a preset neural network based on the historical traffic data of each cell in the corresponding target area for each unit time and the expert demonstration action data for controlling the collaborative sleep of each of the cells, so as to train the neural network into an expert demonstration network for predicting the expert demonstration action data corresponding to each of the cells in the target area;

[0056] A model migration and reinforcement learning module, configured to initialize the current neural network and the target neural network in the reinforcement learning agent with the network parameters of the expert demonstration network, so that the reinforcement learning agent performs joint iterative training on the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the deep reinforcement learning method, so as to train the current neural network into a cell collaborative sleep strategy generation model for predicting the cell collaborative sleep strategy corresponding to each of the cells in the target area.

[0057] The fourth aspect of the present application provides a cell collaborative sleep device, including:

[0058] A traffic acquisition module, configured to acquire the target traffic data of each cell in the target area within a historical unit time;

[0059] A cell collaborative sleep strategy prediction module, configured to input the target traffic data into a preset cell collaborative sleep strategy generation model, so that the cell collaborative sleep strategy generation model outputs the cell collaborative sleep strategy prediction result data within a unit time after the current time for the target area, where the cell collaborative sleep strategy generation model is pre-trained based on the cell collaborative sleep strategy generation model training method provided in the first aspect of the present application;

[0060] A cell collaborative sleep control module, configured to perform corresponding sleep state switching actions on each of the cells in the target area according to the cell collaborative sleep strategy prediction result data.

[0061] The fifth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the cell collaborative sleep strategy generation model training method and / or the cell collaborative sleep method described above.

[0062] The sixth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned method for training a cell cooperative sleep policy generation model and / or the above-mentioned cell cooperative sleep method.

[0063] The seventh aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for training a cell cooperative sleep policy generation model and / or the above-mentioned cell cooperative sleep method.

[0064] The method for training a cell cooperative sleep policy generation model provided by the present application is based on the historical traffic data of each cell in the corresponding target area for each unit time and the expert demonstration action data for controlling the cooperative sleep of each cell, and trains a preset neural network to train the neural network into an expert demonstration network for predicting the expert demonstration action data corresponding to each cell in the target area; migrates the expert demonstration network to the current neural network and the target neural network in the reinforcement learning agent, so that the reinforcement learning agent jointly iteratively trains the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the deep reinforcement learning method, so as to train the current neural network in the reinforcement learning agent into a cell cooperative sleep policy generation model for predicting the cell cooperative sleep policy corresponding to each cell in the target area; considering the dynamics and complexity of the communication environment, the training method provided by the present application introduces the idea of transfer learning. First, the network state and expert demonstration actions are input into the neural network for training, and then the trained network model parameters are transferred to the current neural network and the target neural network in the deep reinforcement learning, so that the reinforcement learning agent can continuously update and optimize the policy based on the expert demonstration results, generate a cell switching policy better than the expert demonstration, and thus can simultaneously solve the problem of poor training effect of reinforcement learning caused by the dynamics and complexity of the environment and the problem of limited decision-making in the imitation learning model, can effectively improve the training effectiveness and reliability of the cell cooperative sleep policy generation model, and can solve the problem of limited decision-making in the imitation learning, and thus can effectively improve the application effectiveness and reliability of the cell cooperative sleep policy generated based on the cell cooperative sleep policy generation model to achieve dynamic and adaptive cell switching decisions.

[0065] The additional advantages, objects, and features of the present application will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following parts, or can be learned from the practice of the present application. The objects and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification and the drawings.

[0066] Those skilled in the art will understand that the objectives and advantages that can be achieved with the present application are not limited to those specifically described above, and the above and other objectives that can be achieved with the present application will be more clearly understood according to the following detailed description. Description of the Drawings

[0067] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and do not limit the present application. The components in the drawings are not drawn to scale, but are only for showing the principles of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:

[0068] Figure 1 It is a first schematic flowchart of the method for training a cell collaborative sleep strategy generation model in an embodiment of the present application.

[0069] Figure 2 It is a second schematic flowchart of the method for training a cell collaborative sleep strategy generation model in an embodiment of the present application.

[0070] Figure 3 It is a third schematic flowchart of the method for training a cell collaborative sleep strategy generation model in an embodiment of the present application.

[0071] Figure 4 It is a fourth schematic flowchart of the method for training a cell collaborative sleep strategy generation model in an embodiment of the present application.

[0072] Figure 5 It is a schematic flowchart of the expert demonstration algorithm based on heuristic search provided by Application Example 3 of the present application.

[0073] Figure 6 It is a schematic flowchart of the expert demonstration algorithm based on cell value degree ranking provided by Application Example 4 of the present application.

[0074] Figure 7 It is a schematic flowchart of the cell collaborative sleep method in an embodiment of the present application.

[0075] Figure 8 It is a schematic structural diagram of the cell collaborative sleep strategy generation model training device in an embodiment of the present application.

[0076] Figure 9 It is a schematic structural diagram of the cell collaborative sleep device in an embodiment of the present application. Detailed Embodiments

[0077] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further elaborates on this application in combination with the embodiments and accompanying drawings. Herein, the illustrative embodiments of this application and their descriptions are used to explain this application, but do not limit this application.

[0078] Herein, it should also be noted that in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution of this application are shown in the drawings, while other details less relevant to this application are omitted.

[0079] It should be emphasized that the term "including / comprising" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0080] Herein, it should also be noted that if not otherwise specified, the term "connection" in this application can not only refer to direct connection, but also represent indirect connection with intermediaries.

[0081] In the following, embodiments of this application will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0082] In order to simultaneously solve the problems of poor training effect of reinforcement learning caused by the dynamics and complexity of the environment and the limited decision-making in the imitation learning model, embodiments of this application respectively provide a method for training a cell collaborative sleep strategy generation model, a cell collaborative sleep method, a cell collaborative sleep strategy generation model training device for executing the method for training a cell collaborative sleep strategy generation model, a cell collaborative sleep device for executing the cell collaborative sleep method, an electronic device, a computer-readable storage medium, and a computer program product, which can realize dynamic and adaptive cell switching decisions, can improve the problem of poor training effect in traditional reinforcement learning due to high environmental complexity, and can solve the problem of limited decision-making in imitation learning.

[0083] Specific details are described in detail through the following embodiments.

[0084] Based on this, embodiments of this application provide a method for training a cell collaborative sleep strategy generation model that can be implemented by a cell collaborative sleep strategy generation model training device. Refer to Figure 1 , the method for training a cell collaborative sleep strategy generation model specifically includes the following content:

[0085] Step 100: Based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the collaborative sleep of each of the cells, train a preset neural network to train the neural network into an expert demonstration network for predicting the expert demonstration action data corresponding to each of the cells in the target area.

[0086] In one or more embodiments of the present application, each unit time refers to a time period with the same duration and continuity.

[0087] It can be understood that, in order to enable the reinforcement learning agent (agent, which can be abbreviated as agent) to adapt to the dynamics of the environment, Step 100 introduces expert demonstrations, enabling the agent to continuously optimize the switching strategy according to the actions of the expert. Considering that only imitating the expert actions will cause the agent to be unable to generate strategies that exceed or equal the expert actions, to solve this problem, in this embodiment, the expert demonstration actions are first input into the neural network for training before the reinforcement learning algorithm training to obtain expert demonstration parameters, and then the following Step 200 can be used to migrate the trained expert demonstration network to the current neural network and the target neural network in reinforcement learning, specifically referring to: migrating the model parameters (which can also be called expert demonstration parameters) corresponding to the trained expert demonstration network to the current neural network and the target neural network in reinforcement learning.

[0088] Step 200: Initialize the current neural network and the target neural network in the reinforcement learning agent with the network parameters of the expert demonstration network, so that the reinforcement learning agent performs joint iterative training on the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the deep reinforcement learning method, to train the current neural network into a cell collaborative sleep strategy generation model for predicting the cell collaborative sleep strategy corresponding to each of the cells in the target area.

[0089] In Step 200, the current neural network generates a strategy through training expert demonstrations, and through joint training with the target neural network, continuously iterates to obtain the final cell collaborative sleep strategy. For the reinforcement learning part, this embodiment adopts the Deep Q Network (DQN) algorithm, and determines the model parameters in the finally converged current neural network as the cell collaborative sleep strategy generation model finally used for online applications.

[0090] In one or more embodiments of the present application, the expert demonstration action data for controlling the collaborative dormancy of each cell refers to the action of switching the dormancy state switch corresponding to each cell in the target area during a unit time demonstrated by an expert; the cell collaborative dormancy strategy includes performing corresponding dormancy state switch actions for each cell in the target area within a unit time after the current time; wherein, the dormancy state switch action includes: an action to turn on the dormancy state of the cell and an action to turn off the dormancy state of the cell.

[0091] As can be seen from the above description, the method for training a cell collaborative dormancy strategy generation model provided by the embodiments of the present application can effectively improve the training effectiveness and reliability of the cell collaborative dormancy strategy generation model, and can solve the problem of limited decision-making in imitation learning. Furthermore, it can effectively improve the application effectiveness and reliability of the cell collaborative dormancy strategy generated based on the cell collaborative dormancy strategy generation model, so as to achieve dynamic and adaptive cell switching decisions.

[0092] In order to further improve the effectiveness and reliability of the process of training an expert demonstration network, in a method for training a cell collaborative dormancy strategy generation model provided by the embodiments of the present application, refer to Figure 2 , step 100 in the method for training a cell collaborative dormancy strategy generation model specifically includes the following content:

[0093] Step 110: Obtain the target expert demonstration action vectors corresponding to each unit time for controlling the collaborative dormancy of each cell.

[0094] It can be understood that the target expert demonstration action vector can also be referred to as the optimal expert demonstration action vector.

[0095] Step 120: Convert the target expert demonstration action vectors corresponding to each unit time into corresponding target expert demonstration action instructions respectively.

[0096] Step 130: Use the target expert demonstration action instructions corresponding to each unit time as the labels for the historical traffic data of each cell corresponding to each unit time to be pre-obtained, so as to obtain the samples corresponding to each unit time and the expert demonstration action data set composed of each sample.

[0097] Step 140: Divide the expert demonstration action data set into a training set and a validation set.

[0098] Step 150: Based on the training set, train a preset neural network using the stochastic gradient descent algorithm, and use the validation set to validate the trained neural network. Determine the expert demonstration network that has passed the validation as the current expert demonstration network for predicting the expert demonstration action data of the target area within one unit of time after the current time, and store the expert demonstration network parameters corresponding to the expert demonstration network.

[0099] In the transfer reinforcement learning method based on expert demonstrations, in order to obtain expert demonstration actions, this application proposes two centralized expert demonstration algorithms, namely the expert demonstration algorithm based on heuristic search and the expert demonstration algorithm based on cell value degree ranking. These two algorithms can turn off cells to the greatest extent, and the second algorithm can greatly reduce the time complexity. Correspondingly, in a method for training a cell collaborative sleep strategy generation model provided in an embodiment of this application, see Figure 3 , step 110 in the method for training the cell collaborative sleep strategy generation model specifically includes the following content:

[0100] Step 111: Use the expert demonstration algorithm based on heuristic search to calculate the first expert demonstration action vector corresponding to each unit of time for controlling the collaborative sleep of each cell.

[0101] And, step 112: Use the expert demonstration algorithm based on cell value degree ranking to calculate the second expert demonstration action vector corresponding to each unit of time for controlling the collaborative sleep of each cell.

[0102] It can be understood that both the first expert demonstration action vector and the second expert demonstration action vector are expert demonstration action vectors. The difference in their expressions is only used to distinguish the result data calculated by the expert demonstration algorithm based on heuristic search or the result data calculated by the expert demonstration algorithm based on cell value degree ranking.

[0103] Step 113: According to the power consumption values of the first expert demonstration action vector and the second expert demonstration action vector corresponding to each unit of time, select one of the power consumption values of the first expert demonstration action vector and the second expert demonstration action vector corresponding to each unit of time as the target expert demonstration action vector corresponding to that unit of time.

[0104] To further improve the effectiveness and reliability of the reinforcement learning process, in a method for training a cell collaborative sleep strategy generation model provided in an embodiment of this application, see Figure 2 or Figure 3 , step 200 in the method for training the cell collaborative sleep strategy generation model specifically includes the following content:

[0105] Step 210: Initialize the current neural network and the target neural network in the reinforcement learning agent using the expert demonstration network parameters.

[0106] Step 220: Initialize the experience replay buffer and the maximum number of iterations corresponding to the reinforcement learning agent, where the experience replay buffer is used to store tuple samples, and the tuple samples include: the current state, action, reward, and the next state.

[0107] Step 230: Take the working state of each cell in a unit time as the action corresponding to this unit time, take the predicted traffic of each cell in this unit time and the sleep state switch action in the previous unit time as the current state corresponding to this unit time, and set the reward function corresponding to the reward; where the working state includes: the active state represented by 1 and the sleep state represented by 0.

[0108] Step 240: Initialize the current neural network and the target neural network of the reinforcement learning agent according to the expert demonstration network parameters.

[0109] Step 250: Select and execute the action according to the current neural network, calculate the corresponding reward based on the reward function, and at this time the environmental state becomes the next state to obtain the corresponding tuple sample, and store the tuple sample in the experience replay buffer.

[0110] Step 260: If the data of the tuple samples in the experience replay buffer meets a preset quantity threshold, randomly select multiple tuple samples from the experience replay buffer, so that in the current iteration round, the reinforcement learning agent, in a deep reinforcement learning manner, jointly trains the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the randomly selected tuple samples. When the number of iterations reaches the maximum number of iterations, stop the training, and use the network parameters of the current neural network as the model parameters corresponding to the cell collaborative sleep strategy generation model for generating the prediction result of the cell collaborative sleep strategy in a unit time after the current time in the target area.

[0111] In order to further improve the application effectiveness and reliability of the reward function, in a method for training a cell collaborative sleep strategy generation model provided in an embodiment of the present application, the reward function in the method for training the cell collaborative sleep strategy generation model is composed of a power consumption saving benefit function, a user service quality satisfaction function, and an action error function for imitating the expert strategy;

[0112] Among them, the value of the power consumption saving benefit function is calculated based on the total power consumption corresponding to all cells in the target area being active in a unit time obtained in advance, the working state of each cell in this unit time, and a power consumption model;

[0113] The value of the user's quality of service satisfaction function is calculated based on the user available service rate, user demand traffic, and switching variable obtained in advance;

[0114] The value of the action error function imitating the expert strategy is calculated based on the target expert demonstration action vector obtained in advance;

[0115] Among them, the user available service rate is calculated based on the bandwidth resources allocated to the associated users by the active cells in a unit time and the signal-to-noise ratio from the active cells to the associated users.

[0116] In order to further improve the calculation effectiveness and reliability of the quality of service satisfaction, in a method for training a cell cooperative sleep strategy generation model provided in an embodiment of the present application, see Figure 4 Before step 100 in the method for training the cell cooperative sleep strategy generation model, the following specific content is further included:

[0117] Step 010: Collect the historical traffic data, total number of users, and total number of cells corresponding to each cell in the target area, and preprocess the historical traffic data to obtain the traffic data of each cell in the target area at each unit time.

[0118] Step 020: Input the traffic data of each cell at each unit time into a preset traffic prediction model, so that the traffic prediction model outputs the traffic prediction results of each cell at one unit time after each unit time.

[0119] Step 030: Sum up the traffic prediction results of each cell at one unit time after each unit time to obtain the predicted total traffic of the target area at one unit time after each unit time, and allocate the user expected traffic and user demand traffic corresponding to each user in the target area at one unit time after each unit time according to the predicted total traffic and the total number of users.

[0120] In order to further improve the calculation effectiveness and reliability of the power consumption saving benefit, in a method for training a cell cooperative sleep strategy generation model provided in an embodiment of the present application, see Figure 4 Before step 100 and after step 030 in the method for training the cell cooperative sleep strategy generation model, the following specific content is further included:

[0121] Step 040: Set the power consumption model. Among them, the power consumption model corresponding to each cell within a unit time consists of the fixed power consumption of the cell within this unit time and the transmission power consumption of the cell within this unit time.

[0122] Step 050: Set the cell working state. The cell working state includes: when the cell is in the active state within a unit time, it is a value of 1, and when the cell is in the sleep state within a unit time, it is a value of 0.

[0123] In order to further improve the application effectiveness and reliability of the expert demonstration algorithm based on heuristic search, in a method for training a cell collaborative sleep strategy generation model provided in an embodiment of the present application, step 111 in the method for training the cell collaborative sleep strategy generation model specifically includes the following contents:

[0124] 1) Define the working state of each cell within the target area as the position of the particle, define the total power consumption of all cells within the target area as the fitness of the particle, define the particle swarm according to various switch combinations for all cells within the target area, and define the update process of the particle.

[0125] 2) Based on the position of the particle, the fitness of the particle, the particle swarm, and the update process of the particle, adopt the expert demonstration algorithm based on heuristic search to search for the optimal solution to the corresponding cell collaborative sleep problem for each cell within the target area, so as to obtain the first expert demonstration action vector for controlling the collaborative sleep of each cell corresponding to each unit time.

[0126] In order to further improve the application effectiveness and reliability of the expert demonstration algorithm based on cell value degree ranking, in a method for training a cell collaborative sleep strategy generation model provided in an embodiment of the present application, step 112 in the method for training the cell collaborative sleep strategy generation model specifically includes the following contents:

[0127] 1) Based on the traffic load of each cell within each unit time obtained in advance, determine the total traffic load of the target area within each unit time.

[0128] 2) According to the traffic load of each cell within each unit time and the total traffic load, determine the cell value degree of each cell within each unit time respectively.

[0129] 3) Input the cell value degrees of each cell within each historical unit time before a unit time for each cell into a preset time series prediction model, so that the time series prediction model outputs the cell value degree prediction result of each cell within each unit time.

[0130] 4) Solve the working state corresponding to each of the cells in the target area according to the predicted results of the cell value degrees of each of the cells in each unit time, so as to obtain a second expert demonstration action vector corresponding to each of the unit times for controlling the collaborative sleep of each of the cells.

[0131] To further illustrate the cell collaborative sleep strategy generation model training method provided in the above embodiments, the present application also provides the following four application examples respectively.

[0132] (1) Application Example 1

[0133] Application Example 1 provides a cell collaborative sleep strategy generation model training method based on transfer reinforcement learning with expert demonstrations, which is used to explain the embodiments of steps 100 to 200, the embodiments of steps 111 to 113, the embodiments of steps 210 to 260, the embodiments of the reward function, the embodiments of steps 010 to 030, and the embodiments of steps 040 to 050 described above. Specifically, Application Example 1 includes the following content:

[0134] Step 1-1: Cell traffic data collection. Collect the historical traffic data of the cell (such as the uplink and downlink traffic of the air interface, the uplink and downlink traffic of the Packet Data Convergence Protocol (PDCP) layer, etc.) through the network management or data collection system, and perform preprocessing such as cleaning, construction, aggregation, and screening on the collected data, such as removing outliers and filling in missing values.

[0135] Step 1-2: Cell traffic prediction. Based on the historical cell traffic obtained in Step 1-1, use machine learning algorithms to predict the future cell traffic. Since the switch decision of the cell is obtained for the cell traffic in the future time period, it is necessary to perform time series prediction on the historical cell traffic collected in Step 1-1. For the time series prediction of cell traffic, traffic prediction models such as the triple exponential smoothing algorithm (Holt-winters) and the long short-term memory network (LSTM) can be used. Assume that a target area contains M cells and N users, and the cell set and user set can be expressed as and where m represents the m-th cell in the target area and n represents the n-th user in the target area. The traffic of all M cells in the t-th unit time can be expressed as where represents the traffic of the m-th cell in the t-th unit time. And in order to predict the cell traffic of cell m in the t-th unit time, the historical traffic sequence of l consecutive unit times before the t-th unit time will be Input the aforementioned traffic prediction model. The traffic prediction model outputs the traffic prediction result at the t-th unit time, defined as This application example can further obtain the predicted traffic sequence of cell m within the next T time

[0136] Step 1-3: User traffic acquisition. By performing traffic prediction on all cells within a certain target area, the predicted traffic of all cells at the t-th unit time can be obtained, that is At the t-th unit time, first sum up the predicted traffic of all cells within the target area to obtain the predicted total traffic within the target area at the next t-th unit time, that is Then, on the premise of knowing the total traffic of the target area, simulate the traffic of each user. Assume that the traffic of all users follows a Poisson distribution. The specific calculation process is as follows:

[0137] Step 1-31: User expected traffic allocation. From Step 1-1, the total number of users within the entire target area can be obtained as N, and the traffic of each user follows a Poisson distribution. Assume that the expected value of the traffic of each user is Allocate the predicted total traffic of the target area To each user. A common method is uniform allocation. The expected traffic of user n can be obtained from formula (1):

[0138]

[0139] If specific user preferences are available, weighted allocation can be performed based on this information.

[0140] Step 1-32: User required traffic acquisition. For each user n, since the Poisson distribution describes the number of events occurring within a fixed time interval, the required traffic of each user n is a Poisson random variable with a parameter of and can be obtained from formula (2):

[0141]

[0142] Finally, the required traffic of all users within the target area can be expressed as

[0143] Steps 1 - 4: Fitting the cell power consumption model. In a wireless communication network, the power consumption of a cell is not only proportional to the traffic load within its coverage area. As a major power consumer in the communication network, the power consumption of a cell is mainly divided into fixed power consumption and transmission power consumption. Among them, the fixed power consumption is related to hardware devices such as refrigeration equipment and cell circuits, and usually does not change much. The transmission power consumption is related to power amplifiers, radio frequency units, etc., and it is related to the transmission load of the cell. In this application example, a general power consumption model is fitted by collecting the power consumption values and traffic loads of the cell at different times. The power consumption model of cell m at the t - th unit time is shown in formula (3):

[0144]

[0145] where P fix represents the fixed power consumption, is the transmission power consumption of cell m, which is non - linearly related to the traffic load of cell m. The formula (3) can be obtained by fitting the relationship between the cell power consumption and the traffic load.

[0146] Steps 1 - 5: Defining the working state of the cell and the association state between the user and the cell. Based on the power consumption model of the cell in the active state obtained in Steps 1 - 4, the power consumption of the cell under different traffic loads can be obtained. Since the power consumption of the cell is huge, this application example aims to reduce the power consumption of the cell by switching low - load cells to the sleep mode. To ensure that each user can be served by a cell in the communication network, a 0 - 1 switching variable is introduced to represent whether user n is associated with cell m at the t - th unit time. If user n is associated with cell m at the t - th unit time, then Conversely, can be expressed as

[0147] Steps 1 - 6: Cell collaborative sleep. To reduce the cell power consumption and improve the network efficiency, this application example adjusts some cells with low traffic loads to the sleep state to reduce unnecessary power consumption overhead. Based on the cell power consumption model and the cell working state defined in the above Steps 1 - 4 to 1 - 5, this application example uses a deep reinforcement learning model to solve the cell collaborative sleep problem. To enhance the adaptability of the agent to the environment and improve the training effect, this application example proposes a cell collaborative sleep method based on transfer reinforcement learning with expert demonstrations, including the following Steps 1 - 61 to Steps 1 - 62:

[0148] Steps 1 - 61: First, convert the cell collaborative sleep optimization problem into a Markov Decision Process (MDP), and define the variables as follows:

[0149] 1) Action: The action determines the working state (active / sleep) of all M cells in the t-th unit time, which can be written as where is 1 when cell m is in the active state in the t-th unit time, and 0 otherwise.

[0150] 2) State: The system state consists of the predicted traffic of all M cells in the entire target area in the t-th unit time and the cell switching actions in the (t - 1)-th unit time, denoted as The predicted traffic of all N users in the entire target area in the t-th unit time is obtained from steps 1 - 3.

[0151] 3) Reward function: To evaluate the current switching strategy of the cell, the instantaneous reward obtained by taking action a t in state s t can be used as the basis for judgment. To ensure that the QoS of users is not reduced while reducing the power consumption of the cell and achieving a balance between power consumption and user service quality, the reward function can be R(s t , a t ) consists of the benefit c p (s t , a t ) of power consumption savings, the satisfaction c s (s t , a t ) of the user's quality of service (QoS), and the action error c exp (s t , a t ) of imitating the expert strategy. The reward function can be obtained through formula (4), where λ, μ, v are the weight parameter values of each part of the reward function:

[0152] R(s t , a t ) = λ·c p (s t , a t ) + μ·c s (s t , a t ) + v·c exp (s t , a t ) (4)

[0153] (3a) The benefit c p (s t , a t ) of power consumption savings in the reward function is shown in formula (5): ​

[0154]

[0155] Among them, represents the total power consumption when all M cells are in the active state at the t-th unit time. Equation (5) represents the benefit of the overall power consumption reduction in the target area due to the sleep strategy adopted by the cells.

[0156] (3b) The satisfaction degree c of the user's QoS in the reward function s (s t , a t ) is shown in Equation (6):

[0157]

[0158] The user's QoS satisfaction degree is determined by the service rate that the user can obtain and the ratio of the user's required service volume. Among them, the service rate that the user can obtain (i.e., the user's available service rate) is the service rate that user n can obtain by associating with active cell m at the t-th unit time; the user's required service volume can be obtained from Steps 1 - 32.

[0159] Among them, indicates that the service rate that user n can obtain by associating with active cell m at the t-th unit time can be calculated using the Shannon formula, as shown in Equation (7):

[0160]

[0161] Among them, refers to the bandwidth resource allocated by active cell m to the associated user n at the t-th unit time. is the signal-to-noise ratio from active cell m to the associated user n at the t-th unit time, denoted as is the power resource allocated by active cell m to the associated user n at the t-th unit time, is the channel gain from active cell m to the associated user n at the t-th unit time, and N0 is the Gaussian noise power spectral density.

[0162] (3c) The action error c of imitating the expert strategy in the reward function exp (s t , a t ), and the specific calculation is introduced in Steps 1 - 62.

[0163] Step 1-62: Transfer Reinforcement Learning Algorithm Based on Expert Demonstration. Based on the Markov decision process defined in Step 1-61, this application example proposes a transfer reinforcement learning algorithm based on expert demonstration to solve the problem of cell collaborative sleep. Specifically, in order to enable the agent to adapt to the dynamics of the environment, this application example introduces expert demonstration, enabling the agent to continuously optimize the switching strategy according to the actions of the expert. Considering that only imitating the expert actions will cause the agent to be unable to generate a strategy that exceeds or equals the expert actions, to solve this problem, this application example first inputs the expert demonstration actions into the neural network for training before training the reinforcement learning algorithm to obtain the expert demonstration parameters, and then transfers the trained expert demonstration parameters of the neural network to the current neural network and the target neural network in the reinforcement learning agent. The current neural network is jointly trained with the target neural network and continuously iterated to obtain the final cell collaborative sleep strategy. For the reinforcement learning part, this application example adopts the Deep Q Network (DQN) algorithm. The process of the transfer reinforcement learning algorithm based on expert demonstration is as follows:

[0164] 1) Acquisition of expert demonstration actions. The expert demonstration action vector of heuristic search is obtained by adopting the algorithm of Application Example 3 The expert demonstration action vector based on and sorted by cell value degree is obtained by adopting Application Example 4

[0165] 2) Construction of the optimal expert demonstration action set. The optimal expert demonstration action set wherein, is the optimal expert demonstration action vector at the t-th unit time (i.e., the target expert demonstration action vector). is the expert demonstration action vector generated according to Application Example 3 and Application Example 4 and By calculating and comparing the power consumption values of the two, the lower one is selected and denoted as

[0166] 3) Calculation of the action error of imitating the expert strategy. The influence of the expert demonstration action on the current action is that for the actions of the same cell at the same moment, if they are the same, a positive reward is given, and if they are different, a negative penalty is given. The process of reward and penalty is shown in formula (8):

[0167]

[0168] 4) Training of the neural network model based on expert demonstration. The optimal expert demonstration action set is used to train the neural network ED of expert demonstration to obtain the trained neural network parameters θ EDThe neural network training process based on expert demonstration is shown in Application Example 2 for details.

[0169] Steps 1-7: Initialize and train DQN training parameters, including the following steps 1-71 to steps 1-74:

[0170] Step 1-71: Initialize the experience replay buffer R exp and the maximum number of iterations E. Among them, R exp stores the current state a t , action s t , reward R(s t ,a t ) = {r 1 ,r 2 ,...,r t ...,r T} and the next state a t+1 and other data.

[0171] Step 1-72: Initialize the neural network. The DQN algorithm usually uses two neural networks, namely the current neural network Q w (s t ,a t ; θ) and the target neural network Q w′ (s t ,a t ; θ′). Here, use the trained neural network parameters θ ED to initialize the current neural network and the target neural network.

[0172] Step 1-73: Update the Q value of the neural network. Update the Q value using the Bellman equation, that is, calculate the mean square error between the target Q value and the current Q value, and update the weights of the neural network through gradient descent, specifically including the following steps (a) to step (e):

[0173] Step (a), according to the current neural network Q w (s t ,a t ; θ) use the ε-greedy strategy to select the action a t ; Execute the action a t , obtain the reward R(s t ,a t ), and at this time the environmental state becomes s t+1 .

[0174] Step (b), store the above data in the form of (s t ,s t ,r t ,s t+1 ) into the experience replay buffer R exp .

[0175] Step (c), when the data in R exp is sufficient, randomly collect B data from it, that is, obtain {(s i , a i , r i , s i+1 )} i=1,2,...,B .

[0176] Step (d), use the target neural network to calculate the collected data, and the calculated value is shown in formula (9):

[0177] y i = r i + γmaxQ w′ (s i+1 , a i ; θ′) (9)

[0178] Step (e), calculate the target loss value, and use the gradient descent algorithm to complete the update of the current neural network parameter θ. The loss value is shown in formula (10):

[0179]

[0180] Step 1-74: Update the target neural network parameter θ′. In this stage, it is necessary to update the parameter θ′ of the target neural network Q w′ (s t , a t ; θ′). Determine whether the target neural network parameter θ′ needs to be updated at the current moment t. If it is necessary, update the target neural network Q w′ (s t , a t ; θ′), θ′ = θ.

[0181] Step 1-75: Repeat the above process until the maximum number of iterations E is reached and the model converges.

[0182] (2) Application Example 2

[0183] Application Example 2 provides a training process of an expert demonstration network for explaining the embodiments of the foregoing Steps 110 to 150. Specifically, Application Example 2 includes the following content:

[0184] Based on Application Example 1, this application example proposes a neural network training method based on expert demonstration to obtain the expert demonstration parameter θ ED . Specifically, the training process of the expert demonstration network is as follows:

[0185] Step 2-1: Optimal expert demonstration action dataset The dataset has a total of T samples, and the samples Among them, the sample d t The attribute is the traffic volume of M cells at the t-th unit time The sample d t The label is the optimal expert demonstration action instruction (i.e., the target expert demonstration action instruction) y t 。Write the action vector at the t-th unit time in the set of actions demonstrated by the optimal expert of each bit (i.e., the switch state 0 or 1 of the target area cell) as each bit of a binary number, and convert this binary number into a decimal number and write it as y t , that is

[0186] Step 2-2: Divide the data set into a training set and a test set. The training set is used to train the parameters of the expert demonstration network ED. The expert demonstration network ED adjusts its own parameters by learning the sample data in the training set to minimize the loss function. The test set is used to evaluate the performance of the expert demonstration network ED during the training process.

[0187] Step 2-3: Randomly initialize the model parameters θ of the expert demonstration network ED ED_ini , initialize the maximum number of learning rounds E ED .

[0188] Step 2-4: Train the expert demonstration network ED using the stochastic gradient descent algorithm, which specifically includes Step 2-41 and Step 2-42:

[0189] Step 2-41: Randomly sample a batch of data as a sample set. The batch refers to a certain number of randomly selected data samples. The size of this batch of data is N d , d t is an empirical sample in the sampled batch data D.

[0190] Step 2-42: Update the parameters θ of the expert demonstration network ED , which specifically includes Step 2-42(a) and Step 2-42(b):

[0191] Step 2-42(a), the loss function of the expert demonstration network parameters θ ED is shown in formula (11):

[0192]

[0193] Step 2-42(b), using the stochastic gradient descent algorithm, update the network parameters θ of the expert demonstration network ED , and the update process is shown in formula (12):

[0194]

[0195] where l ED is the learning rate of the expert demonstration network, and l ED ∈ [0, 1].

[0196] Step 2-5: Check the loss value and prediction accuracy of the expert demonstration network on the validation set.

[0197] Step 2-6: Repeat the above process until the maximum number of learning rounds E ED is reached, the model converges, and the model parameters θ ED are saved.

[0198] (III) Application Example 3

[0199] Application Example 3 provides a heuristic search-based expert demonstration algorithm for explaining the embodiments of calculating the first expert demonstration action vectors corresponding to each of the unit times for controlling the cooperative dormancy of each cell by using the heuristic search-based expert demonstration algorithm. Specifically, Application Example 3 includes the following content:

[0200] As can be seen from Application Example 1, this application uses a transfer reinforcement learning algorithm based on expert demonstration to solve the problem of cell cooperative dormancy. The algorithm enables the intelligent agent to make strategies that exceed expert demonstrations by learning the neural network training parameters based on expert demonstrations. To obtain the expert demonstration strategy, this application proposes two expert demonstration algorithms, namely, the heuristic search-based expert demonstration algorithm (Application Example 3) and the expert demonstration algorithm based on cell value degree ranking (Application Example 4).

[0201] This application example details the heuristic search-based expert demonstration algorithm. The heuristic search refers to the heuristic search (PSO-GA) combining the Particle Swarm Optimization (PSO) and the Genetic Algorithm (GA). In the traditional particle swarm optimization algorithm, the update processes of the particle's position and velocity are continuous. However, the cell cooperative dormancy problem is a solution process of integer programming binary variables. Therefore, in this application example, based on the traditional particle swarm optimization algorithm, the idea of the genetic algorithm is borrowed to improve the update process of the particles.

[0202] The flowchart of the heuristic search-based expert demonstration algorithm is as Figure 5 shown.

[0203] Step 3-1: Define the particle p k . In the traditional particle swarm algorithm, the particle p k has three attributes Among them is the position of particle p k , representing a feasible solution to the optimization problem is the fitness value of particle p k , representing the quality of this feasible solution. In the heuristic search expert demonstration algorithm, for the problem of cell collaborative sleep, in this application example, the attributes of particle p k are redefined as follows:

[0204] Step 3-11: Redefine the position of particle p k represents the working states of M cells in the overall target area. As can be seen from Application Example 1, there are M cells in the overall target area, and the cell set is represented as The position of particle p can be expressed as k where x refers to the working state of cell m. x k,m is a 0-1 integer variable. If cell m is placed in the sleep state, then x k,m = 0, otherwise, x k,m = 1 k,m .

[0205] Step 3-12: Redefine the fitness of particle p k as: the overall energy consumption of all cells in the target area. According to the working states of M cells in the target area indicated by the position k in particle p , the total power consumption of this target area can be expressed as:

[0206]

[0207] Step 3-2: Define the particle swarm P. This algorithm initializes the solutions to the problem of K-cell collaborative sleep, that is, K particles, equivalent to K switch combinations of all M cells. The particle swarm P = {p1, p2,..., p K} composed of K particles is used to search for the optimal solution to the problem of cell collaborative sleep

[0208] Step 3-3: Define the update process of particle p k in the heuristic search expert demonstration algorithm. The heuristic search expert demonstration algorithm proposed in this application example modifies the update process of particles on the basis of the traditional particle swarm optimization algorithm. In this application example, the idea of the genetic algorithm is introduced, and the variable x indicating the working state of a certain cell m in the position of the particle k,m is regarded as the gene locus in the genetic algorithm, and the mutation operation and the crossover operation are introduced into the position update process of the particle. The specific description is as follows:

[0209] 1) Selection: The selection operation refers to particle p k by randomly selecting gene positions from the historical optimal solution and the historical optimal solution of the particle swarm to generate the coordinates of particle p for the next iteration. Here, ε1 and ε2 represent the number of gene positions selected from the historical optimal solution k and the historical optimal solution of the population respectively.

[0210] 2) Crossover: The crossover operation generates the coordinates of particle p for the next iteration by combining the historical optimal solution and the historical optimal solution of the population . k

[0211] 3) Mutation: To avoid the local optimum problem, this application example applies the mutation operation to update the remaining gene positions of particle p k position , that is, randomly change the gene values of the remaining gene positions of particle p k position . λ represents the number of mutated gene positions in particle p k .

[0212] Step 3-4: Search for the optimal solution to the small cell cooperative sleep problem according to the heuristic search expert demonstration algorithm. Taking the solution to the small cell cooperative sleep problem at the t-th unit time as an example, the specific algorithm flow is as follows:

[0213] Step 3-41: Initialize variables. Initialize the position of particle p k (i.e., the switch action), initialize the crossover factors ε1, ε2, and the maximum mutation factor λ , initialize the mutation factor λ = λ max , initialize the number of iterations N ini . r

[0214] Step 3-42: User association. According to the initial position of particle p k indicated, all users in the target area access the nearest active cell. The process of user-cell association is detailed in Step 62 of Application Example 1.

[0215] Step 3-43: Determine the validity of the initial particle. Judge whether the active cells determined by the initial position of particle p k can meet the service rate of users, that is, whether it satisfies ​​​​​ The service rate of the active cell m' to the user n For the definition, please refer to formula (7) in Application Example 1. If the service rate of a user cannot be satisfied, the position of this particle cannot be used as the solution to the cell cooperative dormancy problem.

[0216] Step 3-44: Initialize the optimal solution of the particle. For According to the initial position k of the particle p and the working states of all M cells indicated thereby, as well as the cell power consumption formula (3), calculate the initial fitness value of this particle and set the current position as the optimal solution of the particle p k , that is

[0217] Step 3-45: Initialize the optimal solution of the particle swarm. For the particle swarm P, initialize the optimal solution of the particle swarm

[0218] Step 3-46: Update the particle position. In the r-th iteration process, for respectively select ε1 and ε2 gene positions from the optimal solution k of the particle p and the optimal solution of the particle swarm , and randomly mutate the remaining λ gene positions, thereby updating the position k of the particle p

[0219] Step 3-47: The user reselects the cell. In the r-th iteration process, for according to the position k of the particle p (i.e., the switch action) to indicate the active cells, the user reselects to connect to the adjacent active cells and updates the traffic load of cell m

[0220] Step 3-48: Determine the validity of the particle. For judge whether the active cells indicated by the coordinate position of the particle p k are overloaded and whether the service rate of the user can be satisfied according to the updated traffic load of the active cells. If an active cell is overloaded, or the service rate of the user cannot be satisfied, the position of this particle cannot be used as the solution to the cell cooperative dormancy problem.

[0221] Step 3-49: Update the particle fitness value. For use the traffic load of cell m updated in Step 3-47 ​The power consumption of cell m is calculated using formula (3). Based on the calculated power consumption of all M cells, the particle p is calculated according to formula (13). k Current position Fitness value

[0222] Step 3-410: Update particle p according to formula (14). k The optimal solution in the r-th iteration process. In the r-th iteration, for According to particle p k Current position Fitness value And the historical optimal solutions of the previous r-1 iterations Fitness value Update particle p k The historical optimal solution up to the r-th iteration Particle p k The historical optimal solution up to the r-th iteration The update expression of is:

[0223]

[0224] Step 3-411: Update the optimal solution of the particle swarm P in the r-th iteration process The update expression of the optimal solution of the particle swarm P searching for the optimal solution of the cell collaborative sleep problem in the r-th iteration process is:

[0225]

[0226] Step 3-412: Determine whether the heuristic search algorithm falls into a local optimal solution in the r-th iteration process. If the particle swarm optimal solution Is not updated and the mutation factor λ does not exceed λ max , then increase the value of the mutation factor λ. This is because if the particle swarm optimal solution Is not improved, it indicates that the algorithm may be trapped in a local optimum at this time. Therefore, it is necessary to increase the value of the mutation factor to increase randomness, thereby increasing the tendency to expand the global search.

[0227] Step 3-413: Repeat steps 3-46 to 3-412 until the maximum number of iterations is reached, i.e., r = N r . Take the particle swarm optimal solution of the N r -th iteration As the working states of all M cells in the overall target area at the t-th unit time, and use it as the expert demonstration action vector at the t-th unit time That is

[0228] (4) Application Example 4

[0229] Application Example 4 provides an expert demonstration algorithm based on cell value degree ranking, which is used to explain the foregoing embodiments of calculating the second expert demonstration action vectors corresponding to each of the unit times for controlling the collaborative sleep of each of the cells. Specifically, referring to Figure 6 , Application Example 4 includes the following content:

[0230] This Application Example 4 proposes an expert demonstration algorithm based on cell value degree ranking. This algorithm determines whether the working state of a cell is active or sleeping based on the high or low value degree of the cell. Specifically, this algorithm tries to turn off low-value cells and turn on high-value cells as much as possible. The specific process is as follows:

[0231] Step 4-1: Definition of cell value degree. To better evaluate the energy-saving potential of a cell, this Application Example introduces the concept of cell value degree. In this Application Example, the cell value degree refers to the cell traffic load. The lower the cell value degree, the greater the possibility of it being turned off.

[0232] Step 4-2: Calculation of cell value degree. According to the cell traffic load acquisition process described in Application Example 1, the traffic load of cell m at the t-th unit time can be obtained Thus, based on the target area defined in Application Example 1, the total traffic load of the entire target area is expressed as Next, the cell value degree of cell m at the t-th unit time can be expressed as The calculation process is shown in formula (16):

[0233]

[0234] Step 4-3: Prediction of cell value degree. Since the on / off decision of a cell is based on the cell value degree in the future time period, it is necessary to predict the historical cell value degree obtained in Step 4-2. Based on the historical cell value degree data, a machine learning algorithm is used to predict the value degree in a certain future time period. For the time series prediction of cell value degree, traffic prediction models such as the triple exponential smoothing algorithm (Holt-winters) and the long short-term memory network (LSTM) can be used. To predict the cell value degree of cell m at the t-th unit time, the historical value degree sequence of the continuous l unit times before the t-th unit time of cell m is input into the time series prediction model. After continuous iteration, the predicted output of the prediction model at the t-th unit time is obtained, which is defined as the cell value degree prediction result

[0235] Step 4-4: Solve the sleep status of all M cells in the target area based on the predicted results of cell value. The expert demonstration algorithm based on the sorting of cell value is used to obtain the working status of all M cells in the target area at the t-th unit time wherein, is the working status of the m-th cell at the t-th unit time, when the cell m is in the active state at the t-th unit time is 1, otherwise it is 0. Taking the solution of the cell cooperative sleep problem at the t-th unit time as an example, the specific algorithm process is as follows:

[0236] Step 4-41: At the iteration round l = 0, assume that all M cells are in the active state.

[0237] Step 4-42: At the l-th iteration, according to the predicted results of the M cell values obtained in step 4-3 sort the M cells in descending order.

[0238] Step 4-43: At the l-th iteration, turn off the cell m with the lowest value, and the users accessing the cell m perform cell reselection to update the user connection status of the cells in the active state in the l-th round The cell reselection process of the users is shown in detail in step 62 of Application Example 1.

[0239] Step 4-44: According to the user connection status statistically analyze the service traffic data of all M cells after cell reselection.

[0240] Step 4-45: Determine whether the cell is overloaded. For each cell m' in the active cell set , according to the latest user connection status obtained after cell reselection and all N user traffic demands ρ t , calculate the value of each active cell m' in the active cell set . The value calculation formula of the active cell m' is:

[0241]

[0242] Based on the active cell set obtained according to formula (17) judge whether each active cell m' in the active cell set is overloaded according to the traffic load of each active cell m'.

[0243] Step 4-46: Determine whether the user demand rate is satisfied. For each user n in the user set, determine whether the active cell m' to which the user is currently connected can provide a satisfactory service rate for the user That is, whether it satisfies The definition of the service rate from the active cell m' to the user n is shown in formula (6) in Application Example 1. If the service rate of each user n satisfies the user traffic demand and the power consumption obtained according to formula (1) in Application Example 1 is reduced compared with that before shutdown, then the cell m can be placed in the sleep state; otherwise, it remains in the active state.

[0244] Step 4-47: The algorithm re-sorts the cells m' in the active cell set in descending order of value, and repeats steps 4-42 to 4-46: until there is cell overload in the active cell set or there is a user whose rate cannot be satisfied in the user set. Take the working states of the above-mentioned M cells as the expert demonstration action vector of the M cells in the t-th unit time

[0245] Based on the foregoing embodiments of the method for generating a model for cell cooperative sleep strategy, the present application further provides an embodiment of a cell cooperative sleep method. Refer to Figure 7 The cell cooperative sleep method specifically includes the following contents:

[0246] Step 300: Obtain the target traffic data of each cell in the target area in a historical unit time

[0247] Step 400: Input the target traffic data into a preset cell cooperative sleep strategy generation model, so that the cell cooperative sleep strategy generation model outputs the cell cooperative sleep strategy prediction result data in a unit time after the current time in the target area, where the cell cooperative sleep strategy generation model is pre-trained based on the method for generating a model for cell cooperative sleep strategy

[0248] Step 500: According to the cell cooperative sleep strategy prediction result data, perform corresponding sleep state switch actions for each cell in the target area

[0249] Among them, the embodiment of the method for generating a model for cell cooperative sleep strategy provided in step 400 can specifically adopt the processing flow of the embodiment of the method for generating a model for cell cooperative sleep strategy in the above embodiment, and its function will not be elaborated here. Reference can be made to the detailed description of the embodiment of the method for generating a model for cell cooperative sleep strategy

[0250] As can be seen from the above description, the cell collaborative sleep method provided by the embodiments of the present application can effectively improve the application effectiveness and reliability of the cell collaborative sleep strategy generated based on the cell collaborative sleep strategy generation model, so as to implement dynamic and adaptive cell switching decisions.

[0251] The present application also provides a cell collaborative sleep strategy generation model training device for executing all or part of the content in the cell collaborative sleep strategy generation model training method. See Figure 8 The cell collaborative sleep strategy generation model training device specifically includes the following content:

[0252] The model training module 10 based on expert demonstrations is used to train a preset neural network based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the collaborative sleep of each cell, so as to train the neural network into an expert demonstration network for predicting the expert demonstration action data corresponding to each cell in the target area.

[0253] The model migration and reinforcement learning module 20 is used to initialize the network parameters of the expert demonstration network in the current neural network and the target neural network in the reinforcement learning agent, so that the reinforcement learning agent performs joint iterative training on the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the deep reinforcement learning method, so as to train the current neural network into a cell collaborative sleep strategy generation model for predicting the cell collaborative sleep strategy corresponding to each cell in the target area.

[0254] The embodiments of the cell collaborative sleep strategy generation model training device provided by the present application can specifically be used to execute the processing flow of the embodiments of the cell collaborative sleep strategy generation model training method in the above embodiments, and its functions will not be elaborated here. Reference can be made to the detailed description of the embodiments of the cell collaborative sleep strategy generation model training method above.

[0255] Part of the cell collaborative sleep strategy generation model training of the cell collaborative sleep strategy generation model training device can be executed in the server and can be completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. The present application does not limit this. If all operations are completed in the client device, the client device may further include a processor for specific processing of the cell collaborative sleep strategy generation model training.

[0256] The above-mentioned client device may have a communication module (i.e., communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side. In other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.

[0257] Any suitable network protocol can be used for communication between the above-mentioned server and the client device, including network protocols that have not been developed as of the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may also, for example, include the RPC protocol (Remote Procedure Call Protocol) and REST protocol (Representational State Transfer) used on top of the above-mentioned protocols.

[0258] As can be seen from the above description, the cell cooperative sleep strategy generation model training device provided by the embodiments of this application can effectively improve the training effectiveness and reliability of the cell cooperative sleep strategy generation model, and can solve the problem of limited decision-making in imitation learning. Furthermore, it can effectively improve the application effectiveness and reliability of the cell cooperative sleep strategy generated based on the cell cooperative sleep strategy generation model to achieve dynamic and adaptive cell switching decisions.

[0259] This application also provides a cell cooperative sleep device for executing all or part of the cell cooperative sleep method. Refer to Figure 9 The cell cooperative sleep device specifically includes the following:

[0260] A traffic acquisition module 30, configured to acquire the target traffic data of each cell in the target area within a historical unit time.

[0261] A cell cooperative sleep strategy prediction module 40, configured to input the target traffic data into a preset cell cooperative sleep strategy generation model, so that the cell cooperative sleep strategy generation model outputs the cell cooperative sleep strategy prediction result data of the target area within a unit time after the current time, where the cell cooperative sleep strategy generation model is pre-trained based on the cell cooperative sleep strategy generation model training method.

[0262] The cell collaborative sleep control module 50 is configured to perform corresponding sleep state switching actions for each of the cells in the target area according to the predicted result data of the cell collaborative sleep policy.

[0263] The embodiment of the cell collaborative sleep device provided in this application can specifically be used to execute the processing flow of the embodiment of the cell collaborative sleep method in the above embodiment, and its functions will not be elaborated here. Reference can be made to the detailed description of the embodiment of the cell collaborative sleep method above.

[0264] The part of the cell collaborative sleep device for performing cell collaborative sleep can be executed in the client device or completed in the server. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor for specific processing of cell collaborative sleep.

[0265] As can be seen from the above description, the cell collaborative sleep device provided in the embodiment of this application can effectively improve the application effectiveness and reliability of the cell collaborative sleep policy generated by the cell collaborative sleep policy generation model, so as to achieve dynamic and adaptive cell switching decisions.

[0266] The embodiment of this application also provides an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is configured to execute the cell collaborative sleep policy generation model training method and / or the cell collaborative sleep method mentioned in the above embodiment. Among them, the processor and the memory can be connected through a bus or other means. Taking the bus connection as an example, the receiver can be connected to the processor and the memory in a wired or wireless manner.

[0267] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.

[0268] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the cell cooperative sleep strategy generation model training method and / or the cell cooperative sleep method in the embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, implements the cell cooperative sleep strategy generation model training method and / or the cell cooperative sleep method in the above method embodiments.

[0269] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0270] The one or more modules are stored in the memory and, when executed by the processor, execute the cell cooperative sleep strategy generation model training method in the embodiments.

[0271] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, the memory, the receiver, and the transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transmit and receive signals.

[0272] As an implementation manner, the functions of the receiver and the transmitter in the present application can be considered to be implemented through a transceiver circuit or a dedicated transceiver chip, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit, or a general-purpose chip.

[0273] As another implementation manner, it can be considered to use a general computer to implement the server provided in the embodiments of the present application. That is, the program codes for implementing the functions of the processor, the receiver, and the transmitter are stored in the memory, and the general processor implements the functions of the processor, the receiver, and the transmitter by executing the codes in the memory.

[0274] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing method for training a cell collaborative sleep strategy generation model and / or the cell collaborative sleep method are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.

[0275] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing method for training a cell collaborative sleep strategy generation model and / or the cell collaborative sleep method are implemented.

[0276] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0277] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0278] In the present application, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0279] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for training a small cell collaborative sleep strategy generation model, characterized in that Including: Based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the collaborative dormancy of each cell, train a preset neural network to train the neural network into an expert demonstration network for predicting the expert demonstration action data corresponding to each cell in the target area; the neural network includes: a deep Q network; Use the network parameters of the expert demonstration network to initialize the current neural network and the target neural network in the reinforcement learning agent, so that the reinforcement learning agent performs joint iterative training on the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the deep reinforcement learning method, so as to train the current neural network into a cell collaborative dormancy strategy generation model for predicting the cell collaborative dormancy strategy corresponding to each cell in the target area; The training of the preset neural network based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the collaborative dormancy of each cell to train the neural network into an expert demonstration network for predicting the expert demonstration action data corresponding to each cell in the target area includes: Obtain the target expert demonstration action vectors for controlling the collaborative dormancy of each cell corresponding to each unit time; Convert the target expert demonstration action vectors corresponding to each unit time into their respective target expert demonstration action indications; Respectively use the target expert demonstration action indications corresponding to each unit time as the labels of the historical traffic data of each cell corresponding to each unit time to be pre-obtained, so as to obtain the samples corresponding to each unit time and the expert demonstration action data set composed of each sample; Divide the expert demonstration action data set into a training set and a validation set; Based on the training set, train a preset neural network with a stochastic gradient descent algorithm, and use the validation set to verify the trained neural network, and determine the neural network that passes the verification as the current expert demonstration network for predicting the expert demonstration action data in the next unit time after the current time in the target area, and store the expert demonstration network parameters corresponding to the expert demonstration network; The use of the network parameters of the expert demonstration network to initialize the current neural network and the target neural network in the reinforcement learning agent, so that the reinforcement learning agent performs joint iterative training on the current neural network in the reinforcement learning agent and the target neural network in the reinforcement learning agent based on the deep reinforcement learning method, so as to train the current neural network into a cell collaborative dormancy strategy generation model for predicting the cell collaborative dormancy strategy corresponding to each cell in the target area includes: Use the expert demonstration network parameters to initialize the current neural network and the target neural network in the reinforcement learning agent; Initialize the experience replay buffer and the maximum number of iterations corresponding to the reinforcement learning agent, where the experience replay buffer is used to store tuple samples, and the tuple samples include: the current state, action, reward, and the next state; Take the working state of each cell in a unit time as the action corresponding to this unit time, take the predicted traffic of each cell in this unit time and the sleep state switch action in the previous unit time as the current state corresponding to this unit time, and set the reward function corresponding to the reward; where the working state includes: the active state represented by 1 and the sleep state represented by 0; Initialize the current neural network and the target neural network of the reinforcement learning agent according to the expert demonstration network parameters; Select the action according to the current neural network and execute the action, and calculate the corresponding reward based on the reward function. At this time, the environmental state becomes the next state to obtain the corresponding tuple sample, and store the tuple sample in the experience replay buffer; If the data of the tuple samples in the experience replay buffer meets the preset quantity threshold, randomly select multiple tuple samples from the experience replay buffer, so that in the current iteration round, the reinforcement learning agent, in a deep reinforcement learning manner, jointly trains the current neural network and the target neural network in the reinforcement learning agent based on each randomly selected tuple sample. Stop training until the number of iterations reaches the maximum number of iterations, and use the network parameters of the current neural network as the model parameters corresponding to the cell collaborative sleep strategy generation model for generating the prediction result of the cell collaborative sleep strategy in the unit time after the current time in the target area.

2. The method for training a small cell collaborative sleep strategy generation model according to claim 1, wherein The obtaining of the target expert demonstration action vectors corresponding to each of the unit times for controlling the collaborative sleep of each cell includes: Calculating, by using an expert demonstration algorithm based on heuristic search, the first expert demonstration action vectors corresponding to each of the unit times for controlling the collaborative sleep of each cell; And calculating, by using an expert demonstration algorithm based on cell value degree ranking, the second expert demonstration action vectors corresponding to each of the unit times for controlling the collaborative sleep of each cell; According to the power consumption values of the first expert demonstration action vectors and the second expert demonstration action vectors corresponding to each of the unit times, select one of the power consumption values of the first expert demonstration action vectors and the second expert demonstration action vectors corresponding to each of the unit times as the target expert demonstration action vector corresponding to this unit time.

3. The method for training a small cell collaborative sleep strategy generation model according to claim 1, wherein The reward function is composed of a power consumption saving benefit function, a user service quality satisfaction function, and an action error function for imitating the expert strategy; Among them, the value of the power consumption saving benefit function is calculated based on the total power consumption corresponding to all cells in the target area being in the active state in a unit time obtained in advance, the working state of each cell in this unit time, and the power consumption model; The value of the service quality satisfaction function of the user is calculated based on pre-acquired user available service rate, user demand traffic, and a switching variable; The value of the action error function imitating the expert strategy is calculated based on the pre-acquired target expert demonstration action vector; Among them, the user available service rate is calculated based on the bandwidth resource allocated to the associated user by the active cell in unit time and the signal-to-noise ratio from the active cell to the associated user.

4. The method for training a small cell collaborative sleep strategy generation model according to claim 3, wherein Before training a preset neural network based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the cooperative dormancy of each cell, it further includes: Collecting the historical traffic data, total number of users, and total number of cells corresponding to each cell in the target area, and preprocessing the historical traffic data to obtain the traffic data of each cell in the target area at each unit time; Inputting the traffic data of each cell at each unit time into a preset traffic prediction model, so that the traffic prediction model outputs the traffic prediction results of each cell at one unit time after each unit time; Summing up the traffic prediction results of each cell at one unit time after each unit time to obtain the predicted total traffic of the target area at one unit time after each unit time, and based on the predicted total traffic and the total number of users, allocating the user expected traffic and user demand traffic corresponding to each user in the target area at one unit time after each unit time.

5. The method for training a small cell collaborative sleep strategy generation model according to claim 4, wherein Before training a preset neural network based on the historical traffic data of each cell in the target area corresponding to each unit time and the expert demonstration action data for controlling the cooperative dormancy of each cell, it further includes: Setting a power consumption model, where the power consumption model corresponding to each cell in one unit time consists of the fixed power consumption of the cell in this unit time and the transmission power consumption of the cell in this unit time; Setting the cell working state, where the cell working state includes: the cell is in the active state with a value of 1 in one unit time, and the cell is in the dormant state with a value of 0 in one unit time.

6. The method for training a small cell collaborative sleep strategy generation model according to claim 2, wherein The step of calculating the first expert demonstration action vector for controlling the cooperative dormancy of each cell corresponding to each unit time by using the expert demonstration algorithm based on heuristic search includes: Defining the working state of each cell in the target area as the position of a particle, defining the total power consumption of all cells in the target area as the fitness of the particle, defining a particle swarm according to various switching combinations for all cells in the target area, and defining the update process of the particle; Based on the position of the particle, the fitness of the particle, the particle swarm, and the update process of the particle, an expert demonstration algorithm based on heuristic search is adopted to search for the optimal solution to the corresponding cell collaborative sleep problem for each cell in the target area, so as to obtain the first expert demonstration action vector corresponding to each unit time for controlling the collaborative sleep of each cell.

7. The method for training a small cell collaborative sleep strategy generation model according to claim 2, wherein The second expert demonstration action vector corresponding to each unit time for controlling the collaborative sleep of each cell calculated by using the expert demonstration algorithm based on cell value degree ranking includes: Based on the traffic loads of each cell in each unit time obtained in advance, determine the total traffic load of the target area in each unit time; According to the traffic loads of each cell in each unit time and the total traffic load, determine the cell value degrees of each cell in each unit time respectively; Input the cell value degrees of each cell in each historical unit time before a unit time into a preset time series prediction model, so that the time series prediction model outputs the cell value degree prediction results of each cell in each unit time; Solve the working states corresponding to each cell in the target area according to the cell value degree prediction results of each cell in each unit time, so as to obtain the second expert demonstration action vector corresponding to each unit time for controlling the collaborative sleep of each cell.

8. A method for collaborative dormancy of a cell, characterized in that, It includes: Obtain the target traffic data of each cell in the target area in a historical unit time; Input the target traffic data into a preset cell collaborative sleep policy generation model, so that the cell collaborative sleep policy generation model outputs the cell collaborative sleep policy prediction result data within a unit time after the current time, where the cell collaborative sleep policy generation model is pre-trained based on the cell collaborative sleep policy generation model training method according to any one of claims 1 to 7; According to the cell collaborative sleep policy prediction result data, perform corresponding sleep state switch actions for each cell in the target area.

Citation Information

Patent Citations

  • Electric power communication equipment test resource scheduling method based on reverse deep reinforcement learning

    CN111026548A

  • Reward estimation via state prediction using expert demonstrations

    US20190272465A1