A method and system for secure dynamic resource allocation in power wireless private networks

By constructing a convex optimization model in the power wireless private network and using deep reinforcement learning to train a neural network, power allocation is optimized, solving the problems of low resource allocation efficiency and insufficient energy consumption in the existing technology. This achieves high-efficiency dynamic resource allocation and ensures the stable operation of the power wireless private network.

CN113543225BActive Publication Date: 2025-10-28GLOBAL ENERGY INTERCONNECTION RES INST CO LTD +4
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202010294058.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-15
Publication Date
2025-10-28
Estimated Expiration
2040-04-15

AI Technical Summary

Technical Problem

Existing technologies for dynamic resource allocation in power wireless private networks suffer from low learning efficiency, insufficient energy consumption considerations, and long processing times, making it difficult to achieve high-efficiency dynamic bandwidth and power allocation.

Method used

A dynamic high-efficiency resource allocation model for power wireless private networks is constructed using a convex optimization method. A neural network is trained using deep reinforcement learning (DRL), and the connection relationships, data flow information, and minimum transmission energy are stored in a memory pool to optimize power allocation and achieve optimal energy efficiency.

Benefits of technology

Find the optimal resource allocation method in a short period of time to ensure the lowest overall energy efficiency while meeting the needs of user equipment, and ensure the healthy, safe, reliable and efficient operation of the power wireless private network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113543225B_ABST
    Figure CN113543225B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for secure dynamic resource allocation in a power wireless private network, comprising: using a convex optimization method to obtain the minimum transmission energy corresponding to each connection relationship in a pre-constructed dynamic high-efficiency resource allocation model for the power wireless private network; storing each connection relationship, the data flow information in the connection relationship, and the minimum transmission energy in a set memory pool; selecting sample data from the memory pool for training to obtain the optimal connection relationship, the optimal power allocation value, and the optimal energy efficiency value; wherein, the dynamic high-efficiency resource allocation model for the power wireless private network is constructed based on the service relationship between each base station and each user equipment in the power wireless private network. A dynamic high-efficiency resource allocation model for the power wireless private network has been established, and the optimal resource allocation framework has been found.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power communication, and specifically to a method and system for secure dynamic resource allocation in a power wireless private network. Background Technology

[0002] In recent years, as a key component of future networks and power grids, power wireless private networks have been specifically built to meet power demands. This necessitates a deep integration of technology and business, requiring customized development based on specific needs and demanding that this private network be refined to the highest standards. Furthermore, due to the limited computing resources in power wireless private networks, and considerations for the ecological environment and economic costs, how to dynamically and energy-efficiently allocate bandwidth and power resources while meeting the needs of user equipment has become a crucial issue.

[0003] The integration of reinforcement learning and neural networks has a long history. Benefiting from big data, improved computing power, and new algorithmic technologies, deep learning has achieved a series of exciting successes, especially the combination of deep learning and reinforcement learning, namely deep reinforcement learning. DQN (deep Q-learning) improves upon QL by using a deep convolutional neural network to approximate the value function Q, replacing the Q-table; it utilizes experience replay to train the reinforcement learning process, breaking down correlations between data points and making the neural network training convergent and stable; it uses gradient descent to update network parameters and sets up a separate target network to handle the TD bias in the time difference algorithm, avoiding training instability. Although deep reinforcement learning (DRL) has been applied to many projects related to power wireless private networks, it has not yet been explored in the area of ​​dynamic high-efficiency resource allocation.

[0004] This paper addresses the resource allocation problem in power wireless private networks, focusing on how to allocate resources more dynamically, efficiently, and effectively to meet the power and SINR constraints of each user device.

[0005] To understand existing resource allocation methods for power grid wireless private networks, existing papers and patents were searched, compared, and analyzed, and the following technical information with high relevance to this invention was selected:

[0006] Technical Method 1: Patent publication number CN110062026A, "Joint Optimization Scheme for Resource Allocation and Computation Offloading in Mobile Edge Computing Networks," discloses a joint allocation scheme for wireless bandwidth and computing resources in mobile edge computing networks. Belonging to the fields of wireless communication and mobile edge computing, it solves the resource contention and load balancing problems in multi-user device and multi-mobile edge server deployment scenarios in heterogeneous wireless networks. The scheme specifically includes: a macro base station controller collecting computation offloading request information sent by all user devices within the current time slot and notifying all MEC servers within its jurisdiction to report the current remaining resource status; based on the acquired information, the macro base station controller performs an initial matching of user device computing tasks and MEC server resources; formulates rules for allocating wireless bandwidth and computing resources in MEC servers; and establishes a cooperative game model to output the final matching strategy set. This invention takes into account the characteristics of each user device, effectively reducing the cost of computation offloading and saving energy consumption on mobile user devices. With the same number of MEC servers deployed, this invention can accommodate more computation offloading tasks, balance the communication and computing load between servers, and improve the system's task execution efficiency.

[0007] Technical Method 2: Patent publication number CN110248206A, entitled "A Resource Allocation Method, Apparatus, and Electronic Equipment for a Power Wireless Private Network System," discloses a resource allocation method, apparatus, and electronic equipment for a power wireless private network system. The method includes obtaining a caching scheme, which includes the caching status of each segment at a first base station; determining a transmission delay set based on the caching scheme, which includes the transmission delay when each terminal device obtains each segment; determining a recommended scheme based on the caching scheme and the transmission delay set to minimize the overall latency; wherein the recommended scheme includes recommended content for each terminal device, which includes at least one of the F video files; the overall latency is the sum of the expected latency values ​​of the U terminal devices; and determining an update scheme for the caching scheme based on the transmission delay set and the recommended scheme to minimize the overall latency. This technical solution can alleviate the transmission pressure on the communication link.

[0008] Technical Solution 3: Patent publication number CN109814951A, "Joint Optimization Method for Task Offloading and Resource Allocation in Mobile Edge Computing Networks," discloses a joint optimization method for task offloading and resource allocation in mobile edge computing networks, including the following steps: Step 1: Establish a scenario model based on OFDMA with multiple MEC base stations and multiple user devices, where the MEC base stations support multi-user device access; Step 2: Introduce an offloading decision mechanism; simultaneously construct a local computing model and a remote computing model, select user devices that need to perform computation offloading, and establish a computation task offloading and resource allocation scheme based on minimum energy consumption under the above conditions, satisfying latency constraints; Step 3: Simplify the problem by fusing three mutually constrained optimization variables: offloading decision variables, wireless resource allocation variables, and computation resource allocation variables; Step 4: Obtain the offloading decision and resource allocation result that minimizes the total energy consumption of user devices in the MEC system through a branch-and-bound algorithm. This invention has the advantage of effectively reducing system energy consumption while ensuring strict latency constraints.

[0009] The solutions mentioned above in related technologies all have certain drawbacks:

[0010] Scheme 1 involves the macro base station controller collecting computation offloading request information sent by all user equipment terminals within the current time slot. Based on this information, it performs an initial match between user equipment computation tasks and MEC server resources; formulates rules for allocating wireless bandwidth and computation resources within the MEC server; and establishes a cooperative game model to output the final matching strategy set. However, this scheme has low learning efficiency, resulting in slow allocation speeds when new requests arrive.

[0011] Scheme 2 designs a caching scheme. Based on the caching status of each segment at the first base station, a transmission delay set is determined. This transmission delay set includes the transmission delay when each terminal device acquires each segment. Based on the caching scheme and the transmission delay set, a recommended scheme is determined to minimize the overall latency. However, this scheme lacks consideration for energy consumption, focusing primarily on latency control.

[0012] Option 3 establishes a scenario model based on OFDMA with multiple MEC base stations and multiple user devices, introducing an offload decision mechanism. Simultaneously, it constructs local and remote computing models, selects user devices requiring computational offload, and then uses a branch-and-bound algorithm to obtain the offload decision and resource allocation results that minimize the total energy consumption of user devices in the MEC system. However, this process is time-consuming and slow, and this method is not entirely comprehensive. Summary of the Invention

[0013] To address the shortcomings of existing technologies, this invention proposes a method for secure dynamic resource allocation in power wireless private networks, comprising:

[0014] For each connection relationship in the pre-constructed dynamic high-efficiency resource allocation model of the power wireless private network, the minimum transmission energy corresponding to each connection relationship is obtained by using a convex optimization method; and each connection relationship, the data flow information in the connection relationship, and the minimum transmission energy are stored in a set memory pool.

[0015] Sample data is selected from the memory pool and trained to obtain the optimal connectivity, optimal power allocation value, and optimal energy efficiency value.

[0016] The dynamic high-efficiency resource allocation model for the power wireless private network is constructed based on the service relationship between each base station and each user equipment in the power wireless private network.

[0017] Preferably, the connection relationship between each base station and each user equipment includes:

[0018] Based on the data rate of the stream carried by the link between the base station and the user equipment, the spectral efficiency of the wireless link is determined by using the Shannon limit.

[0019] The data rate capacity implemented by the wireless link is determined based on the spectral efficiency of the wireless link.

[0020] The base station set and the core network are connected via a wired backhaul link, and the user equipment and the base station transmit data via a wireless link.

[0021] Preferably, the step of using a convex optimization method to obtain the minimum wireless transmission energy corresponding to each connection relationship; and storing each connection relationship, the data stream information in the connection relationship, and the minimum wireless transmission energy into a set memory pool includes:

[0022] 1) Set up a state space based on the connection relationship; randomly initialize a state from the state space, initialize the memory pool, set the number of observation steps, and set the number of observation steps as the observation value;

[0023] 2) Based on the current state, select an action, obtain the corresponding reward value, and the state after the action ends. Save the current state, action, reward value, and state after the action ends to the memory pool.

[0024] 3) Check if the amount of data stored in the memory pool exceeds the observed value. If not, go to 4); otherwise, end.

[0025] 4) Determine if the maximum number of search steps has been reached. If the maximum number of search steps has been reached, randomly reset a state; otherwise, set the state after the action ends to the current state s and return to step 2).

[0026] Preferably, the setting of the state space based on the connection relationship includes:

[0027] An array M is constructed based on the permutations and combinations of connections between all user equipment and each base station, where M = {s1, s2, ..., s...}. k}; where k is the number of user devices, s1, s2, ..., s k The connection relationship between user equipment {1,2,…,k} and each base station;

[0028] A state space S is constructed based on the connection relationships formed by arrays of all base stations, where S = {M1, M2, ..., M}. N}; where N is the total number of connections, N = H k H represents the number of base stations.

[0029] Preferably, the step of selecting an action based on the current state, obtaining a corresponding reward value, and the state after the action ends includes:

[0030] The action space is set as A based on the number of connections, where A = {1, 2, ..., N}, and the numbers represent the position in the current state, and N represents the total number of connections.

[0031] The minimum wireless transmission energy in the current state is obtained by using the convex optimization method, and combined with the numbers in the action space A by the greedy policy;

[0032] The state after the action ends is determined based on the position of the number in the current state;

[0033] The reward is set to E. max -E, where E max E represents the maximum energy consumption that the current base station can provide, and E represents the energy consumption after taking the action.

[0034] Preferably, the energy consumption value after the action is calculated using the following formula:

[0035]

[0036] Among them, E T P represents the energy expenditure after the action. ij S represents the power allocated by base station i to user equipment j. ij When a user equipment j is served by a base station i, there is a connection relationship. t0 is the operation time, I is the set of all base stations, and J is the set of all user equipment.

[0037] Preferably, the data stream information is as follows:

[0038]

[0039] in, It is the data rate of stream f carried on base station i and user equipment j. For the streams carried on base station i and user equipment j to have the required data packet size t, 0ij Let be the operation time from base station i to a user equipment j.

[0040] Preferably, the data rate capacity achieved by the wireless link is as follows:

[0041] r ij =x ij B i Υ ij

[0042] Where, r ij x represents the data rate at which user equipment j is served by base station i. ij B represents the allocation ratio of base station i to user equipment j. i γ represents the total available spectrum bandwidth of base station i. ij This represents the spectral effect of the wireless link served by base station i for user j.

[0043] Preferably, the step of selecting sample data from the memory pool and training to obtain the optimal connectivity, optimal power allocation value, and optimal energy efficiency value includes:

[0044] Select sample data from the memory pool;

[0045] Based on each sample data, the convex optimization method is used to calculate the pre-set objective function and constraints to obtain the optimal power allocation solution and the corresponding Q value table under the connection relationship corresponding to the current state s in the sample data, as well as the targetQ value corresponding to the Q value table;

[0046] The optimal connectivity, optimal power allocation, and optimal energy efficiency are obtained by training the neural network using a Q-value table and a target Q-value.

[0047] Preferably, the targetQ value corresponding to the Q value table is calculated using the following formula:

[0048] Q(s,A)=R+γmax[Q(s',all_actions)];

[0049] Where s' is the next state, γ is the reward decay coefficient, and all_actions is all actions.

[0050] Preferably, the objective function is calculated as follows:

[0051]

[0052] Where t0 is the operation time, E C (sij ) represents the energy required for node operations. For fixed power consumption, P C To consume power, s ij When base station i serves user equipment j, there is a connection relationship, r f For the required data rate, I is the set of base stations, J is the set of user equipment, and F is the set containing all flows on the network link.

[0053] Preferably, the constraints include:

[0054] Each user equipment can only be served by one base station;

[0055] The size of the stream carried by the wireless link cannot exceed the data rate capacity implemented by the stream;

[0056] The total transmit power and allocated bandwidth shall not exceed the total power and total bandwidth provided by the total transmit power and allocated bandwidth;

[0057] User equipment receives power constraints and signal-to-noise ratio constraints.

[0058] Preferably, after calculating the optimal solution based on the objective function, the method further includes: performing dynamic bandwidth and power resource allocation based on the optimal solution.

[0059] Based on the same inventive concept, this invention also proposes a system for secure dynamic resource allocation in a power wireless private network, comprising: a construction module and a learning module;

[0060] The construction module: uses a convex optimization method to obtain the minimum transmission energy corresponding to each connection relationship in the pre-constructed dynamic high-efficiency resource allocation model of the power wireless private network; and stores each connection relationship, the data flow information in the connection relationship, and the minimum transmission energy into a set memory pool;

[0061] The learning module selects sample data from the memory pool and trains it to obtain the optimal connectivity, optimal power allocation value, and optimal energy efficiency value.

[0062] The dynamic high-efficiency resource allocation model for the power wireless private network is constructed based on the service relationship between each base station and each user equipment in the power wireless private network.

[0063] Preferably, the learning module includes: a selection unit, a calculation unit, and a training unit;

[0064] The selection unit is used to select sample data from the memory pool;

[0065] The computing unit is used to calculate the pre-set objective function and constraints based on each sample data using a convex optimization method, to obtain the optimal power allocation solution and the corresponding Q value table under the connection relationship corresponding to the current state s in the sample data, as well as the target Q value corresponding to the Q value table;

[0066] The training unit is used to train a neural network using a Q-value table and a target Q-value to obtain the optimal connection relationship, the optimal power allocation value, and the optimal energy efficiency value.

[0067] The beneficial effects of this invention are as follows:

[0068] This invention proposes a method and system for secure dynamic resource allocation in a power wireless private network, comprising: using a convex optimization method to obtain the minimum transmission energy corresponding to each connection relationship in a pre-constructed dynamic high-efficiency resource allocation model for the power wireless private network; storing each connection relationship, the data flow information in the connection relationship, and the minimum transmission energy in a set memory pool; selecting sample data from the memory pool and training it using a deep reinforcement learning (DRL) neural network model to obtain the optimal connection relationship, the optimal power allocation value, and the optimal energy efficiency value; wherein, the dynamic high-efficiency resource allocation model for the power wireless private network is constructed based on the service relationship between each base station and each user equipment in the power wireless private network;

[0069] The technical solution provided by this invention establishes a dynamic high-efficiency resource allocation model for power wireless private networks and finds the optimal resource allocation framework. With the goal of minimizing energy consumption and meeting the needs of user equipment, a deep reinforcement learning method is used to allocate power to all user equipment. Finally, simulation experiments verify the effectiveness of the algorithm, ultimately achieving the ability to find a relatively ideal resource allocation method in a short time, ensuring both the lowest overall energy efficiency and the basic needs of user equipment, thus guaranteeing the healthy, safe, reliable, and efficient operation of the power wireless private network.

[0070] This invention utilizes deep reinforcement learning to first study the operational scenarios and requirements of power wireless private networks; secondly, it analyzes the power allocation and bandwidth allocation methods in power wireless private networks, and establishes a network optimization model for the overall network planning scheme.

[0071] This invention proposes a method for dynamic resource allocation in power wireless private networks based on a Deep Reinforcement Learning (DRL) framework. Under a specific connection relationship, the minimum wireless transmission energy is first obtained using a convex optimization method. Then, DQN is used for iteration. Based on the convex optimization results, the optimal connection relationship and the optimal power allocation value are found to calculate the optimal energy efficiency value. Simulation results show that we obtained the energy efficiency value when the algorithm converges and tends to be stable. Compared with the nearest distance strategy and the user equipment clustering strategy, the energy efficiency value is the lowest. This verifies the efficiency of the DRL-based framework and its effectiveness in meeting user equipment requirements and achieving dynamic high-efficiency resource allocation. Attached Figure Description

[0072] Figure 1 A flowchart of a method for secure dynamic resource allocation in a power wireless private network provided by the present invention;

[0073] Figure 2 This invention provides a method for dynamic resource allocation in a secure power wireless private network, including a resource allocation diagram.

[0074] Figure 3 The present invention provides a method for dynamic resource allocation in a power wireless private network, including a flowchart of the resource allocation algorithm steps.

[0075] Figure 4 This invention provides a reward comparison chart for different numbers of user devices under a DRL strategy;

[0076] Figure 5 This invention provides a comparison chart of losses under different numbers of user devices under a DRL strategy;

[0077] Figure 6 This invention provides an energy efficiency comparison chart for different numbers of user devices under a DRL strategy;

[0078] Figure 7 This invention provides a comparison chart of energy efficiency values ​​under three different strategies;

[0079] Figure 8 This invention provides a comparison chart of power consumption values ​​under three different strategies;

[0080] Figure 9 This invention provides a comparison chart of the received power performance of user equipment under three different strategies;

[0081] Figure 10 A cumulative probability distribution diagram of SINR provided by the present invention;

[0082] Figure 11 This invention provides a system architecture diagram for secure dynamic resource allocation in a power wireless private network. Detailed Implementation

[0083] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0084] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0085] Example 1

[0086] This invention provides a method for secure dynamic resource allocation in a power wireless private network, such as... Figure 1 Shown, including:

[0087] Step 1: For each connection relationship in the pre-constructed dynamic high-efficiency resource allocation model of the power wireless private network, use the convex optimization method to obtain the minimum transmission energy corresponding to each connection relationship; and store each connection relationship, the data flow information in the connection relationship, and the minimum transmission energy into a set memory pool;

[0088] Step 2: Select sample data from the memory pool and train to obtain the optimal connection relationship, optimal power allocation value, and optimal energy efficiency value;

[0089] The dynamic high-efficiency resource allocation model for the power wireless private network is constructed based on the service relationship between each base station and each user equipment in the power wireless private network.

[0090] 1. The dynamic high-efficiency resource allocation model for power wireless private networks of the present invention

[0091] (1) Dynamic high-efficiency resource allocation model for power wireless private networks

[0092] Consider the downlink transmission scenario in a heterogeneous network (HetNet) supporting SDN, consisting of a set of BSs I := {1,…,i,,,,I}, a set of user equipments J := {1,…,j,,,,J}, and a core network N. The BS set I and the core network (core routers) N ​​are connected via a wired backhaul link. User equipment J transmits data with the BS set I via a wireless link.

[0093] The set of flows F = {1, ..., f, ..., F} operates in the network under consideration, where each flow has a required packet size of m. f Required data rate r fThe transmission time is t0.

[0094]

[0095] We assume that multiple flows can reach a single user device (multiple services), and one link supports one flow. Let be the data rate of flow f carried on a certain BS and i and a certain user device j.

[0096] The data stream information is as follows:

[0097]

[0098] in, It is the data rate of stream f carried on base station i and user equipment j. For the streams carried on base station i and user equipment j to have the required data packet size t, 0ij Let be the operation time from base station i to a user equipment j.

[0099] The capacity of a wireless link depends on the ratio of radio resources allocated to that link by the network. Therefore, using the Shannon limit, the spectral efficiency of a wireless link is defined as...

[0100]

[0101] Where g ij It is the large-scale channel gain, which includes path loss and shadowing between transmitting node i (the source of the wireless link) and receiving node j (the destination of the wireless link). P ij (Watts) is the transmitted power on the wireless link. N0 is the noise power spectral density (PSD), B i x is the total available spectrum bandwidth of base station i. ij ∈[0,1] is the allocation ratio of base station i to user equipment j. Therefore, the achievable data rate capacity of the wireless link is

[0102] r ij =x ij B i Υ ij

[0103] We assume that node i ∈ I is equipped with a cache and stores popular content of Si. For example... Figure 2 As shown, in this proposed scheme, the total S content files are stored on a content server (cloud center or internet source), and each piece of content has a standardized size of 1. This assumption is reasonable because we can slice the content into chunks of the same length. Since nodes in a mobile network have only limited storage capacity, Si is much smaller than that in a cloud center.

[0104] Where, rij x represents the data rate at which user equipment j is served by base station i. ij B represents the allocation ratio of base station i to user equipment j. i γ represents the total available spectral bandwidth of source node i. ij This represents the spectral effect of the wireless link served by base station i for user j.

[0105] We assume that the maximum scheduling CPU computation frequency at points BS and i is c. i To process bit information, you need C. i CPU cycle, which means c i r f (Period / second) is the minimum requirement to support flow f starting from BS i. The connection relationship between base station i and user equipment j is represented as S. ij ∈{0,1}, where S ij =1 indicates that base station i has a connection with user equipment j, S ij =0 indicates that user equipment j and base station i have no connection relationship.

[0106] (2) Objective function

[0107] The energy consumed to support flow F in a network consists of two main parts: node operation energy and transmission energy. The node operation energy E is shown below. C (s ij (Depends on fixed power consumption) (e.g., circuitry, control signals, content cache) and computational power consumption p C (watts / cycle), where J represents the set of user equipment and I represents the set of base stations.

[0108]

[0109] Where t0 is the operation time (seconds). Furthermore, we ignore the energy used by the content server because our goal is to minimize the energy efficiency of the system under consideration; t0 is the operation time, E C (s ij ) represents the node operation energy, s ij When base station i serves user equipment j, there is a connection relationship, r f For the required data rate, I is the set of base stations, J is the set of user equipment, and F is the set containing all flows on the network link.

[0110] Transmission power depends on both wireless transmission power and backhaul transmission power. The wireless transmission power is set to p. ij , p ij This refers to the power allocated by BS i to user equipment j. Backhaul transmission power is represented as P. w .

[0111] The wireless transmission energy is:

[0112]

[0113] The return transmission energy is:

[0114]

[0115] Therefore, total energy consumption is

[0116] η EE =(E C (s ij )+E T (s ij , p ij )+E B )

[0117] Therefore, the energy efficiency is

[0118]

[0119] The combined energy-saving problem can be mathematically represented as min S,P E b

[0120] Restricted to:

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128] In constraint (ab), s ij ∈{0,1} is a binary decision variable used to indicate whether user equipment j is connected to a certain base station i. Each user equipment can only be served by one base station.

[0129] Constraint (c) reflects that the size of the stream carried by the wireless link cannot exceed its achievable data rate capacity.

[0130] The constraint (de) reflects that for any BS, its total transmit power and allocated bandwidth cannot exceed its total available power and bandwidth.

[0131] The constraint (fg) reflects the received power constraint and SINR constraint of the user equipment. Among them, ω=-110dBm is the minimum received power, and φ=-3dB is the SINR limit.

[0132] The above model comprehensively considers the computational and caching capabilities of power grid wireless private network nodes, seeking a dynamic, energy-efficient resource allocation framework with the goal of minimizing energy efficiency. Some methods utilize Q-Learning to allocate resources by minimizing the energy consumption of the power grid wireless private network; however, they do not consider the magnitude of energy efficiency. Moreover, compared to Q-learning, DRL has a faster training speed and is more suitable for large action spaces. Compared to heuristic algorithms, DRL can avoid getting trapped in local optima, thus obtaining a globally optimal solution. Therefore, we choose to use DRL to solve this optimization problem.

[0133] 2. Dynamic and Energy-Efficient Resource Allocation Strategy Based on Deep Reinforcement Learning

[0134] (1) Algorithm Implementation

[0135] DRL consists of two parts: an agent and the external environment. Changes in the state of the external environment are achieved by the agent taking different actions, and the agent then receives a reward from the external environment. The goal of DRL is to find the optimal strategy to maximize the reward.

[0136] In this embodiment, we propose a dynamic high-efficiency resource allocation framework for power grid wireless private networks based on DRL. The goal is to minimize the energy efficiency of the power grid wireless private network while meeting the needs of each user device and not exceeding the maximum power and bandwidth capacity of each base station. To reduce the state space size of the framework, under a specific connection relationship, we first use a convex optimization method to obtain the minimum wireless transmission energy, and then use DQN for iteration. Based on the convex optimization results, we find the optimal connection relationship and the optimal power allocation value to calculate the optimal energy efficiency value.

[0137] State Space: As defined above, we use 0 to represent that a user equipment is not served by a base station, and 1 to represent that the user equipment is served by a base station. In this embodiment, we assume there are three base stations, that is, the number of base stations H = 3. Therefore, for any user equipment, there are three possible connection relationships with the base stations: u1 = [1, 0, 0], u2 = [0, 1, 0], u3 = [0, 0, 1]. u1 indicates that the user equipment is served by base station 1, u2 indicates that the user equipment is served by base station 2, and u3 indicates that the user equipment is served by base station 3. Therefore, our state space is the permutation and combination of the three connection relationships of all user equipment. Specifically, assuming the number of user equipment is K, we can use M = {s1, s2, ..., s...} kLet} represent a certain connection relationship between all user equipment and the base station, where s1, s2, ..., s k It is a value among u0, u1, and u2. Therefore, our state space can be represented as S = {M1, M2, ..., M}. N}, where N = 3 k This represents the total number of connection relationships. A state space is defined based on these connection relationships, including: an array M, M = {s1, s2, ..., s...}, constructed based on the permutations and combinations of connections between each user equipment and each base station. k}; where k is the number of user devices, s1, s2, ..., s k The connection relationship between user equipment {1,2,…,k} and each base station;

[0138] A state space S is constructed based on the connection relationships formed by arrays of all base stations, where S = {M1, M2, ..., M}. N}; where N is the total number of connections, N = H k H represents the number of base stations.

[0139] Action Space: Actions are used by the DRL agent to indicate the next state based on the training results and the action space. We define the action space as A = {1, 2, ..., N}. The DRL agent obtains the minimum wireless transmission energy E in the current state through a convex optimization method. t The analysis is performed to select the next state, that is, the number in the action space A is used to indicate the location of the next state in the state space, so as to obtain the next state.

[0140] Reward: The reward represents the degree of conformity with the framework objective; a larger reward indicates a better fit with the optimization objective. In this framework, our goal is to minimize energy efficiency while satisfying constraints. Therefore, the lower the energy consumption, the larger the reward value. We define the immediate reward as E. max -E, where E max This represents the maximum energy consumption the base station can provide, and E represents the energy consumption after taking this action. The formula for calculating the energy consumption after the action is as follows: E T (p ij s ij )=∑ j∈J ∑ i∈ I t0p ij s ij

[0141] Among them, E T P represents the energy expenditure after the action. ij S represents the power allocated by base station i to user equipment j. ijWhen a user equipment j is served by a base station i, there is a connection relationship. t0 is the operation time, I is the set of all base stations, and J is the set of all user equipment.

[0142] DRL comprises two phases: an offline network construction phase and an online deep Q-learning phase. The primary task of the offline phase is to utilize a CNN to obtain the relationship between the state-action pair (s, a) and the value function Q(s, a), where the value function is the cumulative discounted reward for performing action a in state s. Here, s′ represents the next state, and a′ represents the next action.

[0143] Q(s,a)=r(s,a,s′)+λQ(s′,a′)

[0144] Where λ represents the discounted parameter, and r(s, a, s′) represents the reward obtained by performing action a. Offline construction requires accumulating sufficient value estimates and corresponding (s, a) samples, and using memory replay to smooth the training process. Offline DNNs need to accumulate sufficient value estimate samples and corresponding (s, a) to make the DNN sufficiently accurate.

[0145] Deep Q-learning performs online dynamic control based on an offline-built DNN. During the online learning process, at each time interval, the DRL agent uses a CNN to obtain an estimated Q-value and selects action 'a' using a greedy policy. Actions are randomly selected with a probability of 'a' and those with a probability of 1-'a', while actions with the maximum estimated Q-value are selected. After selecting actions, a certain connection relationship is obtained, and the convex optimization framework uses this connection relationship to calculate the minimum value of the following wireless transmission energy.

[0146]

[0147] Specifically, the established connectivity relationships are input into the convex optimization method. Based on this, the convex optimization method can find the optimal power allocation between each base station and each user device according to the objective function and constraints, thus obtaining the minimum wireless transmission energy under a given connectivity relationship. The total energy consumption is then calculated, and the immediate reward *r* and the next state *s′* are observed during interaction with the environment. The state transitions (s, a, r, s′) are stored in memory. Afterward, DQN randomly draws a portion of data from the memory pool to iteratively estimate and update the network parameters. Simultaneously, after a certain number of steps, the estimated network parameters are synchronized to the target network. Because different actions yield different rewards, the network parameters tend towards the optimal level.

[0148] (2) Algorithm Flow

[0149] The steps of the dynamic high-efficiency resource allocation algorithm based on deep reinforcement learning are as follows:

[0150] 1) Randomly initialize a state s, s∈S, for the power wireless private network, initialize the memory pool, and set the observation value (the number of steps observed before sample training);

[0151] 2) Based on the current state s, choose an action a, a∈A, obtain the corresponding reward value r, r∈R, and the state s' after the action ends.

[0152] The action space is defined as A, where A = {1, 2, ..., N}, where the number represents the position in the current state and N represents the total number of connections. The minimum wireless transmission energy in the current state is obtained using a convex optimization method, combined with the number in action space A obtained through a greedy strategy. The state after the action ends is determined based on the position of this number in state space S. Simultaneously, the reward is set to E. max -E, where E max E represents the maximum energy consumption that the current base station can provide, and E represents the energy consumption after taking the action.

[0153] And save the relevant parameters s, a, r, s' to the memory pool;

[0154] 3) Determine if the amount of data stored in the memory pool exceeds the observed value. If not, proceed to step 4). If the data is sufficient, proceed to step 5.

[0155] 4) Determine if the search process has ended (set the maximum number of search steps before starting the search).

[0156] ① If the maximum number of search steps is reached, randomly reset the state of s;

[0157] ②If the search does not reach the maximum number of steps, update the current state s to s';

[0158] Return to step 2)

[0159] 5) Begin training:

[0160] ① Randomly select a certain proportion of data from the memory pool as samples for training;

[0161] ② Take the randomly selected sample state s' as the training sample, use the convex optimization method to calculate the optimal power allocation solution under the current state of this sample, and obtain the Q value table under the corresponding state;

[0162] ③ Calculate the targetQ value corresponding to the Q-value table according to the formula; the formula is:

[0163] Q(s,A)=R+γmax[Q(s',all_actions)]

[0164] Where s' is the next state, γ is the reward decay coefficient, and all_actions is all actions, i.e. the overall action space, which can be replaced by A;

[0165] 6) Train the neural network using the Q-value table and the target Q-value.

[0166] 7) End

[0167] The flowchart of the business channel optimization algorithm based on deep reinforcement learning is as follows: Figure 3 as follows:

[0168] An optimal resource allocation framework. This invention first studies the operation scenarios and requirements of power wireless private networks; secondly, it analyzes the power allocation and bandwidth allocation methods in power wireless private networks, and establishes a network optimization model for the entire network planning scheme. With the goal of minimizing energy consumption and meeting the needs of user equipment, deep reinforcement learning is used to allocate power to all user equipment. Finally, simulation experiments verify the effectiveness of the algorithm, ultimately achieving a relatively ideal resource allocation method in a short time, ensuring both the lowest overall energy efficiency and the basic needs of user equipment, thus guaranteeing the healthy, safe, reliable, and efficient operation of the power wireless private network.

[0169] Example 2:

[0170] In this invention, we selected four scenarios: three user devices, four user devices, five user devices, and six user devices. Figure 4 The chart shows a comparison of rewards under four different scenarios. Figure 5 The chart shows a comparison of losses under four different scenarios. Figure 6 The graph shows a comparison of energy efficiency under four different conditions.

[0171] The approximate step values ​​and result values ​​at convergence in the four cases are shown in the table below:

[0172] Table 1: Approximate values ​​at convergence

[0173]

[0174] From the above, we can see that Figure 6 The four lines from top to bottom represent the scenarios with six user devices, five user devices, four user devices, and three user devices, respectively. The more user devices there are, the more steps are required for convergence, meaning convergence is slower. This is because a larger number of user devices increases the state space and action space of the DRL, leading to a greater number of convergence steps. Furthermore, the energy efficiency also increases with the number of user devices.

[0175] Next, we selected six user equipment (UE) scenarios and compared the energy efficiency and power consumption values ​​under three different strategies. Here, DA represents the UE selecting the nearest base station to receive service, DRL represents the convergence value of the dynamic high-efficiency resource allocation framework for power wireless private networks (DRL) mentioned in this paper, and UC represents the clustering of any two UE and allocation to one of the base stations for service. The energy efficiency and power consumption values ​​at DRL convergence are the lowest, followed by the nearest-distance (DA) strategy, while the clustering of any two UE (UC) method has the highest energy efficiency and power consumption values. Figure 7 , Figure 8 As shown.

[0176] Next, we compared the user equipment receiving power under these three different strategies, such as Figure 9 As shown. And the cumulative probability distribution of SINR, as shown. Figure 10 As shown. UC represents the strategy for clustering any two user equipments, DA represents the strategy of selecting the closest one, and DRL represents the dynamic high-efficiency resource allocation framework for power wireless private networks based on DRL. Figure 9 In the diagram, the three dashed lines from top to bottom represent the UC, DA, and DRL strategies, respectively. The cumulative probability distribution of the user equipment's received power shows that the probability of the received power being between -50dBm and -20dBm is highest with the DRL method, followed by the DA method, and lastly the UC method. That is, the DRL strategy can relatively maximize the received power of the user equipment. Figure 10 In the cumulative probability distribution of SINR, the probability variation of SINR is largest between 2dB and 4bB, followed by the DA strategy, and finally the UC strategy. This means that the DRL method can relatively minimize the SINR value. All of this demonstrates that the DRL method performs best.

[0177] Example 3:

[0178] like Figure 11 As shown. Based on the same inventive concept, the present invention also provides a system for secure dynamic resource allocation in a power wireless private network, comprising:

[0179] Build modules and learn modules;

[0180] The construction module: uses a convex optimization method to obtain the minimum transmission energy corresponding to each connection relationship in the pre-constructed dynamic high-efficiency resource allocation model of the power wireless private network; and stores each connection relationship, the data flow information in the connection relationship, and the minimum transmission energy into a set memory pool;

[0181] The learning module selects sample data from the memory pool and trains it to obtain the optimal connectivity, optimal power allocation value, and optimal energy efficiency value.

[0182] The dynamic high-efficiency resource allocation model for the power wireless private network is constructed based on the service relationship between each base station and each user equipment in the power wireless private network.

[0183] Preferably, the learning module includes: a selection unit, a calculation unit, and a training unit;

[0184] The selection unit is used to select sample data from the memory pool;

[0185] The computing unit is used to calculate the pre-set objective function and constraints based on each sample data using a convex optimization method, to obtain the optimal power allocation solution and the corresponding Q value table under the connection relationship corresponding to the current state s in the sample data, as well as the target Q value corresponding to the Q value table;

[0186] The training unit is used to train a neural network using a Q-value table and a target Q-value to obtain the optimal connection relationship, the optimal power allocation value, and the optimal energy efficiency value.

[0187] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0188] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0189] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0190] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0191] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for secure dynamic resource allocation in a power wireless private network, characterized in that, include: For each connection relationship in the pre-constructed dynamic high-efficiency resource allocation model of the power wireless private network, the minimum transmission energy corresponding to each connection relationship is obtained by using a convex optimization method; and each connection relationship, the data flow information in the connection relationship, and the minimum transmission energy are stored in a set memory pool. Sample data is selected from the memory pool and trained to obtain the optimal connectivity, optimal power allocation value, and optimal energy efficiency value; specifically including: Select sample data from the memory pool; Based on each sample data, the convex optimization method is used to calculate the pre-set objective function and constraints to obtain the optimal power allocation solution and the corresponding Q value table under the connection relationship corresponding to the current state s in the sample data, as well as the targetQ value corresponding to the Q value table; The optimal connectivity, optimal power allocation, and optimal energy efficiency are obtained by training the neural network using the Q-value table and the target Q-value. The dynamic high-efficiency resource allocation model for the power wireless private network is constructed based on the service relationship between each base station and each user equipment in the power wireless private network. The targetQ value corresponding to the Q value table is calculated using the following formula: Q(s,A)=R+γmax[Q(s',all_actions)]; Where s' is the next state, γ is the reward decay coefficient, and all_actions is all actions; The objective function is calculated as follows: Where t0 is the operation time, E C (s ij ) represents the energy required for node operations. For fixed power consumption, P C To consume power, s ij When base station i serves user equipment j, there is a connection relationship, r f For the required data rate, I is the set of base stations, J is the set of user equipment, F is the set containing all flows on the network link, and c i The maximum scheduling CPU calculation frequency at base station i.

2. The method according to claim 1, characterized in that, The connection relationship between each base station and each user equipment includes: Based on the data rate of the stream carried by the link between the base station and the user equipment, the spectral efficiency of the wireless link is determined by using the Shannon limit. The data rate capacity implemented by the wireless link is determined based on the spectral efficiency of the wireless link. The base station set and the core network are connected via a wired backhaul link, and the user equipment and the base station transmit data via a wireless link.

3. The method according to claim 1, characterized in that, The process involves using a convex optimization method to obtain the minimum wireless transmission energy corresponding to each connection relationship; and storing each connection relationship, the data stream information within the connection relationship, and the minimum wireless transmission energy into a designated memory pool, including: 1) Set up a state space based on the connection relationship; randomly initialize a state from the state space, initialize the memory pool, set the number of observation steps, and set the number of observation steps as the observation value; 2) Based on the current state, select an action, obtain the corresponding reward value, and the state after the action ends. Save the current state, action, reward value, and state after the action ends to the memory pool. 3) Check if the amount of data stored in the memory pool exceeds the observed value. If not, go to 4); otherwise, end. 4) Determine if the maximum number of search steps has been reached. If the maximum number of search steps has been reached, randomly reset a state; otherwise, set the state after the action ends to the current state s and return to step 2).

4. The method according to claim 3, characterized in that, The state space setting based on connection relationships includes: Construct an array M based on the permutations and combinations of connections between all user equipment and each base station, M = {s1, s2, ..., s...} k }; where k is the number of user devices, s1, s2, ..., s k The connection relationship between user equipment {1,2,…,k} and each base station; A state space S is constructed based on the connection relationships formed by arrays of all base stations, where S = {M1, M2, ..., M}. N }; where N is the total number of connections, N = H k H represents the number of base stations.

5. The method according to claim 3, characterized in that, The process of selecting an action based on the current state, obtaining a corresponding reward value, and the state after the action ends includes: The action space is set as A based on the number of connections, where A = {1, 2, ..., N}, where the numbers represent the position in the current state and N represents the total number of connections. The minimum wireless transmission energy in the current state is obtained by using the convex optimization method, and combined with the numbers in the action space A by the greedy policy; The state after the action ends is determined based on the position of the number in the current state; The reward is set to E. max -E, where E max E represents the maximum energy consumption that the current base station can provide, and E represents the energy consumption after taking the action.

6. The method according to claim 5, characterized in that, The formula for calculating the energy consumption value after the action is as follows: Among them, E T P represents the energy expenditure after the action. ij S represents the power allocated by base station i to user equipment j. ij When base station i serves user equipment j, there is a connection relationship, t0 is the operation time, I is the set of all base stations, and J is the set of all user equipment.

7. The method according to claim 3, characterized in that, The data stream information is as follows: in, It is the data rate of stream f carried on base station i and user equipment j. For the streams carried on base station i and user equipment j to have the required data packet size t, 0ij Let be the operation time from base station i to a user equipment j.

8. The method according to claim 2, characterized in that, The data rate capacity achieved by the wireless link is as follows: r ij =x ij B i Y ij Where, r ij x represents the data rate at which user equipment j is served by base station i. ij B represents the allocation ratio of base station i to user equipment j. i γ represents the total available spectrum bandwidth of base station i. ij This represents the spectral effect of the wireless link served by base station i for user j.

9. The method as described in claim 1, characterized in that, The constraints include: Each user equipment can only be served by one base station; The size of the stream carried by the wireless link cannot exceed the data rate capacity implemented by the stream; The total transmit power and allocated bandwidth shall not exceed the total power and total bandwidth provided by the total transmit power and allocated bandwidth; User equipment receives power constraints and signal-to-noise ratio constraints.

10. The method as described in claim 1, characterized in that, After calculating the optimal solution based on the objective function, the process further includes: dynamically allocating bandwidth and power resources based on the optimal solution.

11. A system for secure dynamic resource allocation in a power wireless private network, characterized in that, include: Build modules and learn modules; The construction module: uses a convex optimization method to obtain the minimum transmission energy corresponding to each connection relationship in the pre-constructed dynamic high-efficiency resource allocation model of the power wireless private network; and stores each connection relationship, the data flow information in the connection relationship, and the minimum transmission energy into a set memory pool; The learning module selects sample data from the memory pool and trains it to obtain the optimal connectivity, optimal power allocation value, and optimal energy efficiency value; specifically, it includes: Select sample data from the memory pool; Based on each sample data, the convex optimization method is used to calculate the pre-set objective function and constraints to obtain the optimal power allocation solution and the corresponding Q value table under the connection relationship corresponding to the current state s in the sample data, as well as the targetQ value corresponding to the Q value table; The optimal connectivity, optimal power allocation, and optimal energy efficiency are obtained by training the neural network using the Q-value table and the target Q-value. The dynamic high-efficiency resource allocation model for the power wireless private network is constructed based on the service relationship between each base station and each user equipment in the power wireless private network; the optimal connection relationship and the optimal power allocation value are satisfied. The targetQ value corresponding to the Q value table is calculated using the following formula: Q(s,A)=R+γmax[Q(s',all_actions)]; Where s' is the next state, γ is the reward decay coefficient, and all_actions is all actions; The objective function is calculated as follows: Where t0 is the operation time, E C (s ij ) represents the energy required for node operations. For fixed power consumption, P C To consume power, s ij When base station i serves user equipment j, there is a connection relationship, r f For the required data rate, I is the set of base stations, J is the set of user equipment, F is the set containing all flows on the network link, and c i The maximum scheduling CPU calculation frequency at base station i.

12. The system according to claim 11, wherein the learning module comprises: Selection unit, computation unit, and training unit; The selection unit is used to select sample data from the memory pool; The computing unit is used to calculate the pre-set objective function and constraints based on each sample data using a convex optimization method, to obtain the optimal power allocation solution and the corresponding Q value table under the connection relationship corresponding to the current state s in the sample data, as well as the target Q value corresponding to the Q value table; The training unit is used to train a neural network using a Q-value table and a target Q-value to obtain the optimal connection relationship, the optimal power allocation value, and the optimal energy efficiency value.

Citation Information

Patent Citations

  • A joint optimization method for task unloading and resource allocation in a mobile edge computing network

    CN109814951A

  • Resource allocation and computational unloading joint optimization scheme in mobile edge computing network

    CN110062026A

  • Resource allocation method and device for edge network system and electronic equipment

    CN110248206A

  • Wireless self-backhaul small base station access control and resource allocation joint optimization method

    CN108307511A

  • User bandwidth resource allocation method capable of balancing energy consumption and user service quality

    CN109819522A