Wide-narrow band fusion-based power wireless sensor network resource allocation method and system
By adopting wide-narrow band fusion technology and D3QN algorithm in the power wireless sensor network, the joint optimization of subcarriers and power is solved, and the resource utilization efficiency and user satisfaction are significantly improved.
Patent Information
- Application Number
- CN202510333214.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
AI Technical Summary
In power wireless sensing networks, the existing technology is difficult to effectively solve the resource management challenges in limited spectrum resources, limited node power consumption, complex dynamic wireless environments, and multi-tasking and dynamically changing networks. Especially in large-scale multi-node scenarios, how to achieve efficient communication and reliable resource management has become a key research topic.
The resource allocation method of power wireless sensor network based on wide and narrowband fusion is adopted, and the Markov decision-making process (MDP) model and dual-deep Q network (D3QN) algorithm are constructed to realize the joint optimization of subcarriers and power, and the resource allocation strategy is dynamically adjusted to cope with changes in network state.
It significantly improves the utilization efficiency of spectrum resources, ensures the satisfaction of delay-sensitive narrowband tasks and high-speed broadband tasks, improves the overall throughput and data transmission efficiency of the network, enhances the flexibility and adaptability of the network, reduces energy consumption, and improves user satisfaction.
Smart Images

Figure CN120186789A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a resource allocation method and system for a power wireless sensor network based on narrowband and broadband integration, and belongs to the technical field of power system communication. Background Art
[0002] In recent years, with the rapid development of smart grids and the Internet of Things (IoT), the importance of power wireless sensor networks (PWSN) in power systems has become increasingly prominent. These networks support the operation monitoring, energy efficiency optimization, and intelligent management of power systems by collecting and transmitting real-time data. However, the efficient operation of PWSN needs to overcome challenges such as limited spectrum resources, limited node power consumption, and complex dynamic wireless environments. Especially in the face of large-scale multi-node scenarios, how to achieve efficient communication and reliable resource management has become a key research topic.
[0003] Orthogonal frequency division multiplexing technology has become an important means to improve network access efficiency due to its high spectral efficiency and multipath fading resistance. Existing research shows that jointly optimizing subcarrier allocation and power control can significantly improve system capacity compared to single-dimensional optimization. However, such methods have limitations in many aspects: First, traditional optimization algorithms rely on accurate channel state information (CSI) modeling, and are prone to performance degradation due to model mismatch in time-varying channel environments; Second, in OFDMA systems, cross-layer resource allocation is performed by combining queue states and channel conditions to ensure the QoS requirements of delay-sensitive users. However, in large-scale node scenarios, the high computational complexity of the mixed integer non-linear programming problem is difficult to meet the millisecond-level scheduling requirements of large-scale networks.
[0004] Narrowband and broadband integration technology provides a more flexible resource allocation scheme for different data transmission task requirements by combining narrowband and broadband resources. In this technical architecture, narrowband is usually used for low-rate and delay-sensitive tasks, while broadband is used for high-rate and large-volume data tasks. This resource allocation method can dynamically adjust the bandwidth configuration according to the specific requirements of tasks, optimize the bandwidth utilization rate, and avoid the resource waste that may be caused by traditional single-bandwidth strategies.
[0005] By combining narrowband and broadband resources, the system can not only optimize resource allocation in a static environment, but also ensure the high efficiency and robustness of the system when facing a complex dynamic environment. This makes the narrowband and broadband integration technology have unique application potential in high-demand scenarios such as power wireless sensor networks, can effectively solve the limitations of traditional resource allocation methods in multi-task and dynamically changing networks, and significantly improve the overall performance of the system.
[0006] Although deep reinforcement learning (DRL) has been introduced into the resource allocation field in recent years to cope with environmental dynamics, existing DRL methods still have problems such as low policy exploration efficiency and uncoupled multi-objective trade-offs in reward function design. Therefore, how to achieve multi-objective collaborative optimization of subcarrier-power joint allocation while ensuring the real-time performance of the algorithm remains the core challenge in improving the user satisfaction of power sensing networks. Summary of the Invention
[0007] Object of the Invention: The present invention proposes a resource allocation method and system for power wireless sensor networks based on the integration of narrowband and broadband, comprehensively considering user throughput, system energy efficiency and fairness, and realizing the joint optimization of subcarrier and power allocation in power sensor networks.
[0008] Technical Solution: According to the first aspect of the present invention, there is provided a resource allocation method for power wireless sensor networks based on the integration of narrowband and broadband, including the following steps:
[0009] Step 1: Establish a wireless sensor network model of the power system. The model adopts a single-cell orthogonal frequency division multiplexing architecture, and the base station manages multiple orthogonal subcarrier resource blocks, which are divided into narrowband resource blocks and broadband resource blocks;
[0010] Step 2: Obtain the task data streams of all wireless sensor nodes, and classify the tasks into narrowband tasks and broadband tasks according to the data transmission rate requirements of the tasks. Among them, narrowband tasks are defined as tasks with rate requirements lower than a specified threshold and sensitive to delay, and broadband tasks are defined as tasks with rate requirements higher than the specified threshold;
[0011] Step 3: Construct a Markov decision process MDP model, whose state space includes channel state information CSI, the remaining task data volume of each node, the occupancy of narrowband and broadband resources, bandwidth switching overhead, and task priority, and the action space includes subcarrier resource block allocation, power level adjustment, and narrowband and broadband switching decisions;
[0012] Step 4: Use the double deep Q-network D3QN algorithm to solve the MDP model, and generate an optimal resource allocation strategy through the collaborative training of the main network and the target network. The strategy includes dynamic subcarrier allocation, power control, and narrowband and broadband resource switching;
[0013] Step 5: Execute real-time resource scheduling according to the resource allocation strategy, maximize the throughput of broadband tasks while ensuring the reliable transmission of narrowband tasks, and feedback and optimize the network performance through the reward function.
[0014] Further, the construction of the wireless sensor network model in Step 1 includes:
[0015] Step 1.1: Divide N orthogonal subcarriers into M resource blocks, where each resource block contains K subcarriers and K≥1;
[0016] Step 1.2: Deploy L wireless sensor nodes, which are uniformly distributed within the coverage area of the base station, and whose channel model includes large-scale path loss and small-scale Rayleigh fading;
[0017] Step 1.3: Set the maximum transmit power constraint of each node to P max , and the transmit powers of narrowband resource blocks and broadband resource blocks can be adjusted independently.
[0018] Furthermore, the specific conditions for task classification in Step 2 include:
[0019] The narrowband task satisfies the bit error rate < 1e-5 and the maximum transmission delay ≤ 50 ms, and the number of occupied resource blocks is 1 - 2;
[0020] The broadband task allows a delay ≤ 200 ms, the number of occupied resource blocks ≥ 3, and the throughput requirement is higher than the preset threshold.
[0021] Furthermore, the state space of the MDP model in Step 3 is defined as: state s_t = {H(t), Q_n(t), Q_w(t), R_occ(t), S_hist(t), P_rank(t)},
[0022] where:
[0023] H(t) is the channel gain matrix of all resource blocks at time t, with dimensions L×M, where L is the number of sensor nodes and M is the number of resource blocks;
[0024] Q_n(t) and Q_w(t) are the remaining data volumes of the narrowband tasks and broadband tasks of each node, respectively;
[0025] R_occ(t) is the occupancy status of each resource block, including three states: unoccupied, narrowband occupied, and broadband occupied;
[0026] S_hist(t) is the bandwidth switching record and cumulative switching times of each node in the most recent N_sw times;
[0027] P_rank(t) is the priority weight of each task;
[0028] The action space is defined as: action a_t = {A_sb(t), A_pow(t), A_sw(t)},
[0029] where:
[0030] A_sb(t) is the subcarrier resource block allocation matrix, with dimensions L×M, and the element values are 0 for unallocated and 1 for allocated;
[0031] $A\_pow(t)$ is the transmission power level of each resource block, discretized into three levels: low, medium, and high, corresponding to the power values $P\_low$, $P\_mid$, and $P\_high$;
[0032] $A\_sw(t)$ is the bandwidth switching decision matrix, with a dimension of $L×1$, and the element value of 0 indicates maintaining the current mode, and 1 indicates switching to another mode.
[0033] Furthermore, the implementation of the D3QN algorithm in step 4 includes:
[0034] Step 4.1: Construct a deep neural network including a dueling network architecture, which includes a parallel state value function branch and action advantage function branch;
[0035] Step 4.2: Use the experience replay mechanism to store state transition samples, and randomly extract mini-batch samples from the replay buffer during each training;
[0036] Step 4.3: Calculate the target Q value through the target network, and the update formula is:
[0037]
[0038] where $a$ * is the optimal action selected through the main network, and $\gamma$ is the discount factor;
[0039] Step 4.4: Update the main network parameters using the mean squared error loss function and synchronize the target network parameters regularly.
[0040] Furthermore, the dueling network architecture includes:
[0041] The input layer receives the normalized state vector and maps it to the hidden layer through the fully connected layer;
[0042] The state value function branch outputs a scalar value $V(s)$, and the action advantage function branch outputs an advantage value $A(s,a)$ with the same dimension as the action space;
[0043] The final Q value is calculated as:
[0044]
[0045] Furthermore, the reward function in step 5 is defined as:
[0046]
[0047] where:
[0048] is the amount of data successfully transmitted by node $i$ at time slot $t$;
[0049] The data volume to be transmitted for node i;
[0050] ω i Is the task priority weight, which is positively correlated with P_rank(t);
[0051] Is the bandwidth switching penalty term for node j, and a fixed coefficient α is deducted each time of switching;
[0052] Is the power consumption cost of resource block m, and the coefficient β is used to balance energy efficiency.
[0053] According to the second aspect of the present invention, a power wireless sensor network resource allocation system based on narrowband and broadband integration is provided, including:
[0054] A wireless sensor network model construction module for establishing a wireless sensor network model of the power system. The model adopts a single-cell orthogonal frequency division multiplexing architecture, and the base station manages multiple orthogonal subcarrier resource blocks, and the subcarrier resource blocks are divided into narrowband resource blocks and broadband resource blocks;
[0055] A narrowband and broadband task classification module for obtaining the task data streams of all wireless sensor nodes, and classifying the tasks into narrowband tasks and broadband tasks according to the data transmission rate requirements of the tasks. Among them, the narrowband tasks are defined as tasks with rate requirements lower than the specified threshold and latency-sensitive, and the broadband tasks are defined as tasks with rate requirements higher than the specified threshold;
[0056] An MDP construction module based on narrowband and broadband integration for constructing a Markov decision process MDP model. Its state space includes channel state information CSI, the remaining task data volume of each node, narrowband and broadband resource occupancy, bandwidth switching overhead, and task priority. The action space includes subcarrier resource block allocation, power level adjustment, and narrowband and broadband switching decisions;
[0057] An MDP solving module for solving the MDP model by using the double deep Q network D3QN algorithm, and generating an optimal resource allocation strategy through the collaborative training of the main network and the target network. The strategy includes dynamic subcarrier allocation, power control, and narrowband and broadband resource switching;
[0058] A strategy execution module for performing real-time resource scheduling according to the resource allocation strategy, maximizing the throughput of broadband tasks while ensuring the reliable transmission of narrowband tasks, and feedback optimizing the network performance through a reward function.
[0059] In a third aspect, the present invention further provides an electronic device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the program is executed by the processor, it implements the resource allocation method for the power wireless sensor network based on the integration of narrowband and broadband as described in the first aspect of the present invention.
[0060] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the resource allocation method for the power wireless sensor network based on the integration of narrowband and broadband as described in the first aspect of the present invention.
[0061] Beneficial effects: (1) Through the narrowband and broadband integration technology, the present invention realizes the effective differentiation of narrowband and broadband tasks and the collaborative resource allocation in the power wireless sensor network. This optimization strategy significantly improves the utilization efficiency of spectrum resources, ensures that both delay-sensitive narrowband tasks and high-rate broadband tasks can be satisfied, thereby improving the overall throughput and data transmission efficiency of the network while guaranteeing the quality of service. (2) The present invention adopts the Double Deep Q-Network (D3QN) algorithm, which can adjust the resource allocation strategy in real time and dynamically to cope with the continuous changes in the network state. This intelligent decision-making support mechanism enhances the flexibility and self-adaptability of the network, reduces manual intervention, improves the automation level of network management, and also provides strong technical support for the continuous optimization of network performance, especially suitable for multi-task and high-dynamic power Internet of Things scenarios. (3) The present invention considers power control in the resource allocation process. By optimizing the transmission power, it effectively reduces the energy consumption of network devices and achieves the dual improvement of energy efficiency and performance. This not only helps to extend the service life of network devices and reduce operating costs, but also provides a practical solution for building a green and efficient power wireless sensor network. (4) Compared with the existing single-bandwidth optimization methods, the present invention can ensure the stability of low-rate tasks while making full use of high-bandwidth resources to improve the overall throughput of the network and effectively enhance the user satisfaction of the system. By jointly optimizing subcarrier selection and power allocation and performing narrowband and broadband switching, the system can achieve efficient utilization of resources in a dynamically changing environment. Description of the Drawings
[0062] Figure 1 It is the architecture diagram of the power wireless sensing OFDM network in the embodiment of the present invention;
[0063] Figure 2 It is the flowchart of the resource allocation optimization method for the sensor network based on the integration of narrowband and broadband in the embodiment of the present invention;
[0064] Figure 3 It is the architecture diagram of the D3QN algorithm in the embodiment of the present invention. Specific Embodiments
[0065] The present invention will be further illustrated below in conjunction with the accompanying drawings and specific embodiments.
[0066] The present invention proposes a communication method for a power wireless sensor network based on the integration of wideband and narrowband networks, which is applicable to a single-cell orthogonal frequency division multiplexing (OFDM) network architecture with a single base station and multiple sensor nodes. Figure 1 It is a schematic diagram of the power wireless sensor network architecture. The present invention considers a single-cell OFDM network with a single base station and L sensor nodes. The system adopts a multi-subcarrier block allocation mechanism. Each subcarrier block consists of several subcarriers. The subcarrier set composed of N subcarriers is denoted as The communication time between the nodes and the base station is divided into T time slots. The time slot set is denoted as The system scheduling operates within each time slot, assuming that the channel conditions remain unchanged within a single time slot.
[0067] In order to more accurately optimize resource allocation, when modeling the system, the present invention differentiates and considers the different characteristics of wideband and narrowband. Among them, narrowband is suitable for low-rate and delay-sensitive monitoring tasks, such as power equipment status monitoring, environmental sensing data collection, etc., with stronger anti-interference ability and lower power consumption. Wideband is suitable for high-rate and large-data-volume tasks, such as remote control instruction transmission, large-scale data reporting, etc., and can transmit a large amount of information in a short time. Based on this feature, when modeling, the present invention separately establishes data flow descriptions and resource demand models for wideband and narrowband tasks, and comprehensively considers the coordinated scheduling of the two in the optimization process to improve the overall system performance.
[0068] The idea of the method of the present invention is as follows: establish a wireless sensor network model for the power system, obtain the data flows of all tasks generated by all wireless sensor nodes. The data flows generated by different nodes are regarded as different tasks, and the size of the data flow is regarded as the task volume. In response to the diverse requirements of different tasks in the power wireless sensor network, a coordinated transmission mechanism of wideband and narrowband is adopted. Combining the data transmission requirements of each task and the network resource situation, the system divides the task data flows generated by the sensor nodes into narrowband and wideband tasks according to the rate requirements, and combines subcarrier selection and power control to achieve resource coordinated optimization and efficient and reliable inter-node communication in a multi-task scenario.
[0069] Refer to Figure 2 The resource allocation optimization method for a sensor network based on the integration of wideband and narrowband of the present invention includes the following steps:
[0070] Step 1: Establish a wireless sensor network model for the power system. The model adopts a single-cell orthogonal frequency division multiplexing architecture. The base station manages multiple orthogonal subcarrier resource blocks, and the subcarrier resource blocks are divided into narrowband resource blocks and wideband resource blocks.
[0071] According to an embodiment of the present invention, a wireless sensor network model of a power system is established, which specifically includes the following sub-steps:
[0072] Step 1.1: Divide N orthogonal sub-carriers into M resource blocks, and each resource block contains K sub-carriers, where K≥1;
[0073] Step 1.2: Deploy L wireless sensor nodes, and these nodes are uniformly distributed within the coverage area of the base station. Its channel model includes large-scale path loss and small-scale Rayleigh fading;
[0074] Multiple wireless sensor nodes are deployed around the base station. The base station is responsible for managing the communication resource allocation of each node. By allocating appropriate sub-carrier blocks to each node, it ensures that all nodes can efficiently complete the data transmission task. Figure 2 The processing unit shown in the figure refers to the base station, that is, the method of the present invention is executed on the base station. The base station dynamically adjusts the sub-carrier allocation and power control strategies according to the requirements of the nodes and the channel conditions. In order to optimize the resource utilization rate of the system, the base station obtains the data streams of all wireless sensor nodes communicating with the same base station, and simulates the communication process of each node in the wireless sensor network according to these task data streams.
[0075] In the wide and narrow network fusion network, the uplink communication in which the sensor nodes send monitoring data to the base station is mainly considered. Assuming that the non-line-of-sight channel is mainly large-scale fading and small-scale Rayleigh fading, its channel gain is:
[0076]
[0077] Where: h k,n,t represents the channel gain of node k on sub-carrier block n at time slot t, p is the channel power gain per unit distance, d k is the distance between node k and the base station, q is the path loss exponent, and g is a random variable with an exponential distribution, representing small-scale fading.
[0078] The transmit power of node k on sub-carrier block n at time slot t is p k,n,t , and its transmission rate is expressed according to the Shannon formula as:
[0079]
[0080] Where B is the bandwidth of the sub-carrier block, and σ 2 is the noise power.
[0081] Step 1.3: Set the maximum transmit power constraint of each node to P max , and the transmit powers of narrowband resource blocks and broadband resource blocks can be adjusted independently.
[0082] The allocation of broadband resources and narrowband resources satisfies the following constraints: (1) Each subcarrier block can only be allocated to one sensing node; (2) The total transmission power of each node shall not exceed its power limit; (3) In a single time slot, a node can allocate broadband and narrowband resources simultaneously, but it is necessary to ensure the non-interference of communication.
[0083] Step 2: Obtain the task data streams of all wireless sensing nodes, and classify the tasks into narrowband tasks and broadband tasks according to the data transmission rate requirements of the tasks. Among them, narrowband tasks are defined as tasks with rate requirements lower than the specified threshold and sensitive to delay, and broadband tasks are defined as tasks with rate requirements higher than the specified threshold.
[0084] The narrowband and broadband fusion modes in the present invention include any one of the following: (1) Broadband is used for high data rate transmission, and narrowband is used for low data rate or low power consumption transmission; (2) Broadband and narrowband simultaneously provide communication services for a single task to improve the reliability of the task; (3) Dynamically switch broadband or narrowband resources during the communication process to adapt to the dynamic changes of network resources.
[0085] The data stream includes task data generated by wireless sensing nodes sensing external environmental changes. Such as voltage data, current data, power, frequency data, status data, monitoring data, etc. These data are sent to the power system via the base station.
[0086] Based on the analysis of the historical data of the power sensing network, set the rate demarcation threshold R th (such as 1 Mbps).
[0087] Narrowband task: Required rate R req ≤R th , and it is necessary to meet the delay constraint D max ≤50 ms, bit error rate BER < 10 -5 ;
[0088] Broadband task: Required rate R req >R th , allowable delay D max ≤200 ms, with the goal of maximizing throughput.
[0089] Step 3: Construct a Markov decision process MDP model, whose state space includes channel state information CSI, the remaining task data volume of each node, the occupancy of narrowband and broadband resources, the bandwidth switching overhead, and the task priority, and the action space includes subcarrier resource block allocation, power level adjustment, and narrowband and broadband switching decisions.
[0090] In the power wireless sensor network of the present invention, the base station analyzes the channel state information (CSI) and data transmission requirements of each node, and designs a reasonable subcarrier selection and power allocation scheme to optimize the overall performance of the system. The channel state information of each node is reflected by detecting and recording the channel gain between the node and the base station. According to the distribution of each wireless sensor node, the channel gain of each node on each subcarrier block is calculated using a channel model. The channel gain comprehensively considers factors such as channel quality, path loss, and frequency selectivity, and can accurately reflect the communication ability of the node on a specific subcarrier block. The data transmission requirement is determined by determining the task demand volume of each node in the current communication cycle. In addition, in order to more accurately match the dynamic changes of the system, the priorities of narrowband and broadband tasks and the current network resource situation are also comprehensively considered. While ensuring the reliable transmission of low-rate tasks, the data transmission rate of high-throughput tasks is maximized to improve the overall user experience.
[0091] The present invention models the resource allocation problem as a Markov decision process (MDP), which optimizes the system resource utilization efficiency, improves the communication quality of users, and achieves the goal of maximizing throughput by dynamically adjusting the subcarrier block and power allocation scheme of each node in each time slot.
[0092] The Markov decision process is a tuple that includes a state space, an action space, and a reward function.
[0093] State space S: The state space represents the specific configuration of the system at a certain moment, including the current channel state information of each user, the remaining task data volume of each node, and the occupancy status of narrowband and broadband resources. Each state s ∈ S reflects the resource allocation situation and user demand situation in the network.
[0094] According to the implementation manner of the present invention, the state space is defined as:
[0095] s_t = {H(t), Q_n(t), Q_w(t), R_occ(t), S_hist(t), P_rank(t)}
[0096] Where:
[0097] H(t) is the channel gain matrix of all resource blocks at time t, with a dimension of L × M, where L is the number of sensor nodes and M is the number of resource blocks;
[0098] Q_n(t) and Q_w(t) are the remaining data volumes of the narrowband tasks and broadband tasks of each node, respectively;
[0099] R_occ(t) is the occupancy status of each resource block, including three states: unoccupied, narrowband occupied, and broadband occupied;
[0100] $S\_hist(t)$ is the record of the most recent $N\_sw$ bandwidth switching times and the cumulative switching times for each node;
[0101] $P\_rank(t)$ is the priority weight of each task, with a value range of levels 1 - 5; it is determined by task characteristics and system policies;
[0102] Action space $A$: The action space contains all possible joint sub - carrier selection and power allocation strategies that can be taken in each state, including sub - carrier selection and transmission power for both narrowband and wideband. Each action $a\in A$ corresponds to a specific sub - carrier allocation scheme and the corresponding power control scheme.
[0103] According to the embodiments of the present invention, the definition of the action space is:
[0104] $a_t=\{A\_sb(t),A\_pow(t),A\_sw(t)\}$
[0105] where:
[0106] $A\_sb(t)$ is the sub - carrier resource block allocation matrix, with a dimension of $L\times M$, and the element value of 0 indicates unallocated, and 1 indicates allocated;
[0107] $A\_pow(t)$ is the transmission power level of each resource block, discretized into three levels: low, medium, and high, corresponding to power values $P\_low$, $P\_mid$, and $P\_high$;
[0108] $A\_sw(t)$ is the bandwidth switching decision matrix, with a dimension of $L\times1$, and the element value of 0 indicates maintaining the current mode, and 1 indicates switching to another mode.
[0109] The reward function measures the improvement in system satisfaction after taking a specific action $a$ in a given state $s$. The goal is to maximize the total system satisfaction by selecting appropriate actions (i.e., jointly optimizing sub - carrier and power allocation and narrow - band and wide - band switching).
[0110] According to the embodiments of the present invention, the reward function is defined as:
[0111]
[0112] where:
[0113] is the amount of data successfully transmitted by node $i$ at time slot $t$;
[0114] is the amount of data that node $i$ needs to transmit;
[0115] $\omega$ i is the task priority weight, which is positively correlated with $P\_rank(t)$;
[0116] It is the bandwidth switching penalty term for node j, and a fixed coefficient α is deducted each time of switching.
[0117] It is the power consumption cost of resource block m, and the coefficient β is used to balance energy efficiency.
[0118] Among them, The traffic satisfaction situation can be seen. Define the required traffic of node k at time slot t as q k,t , and its traffic satisfaction rate is:
[0119]
[0120] In a specific scenario, the traffic satisfaction rate may be the first priority guarantee index. Then, by controlling the coefficients α and β, the wide - narrow band switching and power consumption cost can be ignored to maximize the satisfaction of users for traffic transmission.
[0121] The modeling also needs to meet certain constraints to ensure usability. The constraint conditions are as follows:
[0122] (1) Sub - carrier allocation constraint: Define the binary variable x k,n,t to indicate whether node k transmits data on sub - carrier block n at time slot t:
[0123]
[0124] (2) Each sub - carrier block can only be used by one node within any time slot, that is:
[0125]
[0126] (3) Power allocation constraint: The transmit power of node k is limited by the maximum power P max , that is:
[0127]
[0128] Compared with the traditional single - dimension optimization method, when the present invention conducts state modeling, it expands the definition of the state space to more comprehensively describe the resource allocation situation in the wide - narrow band fusion network. Specifically, the state space not only includes the current channel state information (CSI) and the remaining task data volume of each node, but also additionally introduces the occupancy of wide - narrow band resources, the bandwidth switching overhead, and the priority information of transmission tasks, so as to more accurately reflect the dynamic changes of the system. At the same time, the action space is no longer limited to simple sub - carrier allocation and power control, but includes the dynamic switching and joint optimization strategy of wide - narrow band resources, so as to achieve a better allocation scheme in the multi - task and multi - resource collaborative environment.
[0129] Step 4: Solve the MDP model using the Double Deep Q-Network (D3QN) algorithm, and generate an optimal resource allocation strategy through the collaborative training of the main network and the target network. The strategy includes dynamic subcarrier allocation, power control, and wideband / narrowband resource switching.
[0130] The present invention uses the D3QN algorithm to solve the optimization problem, and introduces an experience replay mechanism and a dueling network architecture during the training process to enhance the stability and computational efficiency of policy learning. The overestimation problem of the traditional Q-learning algorithm is reduced through a dual-network structure (main network and target network), and the state value function and the advantage function are combined to improve the accuracy and adaptability of resource allocation decisions.
[0131] Refer to Figure 3 , the D3QN algorithm combines Double Q-Learning and the Dueling Network architecture. By introducing the Double Q-Learning mechanism, D3QN reduces the overestimation bias of the traditional Q-learning, effectively improving the accuracy of Q-value estimation and the stability of policy learning.
[0132] In the D3QN algorithm, two Q-networks - the main network (Online Network) and the target network (Target Network) - are respectively used for the estimation of action values and the calculation of target values. The main network is used to select the optimal action in the current state, while the target network calculates the target Q-value based on the samples in the Experience Replay to update the parameters of the main network. This mechanism effectively reduces the estimation bias in Q-value updates and improves the performance of the algorithm in dynamic environments.
[0133] The Q-network adopts the Dueling architecture, which decomposes the Q-value into a state value function V(s) and an advantage function A(s, a):
[0134]
[0135] Based on this architecture, in each iteration, for each sample (s j , a j , r j , s′ j ), first use the main network to select the optimal action in the current state:
[0136] a t = argmax a Q online (s t , a; θ online ).
[0137] The target network calculates the corresponding target value y through the samples in the experience replay pool for updating the parameters of the main network:
[0138] y t = R t + γ max a′ Q target (s t+1 , a′; θ target ).
[0139] where γ is the discount factor and θ′ is the parameter of the target Q network.
[0140] To minimize the difference between the current Q value and the target Q value, a loss function is defined:
[0141]
[0142] Through backpropagation and the gradient descent method, the parameters θ of the main Q network are iteratively updated:
[0143]
[0144] where α is the learning rate. To ensure the stability of the algorithm, the parameters θ of the main network Q are periodically updated to the parameters θ′ of the target Q network.
[0145] Update the network parameters of the estimated Q neural network through the gradient descent method to minimize the loss function; when the number of training rounds reaches a certain number, samples are drawn from the experience pool parameters, and rewards, value function values, and action function values are calculated from these samples. The network parameters are updated by minimizing the loss function to obtain the final estimated Q neural network and target Q neural network.
[0146] Through the above steps, the network parameters are continuously optimized, and finally, the optimal resource allocation strategy for the integration of wide and narrow networks is obtained.
[0147] Step 5: Perform real-time resource scheduling according to the resource allocation strategy, maximize the throughput of broadband tasks while ensuring the reliable transmission of narrowband tasks, and feedback and optimize the network performance through the reward function.
[0148] According to the present invention, the communication process of wide and narrowband integration is as follows: Initialize the resource pools of broadband and narrowband, set the task requirements and power limits of each node; use narrowband resources to complete the transmission of low data rate tasks; use broadband resources to complete the transmission of high data rate tasks; during the resource allocation process, dynamically adjust the allocation ratio of broadband and narrowband resources through the D3QN algorithm; during the communication process, dynamically switch between broadband or narrowband resources according to the real-time state of the network to improve the overall performance of the system.
[0149] The dynamic switching strategy for wide and narrowband resources includes: when high-priority tasks in the network increase, broadband resources are preferentially allocated to ensure the transmission rate of tasks; when network resources are limited, narrowband resources are preferentially allocated to ensure the minimum communication requirements of nodes; in the multi-node cooperation scenario, the joint transmission of wide and narrowband resources is used to improve the reliability of tasks.
[0150] Based on the same technical concept as the method embodiment, according to another embodiment of the present invention, a power wireless sensor network resource allocation system based on wide and narrowband integration is provided, including:
[0151] A wireless sensor network model construction module, which is used to establish a wireless sensor network model of the power system. The model adopts a single-cell orthogonal frequency division multiplexing architecture, and the base station manages multiple orthogonal subcarrier resource blocks, and the subcarrier resource blocks are divided into narrowband resource blocks and broadband resource blocks;
[0152] A wide and narrowband task classification module, which is used to obtain the task data streams of all wireless sensor nodes, and classify the tasks into narrowband tasks and broadband tasks according to the data transmission rate requirements of the tasks. Among them, narrowband tasks are defined as tasks with rate requirements lower than a specified threshold and sensitive to delay, and broadband tasks are defined as tasks with rate requirements higher than the specified threshold;
[0153] An MDP construction module based on wide and narrowband integration, which is used to construct a Markov decision process MDP model. Its state space includes channel state information CSI, the remaining task data volume of each node, the occupancy of wide and narrowband resources, the bandwidth switching overhead, and the task priority. The action space includes subcarrier resource block allocation, power level adjustment, and wide and narrowband switching decisions;
[0154] An MDP solving module, which is used to solve the MDP model by using the double deep Q network D3QN algorithm, and generate an optimal resource allocation strategy through the collaborative training of the main network and the target network. The strategy includes dynamic subcarrier allocation, power control, and wide and narrowband resource switching;
[0155] A strategy execution module, which is used to perform real-time resource scheduling according to the resource allocation strategy, maximize the throughput of broadband tasks while ensuring the reliable transmission of narrowband tasks, and feedback and optimize the network performance through a reward function.
[0156] It should be understood that the power wireless sensor network resource allocation system based on wide and narrowband integration in the embodiment of the present invention can implement all the technical solutions in the above method embodiment. The functions of its respective functional modules can be specifically implemented according to the methods in the above method embodiment, and the specific implementation process can refer to the relevant descriptions in the above embodiments, which will not be elaborated here.
[0157] Based on the same technical concept as the method embodiments, according to another embodiment of the present invention, there is provided an electronic device, the device comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the programs are executed by the processors, each step of the above-mentioned power wireless sensor network resource allocation method based on narrowband and broadband convergence is implemented.
[0158] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned power wireless sensor network resource allocation method based on narrowband and broadband convergence is implemented.
[0159] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0160] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0161] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0162] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or boxes Figure 1 in one flow or a plurality of flows and / or boxes Figure 1 steps for implementing the functions specified in one box or a plurality of boxes.
Claims
1. A method for allocating resources in a power wireless sensor network based on wide- and narrow-band fusion, characterized in that: The following steps are involved: Step 1: Establish a wireless sensor network model of the power system, wherein the model adopts a single-cell orthogonal frequency division multiplexing architecture, and a base station manages multiple orthogonal subcarrier resource blocks, wherein the subcarrier resource blocks are divided into narrowband resource blocks and broadband resource blocks; Step 2: Obtain the task data streams of all wireless sensor nodes, and classify the tasks into narrowband tasks and broadband tasks according to the data transmission rate requirements of the tasks. Narrowband tasks are defined as tasks with a rate requirement lower than a specified threshold and are delay-sensitive, and broadband tasks are defined as tasks with a rate requirement higher than a specified threshold. Step 3: Construct a Markov decision process MDP model, whose state space includes channel state information CSI, the amount of remaining task data of each node, broadband and narrowband resource occupancy, bandwidth switching overhead and task priority, and the action space includes subcarrier resource block allocation, power level adjustment and broadband and narrowband switching decision; Step 4: Using a dual deep Q network D3QN algorithm to solve the MDP model, and generating an optimal resource allocation strategy through collaborative training of the main network and the target network, the strategy includes dynamic subcarrier allocation, power control, and wide- and narrow-band resource switching; Step 5: Perform real-time resource scheduling according to the resource allocation strategy, maximize the throughput of broadband tasks while ensuring reliable transmission of narrowband tasks, and optimize network performance through reward function feedback.
2. The method according to claim 1, characterized in that The construction of the wireless sensor network model in step 1 includes: Step 1.1: Divide N orthogonal subcarriers into M resource blocks, each resource block contains K subcarriers, where K ≥ 1; Step 1.2: deploy L wireless sensor nodes, the nodes are uniformly distributed in the base station coverage area, and the channel model includes large-scale path loss and small-scale Rayleigh fading; Step 1.3: Set the maximum transmit power constraint of each node to P max , and the transmission power of narrowband resource blocks and broadband resource blocks can be adjusted independently.
3. The method according to claim 1, characterized in that The specific conditions for task classification in step 2 include: The narrowband task meets the bit error rate <1e-5 and the maximum transmission delay ≤50ms, and the number of resource blocks occupied is 1-2; The broadband task allows a latency of ≤200ms, occupies ≥3 resource blocks, and has a throughput requirement higher than the preset threshold.
4. The method according to claim 1, characterized in that: The state space of the MDP model in step 3 is defined as: state s_t = {H(t), Q_n(t), Q_w(t), R_occ(t), S_hist(t), P_rank(t)}, in: H(t) is the channel gain matrix of all resource blocks at time t, with dimension L×M, where L is the number of sensor nodes and M is the number of resource blocks; Q_n(t) and Q_w(t) are the remaining data amounts of narrowband tasks and broadband tasks of each node respectively; R_occ(t) is the occupancy status of each resource block, including unoccupied, narrowband occupied, and broadband occupied. S_hist(t) is the record of the most recent N_sw bandwidth switching of each node and the cumulative number of switching times; P_rank(t) is the priority weight of each task; The action space is defined as: action a_t = {A_sb(t), A_pow(t), A_sw(t)}, in: A_sb(t) is the subcarrier resource block allocation matrix with dimension L×M. The element value 0 indicates unallocated and 1 indicates allocated. A_pow(t) is the transmit power level of each resource block, which is discretized into three levels: low, medium, and high. The corresponding power values are P_low, P_mid, and P_high. A_sw(t) is the bandwidth switching decision matrix with a dimension of L×1. An element value of 0 indicates maintaining the current mode, and a value of 1 indicates switching to another mode.
5. The method according to claim 4, characterized in that The implementation of the D3QN algorithm in step 4 includes: Step 4.1: construct a deep neural network including a dueling network architecture, wherein the network includes a state-value function branch and an action-advantage function branch in parallel; Step 4.2: Use the experience replay mechanism to store state transition samples, and randomly extract mini-batch samples from the replay buffer during each training; Step 4.3: Calculate the target Q value through the target network, and the update formula is: where a * is the optimal action selected by the main network, γ is the discount factor; Step 4.4: Use the mean square error loss function to update the main network parameters and synchronize the target network parameters regularly.
6. The method according to claim 5, characterized in that The duel network architecture includes: The input layer receives the normalized state vector and maps it to the hidden layer through the fully connected layer; The state value function branch outputs a scalar value V(s), and the action advantage function branch outputs an advantage value A(s,a) with the same dimension as the action space; The final Q value is calculated as:
7. The method according to claim 5, characterized in that The reward function in step 5 is defined as: in: is the amount of data successfully transmitted by node i in time slot t; is the amount of data that node i needs to transmit; ω i is the task priority weight, which is positively correlated with P_rank(t); is the bandwidth switching penalty term of node j, and a fixed coefficient α is deducted for each switching; is the power consumption cost of resource block m, and the coefficient β is used to balance energy efficiency.
8. A power wireless sensor network resource allocation system based on wide-narrowband fusion, characterized in that: include: A wireless sensor network model building module is used to build a wireless sensor network model of a power system, wherein the model adopts a single-cell orthogonal frequency division multiplexing architecture, and a base station manages multiple orthogonal subcarrier resource blocks, wherein the subcarrier resource blocks are divided into narrowband resource blocks and broadband resource blocks; The narrowband and wideband task classification module is used to obtain the task data streams of all wireless sensor nodes and classify the tasks into narrowband tasks and broadband tasks according to the data transmission rate requirements of the tasks. Narrowband tasks are defined as tasks with a rate requirement lower than a specified threshold and are delay-sensitive, and broadband tasks are defined as tasks with a rate requirement higher than a specified threshold. The MDP building module based on broadband and narrowband fusion is used to build the Markov decision process MDP model. Its state space includes channel state information CSI, the amount of remaining task data of each node, broadband and narrowband resource occupancy, bandwidth switching overhead and task priority. The action space includes subcarrier resource block allocation, power level adjustment and broadband and narrowband switching decision. An MDP solving module is used to solve the MDP model using a dual deep Q network D3QN algorithm, and generate an optimal resource allocation strategy through collaborative training of a main network and a target network, wherein the strategy includes dynamic subcarrier allocation, power control, and wide- and narrow-band resource switching; The strategy execution module is used to execute real-time resource scheduling according to the resource allocation strategy, maximize the throughput of broadband tasks while ensuring the reliable transmission of narrowband tasks, and optimize network performance through reward function feedback.
9. An electronic device, characterized in that: include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the power wireless sensor network resource allocation method based on wide-narrowband fusion as described in any one of claims 1-7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for allocating resources in a power wireless sensor network based on wide-narrowband fusion as described in any one of claims 1 to 7 is implemented.