Body area network power control method based on hierarchical reinforcement learning

By using a hierarchical reinforcement learning framework, the power control decision-making process of wireless body area networks is decomposed into high-level and low-level layers, which solves the problem of balancing energy saving and reliability caused by severe channel attenuation and complex time-varying characteristics in traditional methods, and achieves efficient power control and energy consumption optimization.

CN121126498APending Publication Date: 2025-12-12BEIJING UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511428519.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-01
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Traditional wireless body area network power control methods struggle to achieve a dynamic optimal balance between energy saving and reliability when faced with severe channel attenuation caused by human activity and complex time-varying characteristics. Single-layer reinforcement learning algorithms suffer from the curse of dimensionality and time scale mismatch in high-dimensional state spaces, resulting in low learning efficiency and difficulty in policy convergence.

Method used

A hierarchical reinforcement learning framework is adopted to decompose the power control decision-making process into two levels: a high level and a low level. The high level is responsible for sensing human activity patterns and macroscopic changes in the channel and selecting a global control strategy, while the low level is responsible for adjusting the specific transmit power on a fast time scale. The power control decision process is optimized in a coordinated manner through a specially designed hierarchical reward function and a Markov decision process (MDP) model.

Benefits of technology

It achieves reduced system energy consumption while ensuring communication quality, improves decision intelligence and environmental adaptability, solves the problems of dimensionality curse and time scale mismatch in traditional single-layer algorithms, and improves learning efficiency and policy stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121126498A_ABST
    Figure CN121126498A_ABST
Patent Text Reader

Abstract

The invention discloses a body area network power control method based on hierarchical reinforcement learning, and belongs to the technical field of wireless communication. A complex power control decision process is decomposed into two levels by introducing a hierarchical reinforcement learning framework. The high-level element strategy is responsible for sensing human body activity modes and channel macroscopic changes on a slow time scale and selecting an optimal global control strategy according to the human body activity modes and the channel macroscopic changes; and according to the low-layer execution strategy, the specific transmitting power is finely adjusted on the fast time scale according to the current instantaneous state and the high-layer instruction. The hierarchical decision-making architecture decomposes a single high-dimensional decision-making problem into two relatively simple sub-problems, and is assisted by a specially designed hierarchical reward function, so that inherent defects of a traditional single-layer algorithm are effectively overcome, and intelligent power adaptive control is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wireless communication, and particularly relates to a sensor node adaptive power control method based on hierarchical reinforcement learning in a wireless body area network. BACKGROUND

[0002] With the wide popularity of human health monitoring and wearable devices, the wireless body area network (WBAN) as a key network technology for realizing continuous acquisition and transmission of personal physiological information plays an irreplaceable role in the fields of remote medical treatment, health management, sports science, etc. The WBAN is usually composed of micro-sensor nodes deployed on the body surface or in the body and a central coordinator, and is formed by short-range wireless communication, and has the characteristics of low power consumption, high portability and high real-time performance.

[0003] However, the WBAN node is usually powered by a battery with limited capacity, and the energy constraint is the core bottleneck restricting the long-term stable operation of the network. The power control technology is one of the most effective means to optimize the network energy consumption and prolong the device endurance by dynamically adjusting the node transmission power on the premise of ensuring the reliability of the communication link. The traditional power control strategy (such as power adjustment based on a fixed threshold or feedback control based on received signal strength) is difficult to cope with the severe channel attenuation and complex time-varying characteristics caused by human activities, and cannot achieve a dynamic optimal balance between energy saving and reliability.

[0004] At present, intelligent algorithms such as reinforcement learning have been introduced to solve this problem, and the optimal power strategy is learned autonomously through interaction with the environment. However, the existing research schemes mostly use single-layer reinforcement learning architecture (such as standard deep Q network algorithm, double-delay deep deterministic policy gradient algorithm, etc.), which generally has the problem of dimension disaster when dealing with the high-dimensional state space of WBAN (including channel, energy, queue, posture and other multi-source information), resulting in low learning efficiency and difficulty in strategy convergence. At the same time, there is a huge difference between the slow time scale of human posture change and the fast time scale of channel rapid fading, and the single-layer agent is difficult to coordinate the multi-time scale decision, and cannot simultaneously consider long-term energy efficiency optimization and instantaneous link quality guarantee. The present application proposes an intelligent power control method based on hierarchical reinforcement learning, which builds a two-layer decision architecture, the high layer can formulate long-term strategy planning according to the network macro state, and the low layer can adjust the data transmission power of the sensor node in time, and through the cooperation of the high-level meta-controller and the low-level option controller, the complex joint decision problem is decomposed into two levels of environment perception and power control, thereby avoiding the inherent problems of single-layer architecture, and improving the decision intelligence, environmental adaptability and energy efficiency balance performance. This design provides a scientific and feasible new idea for solving the dimension disaster and time scale mismatch problems faced by traditional single-layer reinforcement learning algorithms in the application process of this research direction, which can reduce system energy consumption while ensuring communication quality. SUMMARY

[0005] In order to solve the problems of dimension disaster and time scale mismatch faced by the traditional power control method in the existing wireless body area network, the present application proposes an adaptive power control method based on a hierarchical reinforcement learning framework, which aims to realize the collaborative optimization of energy consumption, reliability and service quality in a complex dynamic human body channel environment.

[0006] The core idea of the present application is to decompose the complex power control decision process into two levels by introducing a hierarchical reinforcement learning framework. The high-level meta-strategy is responsible for perceiving human activity patterns and channel macro changes on a slow time scale, and selecting the optimal global control strategy accordingly. The low-level execution strategy adjusts the specific transmission power on a fast time scale based on the current instantaneous state and high-level instructions. This hierarchical decision architecture decomposes a single high-dimensional decision problem into two relatively simple sub-problems, and is supplemented by a specially designed hierarchical reward function, thereby effectively overcoming the inherent defects of traditional single-layer algorithms and realizing intelligent power adaptive control.

[0007] The body area network power control method based on hierarchical reinforcement learning specifically includes the following steps:

[0008] S1: Based on the CM3 and CM4 channel models, an improved path loss model for human body channels is established combined with measured data; the path loss model is expressed as:

[0009]

[0010] in, For frequency Down and node position Related at reference distance =Path loss at 10cm. This is the location- and frequency-dependent path loss exponent. This is zero-mean Gaussian shadowing fading that is attitude- and position-dependent. In order to match human posture Additional losses related to location. This refers to the dynamic loss during attitude transition. This includes losses due to environmental changes, such as temperature-related losses, humidity-related losses, and clothing-related losses.

[0011] S2: To ensure WBAN communication reliability, the overall network energy consumption is minimized by adjusting the transmit power sequence of each node within a time window T in real time. Assume there are N sensor nodes in the network, i∈{1,2,...,N}, and time is discretized into T time slots, t∈{1,2,...,T}. The transmit power levels selected by each sensor node are a discrete set P = { , ,..., }, where M=8; node i selects the transmit power ∈ P and normalized accordingly, the corresponding received signal strength is:

[0012]

[0013] in, For time slots In posture The path loss is as follows. =2dBi and =2dBi represents the transmit and receive antenna gains, respectively. Noise is measured for RSSI.

[0014] Link reliability constraints require that the received signal strength not be lower than a preset threshold:

[0015]

[0016] in, Determined based on the data type of node i; the energy consumption over a period of time is equivalent to the sum of the products of the transmission power and the time.

[0017] Energy constraints ensure long-term network operation:

[0018]

[0019] in Let i be the initial energy of node i. This represents the minimum remaining energy required for node i to function properly.

[0020] QoS constraints take into account the latency requirements of different data types:

[0021]

[0022] in, Let i be the data transmission delay of node i in time slot t. This represents the maximum tolerable delay.

[0023] Under the premise of satisfying the above constraints, the multi-objective optimization problem is formulated as follows:

[0024] Minimize WBAN transmission power consumption:

[0025]

[0026] Maximizing reliability:

[0027]

[0028] Constraints:

[0029]

[0030]

[0031]

[0032]

[0033] in, Let represent the importance weight of node i, and (·)⁺ denote the positive part function.

[0034] S3: Hierarchical Markov Decision Process (MDP) Modeling;

[0035] The high-level MDP focuses on detecting environmental changes and selecting adaptive strategies. Its decision cycle is set at the second level, and it is responsible for monitoring changes in channel statistical characteristics and adjusting lower-level control strategies accordingly. Using a criticality analysis method based on the value function gradient, the gradient norm of each state variable with respect to the Q-value function is calculated to evaluate its criticality, compressing the complete five-dimensional state space of the high-level layer into a three-dimensional vector. = { , (t), }.in The time correlation coefficient of the channel at the current moment is used to quantify the smoothness of channel changes. It is calculated by hysteresis of continuous RSSI measurements. Autocorrelation analysis over 50ms yielded the following results:

[0036]

[0037] The calculation formula in step S2 is used to obtain the result; (t) represents the RSSI measurement value at a certain moment within the time window. = The standard deviation within 1 second reflects the severity of channel fluctuations and is a key indicator of environmental dynamics; The average energy consumption level of the network at the current decision-making moment is represented by an exponentially weighted moving average:

[0038]

[0039] Calculation, where = 0.1 is a smoothing factor used to evaluate the long-term energy efficiency performance of the current strategy. Indicates the time since the last decision. -1 to the current decision time The actual measured average energy consumption within this cycle.

[0040] The high-level action space is defined as three discrete options. = {Hold, Switch, Parameter Tuning}, corresponding to different degrees of policy adjustment. The "Hold" action means the current environment is stable, and the low-level policy does not need to be changed. The "Switch" action is triggered when a significant change in the environment is detected. It requires the low-level actor-critic network based on the TD3 algorithm to switch between pre-trained policies across multiple covered environment modes to adapt to the new environment and ensure the policy reaches its optimal state. The "Parameter Tuning" action is in between. When key indicators such as average reward, policy entropy, and TD error variance show abnormal fluctuations, it adjusts key parameters of the low-level network, including the learning rate and exploration rate, to achieve a balance between adaptability and computational efficiency.

[0041] The lower-level MDP consists of a quintuple (S, A, P, r, γ), where: the state (S) describes the current state of the system and involves a multi-dimensional feature vector, which specifically includes: features related to channel state, network performance, environmental context, and historical performance.

[0042] Action A represents the action that the agent can take in the current state. In this problem, the action is defined as the selection of the transmit power level for each time slot, with the specific action space being: A = {-25, -20, -15, -10, -5, 0} dBm, a total of 6 discrete power levels. The transmit power remains constant within a single communication time slot. The specific mapping method is as follows:

[0043]

[0044] in, This represents the transmit power of the sensor node before normalization. This represents the normalized transmit power. This represents the upper limit of the transmission power, with a lower limit of 0.

[0045] The state transition probability P describes the probability distribution of transitioning to the next state after performing an action in a certain state.

[0046] A model-free reinforcement learning method is adopted; the reward r reflects the immediate reward obtained by the agent after performing an action in a specific state; the discount factor γ is used to adjust the agent's emphasis on long-term rewards when making decisions.

[0047] S4: Hierarchical reinforcement learning architecture design;

[0048] The high-level learning algorithm adopts the Option-Critic method, and the design uses an Actor-Critic low-level network architecture based on the TD3 algorithm. The Actor network is responsible for selecting specific power actions based on the current state and high-level Options, while the Critic network is responsible for evaluating the value of the actions. Stable training is achieved through delayed updates and target policy smoothing.

[0049] Communication between higher and lower layers is achieved through action switching signals. When a higher layer makes an action switching decision, it sends corresponding parameter adjustment instructions to the lower layer. Communication between lower and higher layers relies on performance monitoring metrics. Key metrics such as average reward, policy entropy, and TD error variance are calculated periodically and fed back to the higher layer. When these metrics show abnormal fluctuations, the system immediately reports to the higher layer, triggering the corresponding adjustment mechanism.

[0050] S5: Design of a hierarchical reward function;

[0051] Short-term performance metrics are introduced as supplementary rewards, including improvements in average package success rate and energy efficiency. The complete high-level reward function is as follows:

[0052]

[0053] in This represents the actual measured network lifetime under the current strategy, estimated based on the current power consumption rate and remaining power, expressed as follows:

[0054]

[0055] in Represents a node The remaining energy, This represents the average power consumption of a node over a recent period of time. This represents the total number of nodes in the network.

[0056] The adaptive baseline lifetime is represented by taking the most recent historical window. arrive Using the median of the measured lifetime as a benchmark ensures the performance of the current strategy. It compares the performance of other strategies under similar recent environmental conditions to assess the current network lifetime performance. The specific formula for this variable is:

[0057]

[0058] in Indicates from time arrive Within a time window, the actual measured historical network lifetime data set. The tanh function ensures that the reward value is always within the range [-1, 1], avoiding the reward explosion problem.

[0059] The expression for the package success rate improvement item is:

[0060]

[0061] in The current packet success rate is represented by the following formula, where Represents a node The number of data packets successfully delivered. Represents a node Total number of data packets sent.

[0062]

[0063] This represents the packet success rate of the baseline strategy. This represents the amplification factor, used to enhance sensitivity to changes in packet success rate.

[0064] The energy efficiency improvement item is expressed as follows:

[0065]

[0066] in The current energy efficiency is expressed by the following formula, where Represents a node Number of data bits successfully delivered, Represents a node Energy consumed.

[0067]

[0068] This represents the energy efficiency value of the benchmark strategy. Sensitivity technology for energy efficiency items.

[0069] The adaptive weighting coefficients should be determined based on the characteristics of wireless body area networks in different scenarios. In static scenarios, energy efficiency should be given more importance; in dynamic transition scenarios, environmental adaptability should be given more importance; and in periodic motion scenarios, long-term stability should be given more importance, so as to ensure that the contribution of each indicator matches its importance.

[0070] The low-level reward function focuses on optimizing real-time power control, aiming to minimize energy consumption while satisfying communication quality constraints. Unlike higher-level reward functions, the low-level reward function needs to provide timely feedback signals on a millisecond-level timescale to guide the power adjustment decisions of each node. We adopt a single-objective optimization approach, transforming the multi-objective problem into a constrained single-objective problem.

[0071] The core form of the low-level reward function is:

[0072]

[0073] in This parameter represents a dynamic trade-off between energy consumption and reliability, aiming to enable lower-level agents to autonomously adjust the importance of energy saving based on the wireless body area network status. In practical applications, this is achieved by dynamically adjusting the parameter based on online estimation of the marginal cost of constraint violations. Value. Specifically, when When constraints are frequently violated, the system will decrease. This increases the emphasis on communication quality; when energy consumption is too high, the system will increase... This is to strengthen the energy-saving orientation. The parameter update formula is as follows:

[0074]

[0075] Value range restrictions (clipping function):

[0076]

[0077] Among them, For learning rate, For the target constraint violation rate, The actual constraint violation rate measured at the current moment. Define constraint violation rate:

[0078]

[0079] in To limit the number of violations, This represents the total number of decisions.

[0080] It is an energy urgency function, which is inversely proportional to the remaining energy of the network or node. The lower the remaining energy, the higher the urgency.

[0081]

[0082] in, This represents the critical energy threshold. This represents the lowest node energy at the current moment. This represents the total network energy at the current moment.

[0083] , These represent the importance of constraint violation and energy urgency, respectively. The total power consumption of the network is the sum of the power consumption of each node. By normalizing, the value is compressed to the range of [0,1], so that the reward value will not fluctuate by orders of magnitude due to the absolute value being too large or too small, which is conducive to the stable training of reinforcement learning algorithms.

[0084] (·) represents the smoothed RSSI reward function. This represents the minimum requirement for communication quality. The function transforms a rigid, uninformative communication quality constraint into a continuous, differentiable scoring signal that accurately measures the current network state and clearly guides the agent on how to optimize power. This solves one of the core challenges in power control during reinforcement learning—reward sparsity and lack of gradients—significantly accelerating convergence and improving final performance. The mathematical expression of this function is as follows:

[0085]

[0086] in, . Attached Figure Description

[0087] Figure 1 This is a schematic diagram of the complex channel propagation in the human body.

[0088] Figure 2 This is a flowchart of the implementation of this method.

[0089] Figure 3 This is a simulation result diagram of the low-level cumulative reward curve of the present invention;

[0090] Figure 4 This is a simulation result diagram of the cumulative reward curve for the high-level phases of this invention; Detailed Implementation

[0091] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0092] Figure 1 This is a schematic diagram of the complex channel propagation in the human body.

[0093] Figure 2 This is a flowchart of the implementation of this method.

[0094] The main work of this invention is to systematically verify the effectiveness and feasibility of hierarchical reinforcement learning (HRL) in sensor power control of wireless body area networks (WBANs). On the network side, an end-to-end simulation including MAC / PHY was built using the Castalia-3.2-Augment / OMNeT++ platform, and human body path loss curves from BANmodels were loaded. The channel was simulated by superimposing measurement noise and mild time-dependent perturbations on the path loss to approximate the effects of human activities such as posture changes, occlusion, and reflections. Data processing and visualization were performed using the MATLAB toolkit. Given the periodic characteristics of human movement, this study categorized movement patterns into three types: the first type is static monitoring, using a seated scenario as an example, where the posture is relatively stable and RSSI fluctuations are small; the second type is periodic movement, using a walking scenario as an example, exhibiting approximately periodic channel fluctuations; and the third type is dynamic transitions, such as the transition from a seated to a standing posture, where changes in human activity cause abrupt changes in channel data statistics. Under different motion modes, the channel correlation time is approximately 100ms, 50ms, and 20ms, respectively, with corresponding maximum Doppler frequency shifts of 2Hz, 8Hz, and 15Hz. These quantized parameters can provide a time-scale reference for the adaptive adjustment of the power control algorithm.

[0095] Specifically, the following steps are included:

[0096] S1: Channel and Propagation Modeling

[0097] To accurately model the complex channel propagation environment of the human body, an improved path loss model for the human body channel was established based on the CM3 (Body-to-Body) and CM4 (Body Surface-to-Body Surface) channel models in the IEEE 802.15.6 standard, combined with measured data. The path loss can be expressed as:

[0098]

[0099] in, For frequency Down and node position Related at reference distance The path loss at 10cm represents the path loss at a specific frequency under ideal conditions (ignoring attitude, dynamics, and environmental changes). Signals at specific locations in the human body The basic loss suffered when propagating 10cm. This is a relatively fixed value determined by the equipment and the basic structure of the human body. The path loss index is a location- and frequency-dependent factor, which can be obtained by curve fitting using previously conducted experimental measurement data. Zero-mean Gaussian shading fading (standard deviation) related to attitude and position =3-8dB), the specific value of which can be found in the preset table according to the different node positions and attitude angles. In order to match human posture Location-related additional losses can be determined by querying a pre-established loss mapping table based on different human poses and node positions. The dynamic loss during attitude transition can be calculated in real time based on the human motion acceleration estimated from inertial measurement unit (IMU) sensor data. This includes losses due to environmental changes, such as temperature-related losses, humidity-related losses, and clothing-related losses.

[0100] S2: Problem Modeling; To minimize the overall network energy consumption by adjusting the transmit power sequence of each node within a time window T in real time, while ensuring WBAN communication reliability. This problem is a dynamic multi-objective constrained optimization problem, and the solution process requires finding the optimal balance between factors such as energy consumption, reliability, and service quality. Assume there are N sensor nodes in the network (i∈{1,2,...,N}), time is discretized into T time slots (t∈{1,2,...,T}), and the transmit power levels selected by each node are a discrete set P = { , , ..., }, where M=8, and the power range is {-25,-20,-15,-10,-5,0} dBm. In time slot t, node i selects the transmit power. ∈ P and normalized accordingly, the corresponding received signal strength is:

[0101]

[0102] in, For time slots In posture The path loss is as follows. =2dBi and =2dBi represents the transmit and receive antenna gains, respectively. Noise is measured for RSSI.

[0103] Link reliability constraints require that the received signal strength not be lower than a preset threshold:

[0104]

[0105] in, The value is determined based on the data type of node i: for example, ECG node is -85dBm, blood oxygen node is -80dBm, and temperature node is -75dBm.

[0106] We consider only the energy consumption caused by the data transmission process and ignore other forms of energy loss, that is, the energy consumption over a period of time is equivalent to the sum of the product of the transmission power and the time.

[0107] Energy constraints ensure long-term network operation:

[0108]

[0109] in Let i be the initial energy of node i. This represents the minimum remaining energy required for node i to function properly.

[0110] QoS constraints take into account the latency requirements of different data types:

[0111]

[0112] in, Let i be the data transmission delay of node i in time slot t. The maximum tolerable delay is (e.g., ECG: 100ms, blood oxygen: 500ms, temperature: 5s).

[0113] Under the premise of satisfying the above constraints, the multi-objective optimization problem can be expressed as:

[0114] Minimize WBAN transmission power consumption:

[0115]

[0116] Maximize reliability (maximize link quality margin, and improve the link's anti-interference capability and reliability as much as possible while meeting basic communication requirements):

[0117]

[0118] Constraints:

[0119]

[0120]

[0121]

[0122]

[0123] in, Let represent the importance weight of node i, and (·)⁺ denote the positive part function.

[0124] Due to the dynamic, multi-constraint, and large-scale state space characteristics of the research problem, it is difficult to achieve an effective solution using traditional optimization methods (such as linear programming and greedy algorithms). In contrast, hierarchical reinforcement learning methods decompose the complex problem into two sub-problems: high-level posture adaptation and low-level power optimization. The high-level policy is responsible for detecting changes in human motion state and predicting channel change trends, while the low-level policy performs specific power adjustments based on the current channel state and the guidance from the high-level policy. This hierarchical architecture effectively reduces problem complexity while rapidly adapting to dynamic environments, providing a feasible solution for WBAN power control.

[0125] S3: Modeling of Hierarchical Markov Decision Processes (MDPs)

[0126] In the field of reinforcement learning, Markov decision processes are generally used to describe environmental information. Intelligent agents achieve self-learning and decision optimization by constantly interacting with the environment. In the scenario studied in this paper, the central coordinator is regarded as an intelligent agent.

[0127] The higher-layer MDP focuses on detecting environmental changes and selecting adaptive strategies. Its decision cycle is set at the second level, primarily responsible for monitoring changes in channel statistical characteristics and adjusting lower-layer control strategies accordingly. Through a criticality analysis method based on the value function gradient, the criticality of each state variable is evaluated by calculating the gradient norm of the Q-value function, compressing the complete five-dimensional state space of the higher layer (including packet loss rate and average delivery delay) into a three-dimensional vector. = { , (t), }.in The time correlation coefficient of the channel at the current moment is used to quantify the smoothness of channel changes. It is calculated by hysteresis of continuous RSSI measurements. = Autocorrelation analysis over 50ms yielded the following results:

[0128]

[0129] It can be obtained from the calculation formula in step S2; (t) represents the RSSI measurement value at a certain moment within the time window. = The standard deviation within 1 second (i.e., one high-level decision-making cycle) reflects the severity of channel fluctuations and is a key indicator of environmental dynamics. The average energy consumption level of the network at the current decision-making moment is represented by an exponentially weighted moving average:

[0130]

[0131] Calculation, where = 0.1 is a smoothing factor used to evaluate the long-term energy efficiency performance of the current strategy. Indicates the time since the last decision ( -1) to the current decision time ( The actual measured average energy consumption within this cycle (1 second).

[0132] The high-level action space is defined as three discrete options. = {Hold, Switch, Parameter Tuning}, corresponding to different degrees of policy adjustment. The "Hold" action means the current environment is stable, and the low-level policy does not need to be changed. The "Switch" action is triggered when a significant change in the environment is detected. At this time, the low-level actor-critic network based on the TD3 algorithm needs to switch between pre-trained policies covering multiple main environment modes to adapt to the new environment and ensure the policy reaches its optimal state. The "Parameter Tuning" action is in between. When key indicators such as average reward, policy entropy, and TD error variance show abnormal fluctuations, fine adjustments are made to key parameters of the low-level network, such as learning rate and exploration rate, to achieve a balance between adaptability and computational efficiency.

[0133] The lower-level MDP consists of a quintuple (S, A, P, r, γ), where: the state (S) describes the current state of the system, which involves a multi-dimensional feature vector, specifically including: features related to channel state, network performance, environmental context, and historical performance.

[0134] Action A represents the action the agent can take in the current state. In this problem, the action is defined as the selection of the transmit power level for each time slot, specifically the action space: A = {-25, -20, -15, -10, -5, 0} dBm, with a total of 6 discrete power levels. Within a single communication time slot (5ms), the transmit power remains constant. Based on this, we investigate the normalization of the action space to ensure the stability of agent training, thereby improving the efficiency of the algorithm's policy exploration. The specific mapping method is shown below:

[0135]

[0136] in, This represents the transmit power of the sensor node before normalization. This represents the normalized transmit power. This represents the upper limit of the transmission power, with a lower limit of 0.

[0137] The state transition probability (P) describes the probability distribution of the system transitioning to the next state after performing an action in a given state. This study employs a model-free reinforcement learning method, in which the algorithm learns the optimal policy through direct interaction with the environment, without needing to pre-construct an accurate state transition probability model.

[0138] The reward(r) reflects the immediate reward obtained by the agent after performing an action in a specific state. This reward takes into account factors such as energy efficiency, communication reliability, and service quality requirements. The reward function uses a dynamic weighting mechanism to adaptively adjust the trade-off between energy consumption and reliability based on the current network state, thereby enabling the agent to quickly learn and optimize decision-making strategies.

[0139] The discount factor γ is used to adjust the agent's emphasis on long-term returns when making decisions.

[0140] S4: Hierarchical Reinforcement Learning Architecture Design

[0141] The high-level learning algorithm employs the Option-Critic method, which simultaneously learns both the Option selection strategy and the termination function. The Option selection strategy determines which action mode should be selected in the current network state, while the termination function determines when to terminate the current mode and switch to another mode. By jointly optimizing these three components, the system can learn a suitable hierarchical strategy.

[0142] Because TD3 effectively mitigates overestimation bias through its dual Critic network, improves training stability through delayed policy updates, and enhances policy robustness through target policy smoothing, these characteristics are well-suited to the dynamic environment of WBAN. Therefore, this invention adopts an Actor-Critic low-level network architecture based on the TD3 algorithm. The Actor network is responsible for selecting specific power actions based on the current state and high-level Options, while the Critic network (dual network structure) is responsible for evaluating the value of the actions. Delayed updates and target policy smoothing stabilize training. The key parameters of the TD3 learning algorithm are determined through systematic hyperparameter optimization.

[0143] Communication between higher and lower layers is primarily achieved through action switching signals. When the higher layer makes an action switching decision, it sends corresponding parameter adjustment instructions to the lower layer. Communication between lower and higher layers relies mainly on performance monitoring metrics. The system periodically calculates key metrics such as average reward, policy entropy, and TD error variance, and feeds this information back to the higher layer. When these metrics show abnormal fluctuations, the system immediately reports to the higher layer, triggering the corresponding adjustment mechanism. Given that the channel correlation time is in the range of 10-100 milliseconds, while the human posture change time is in the range of 1-10 seconds, the decision cycle for the higher layer is set to 1 second, and the decision cycle for the lower layer is set to 10 milliseconds, thus forming a timescale ratio of 100:1. According to hierarchical reinforcement learning theory, when the timescale separation ratio between the higher and lower layers is greater than 10, hierarchical learning can achieve effective convergence without significantly affecting performance.

[0144] S5: Design of a hierarchical reward function;

[0145] In the two-layer reinforcement learning architecture of WBAN power control, the design of the reward function faces many challenges: the higher and lower layers have different optimization objectives and time scales, requiring a balance between multi-objective optimization of energy consumption and communication quality, and addressing the reward sparsity problem prevalent in WBAN environments. This section details the layered reward function designed to address these challenges.

[0146] The high-level reward function is designed to guide the system to learn effective environmental adaptation strategies, enabling the entire network to maintain good performance under different application scenarios and environmental conditions. The high-level reward function focuses on long-term performance metrics, with the core being the normalized network lifetime extension rate. To alleviate the reward sparsity caused by long-term metrics, short-term performance metrics are introduced as auxiliary rewards, including improvements in average packet success rate and energy efficiency. The complete high-level reward function is as follows:

[0147]

[0148] in This represents the actual measured network lifetime under the current strategy. It is estimated based on the current power consumption rate and remaining power, and is expressed as follows:

[0149]

[0150] in Represents a node The remaining energy, This represents the average power consumption of a node over a recent period of time. This represents the total number of nodes in the network.

[0151] The adaptive baseline lifetime is represented by taking the most recent historical window ( arrive Using the median of the measured lifetime within a given period as a benchmark, the performance of the current strategy is ensured. This variable is used to compare the performance of other strategies under similar recent environmental conditions to assess the current network lifetime performance. The specific formula for this variable is:

[0152]

[0153] in Indicates from time arrive Within the time window, it is the set of historical network lifetime data actually measured. The tanh function ensures that the reward value is always within the range of [-1, 1], avoiding the reward explosion problem.

[0154] The expression for the package success rate improvement item is:

[0155]

[0156] in The success rate of the current packet can be calculated using the following formula, where Represents a node The number of data packets successfully delivered. Represents a node Total number of data packets sent.

[0157]

[0158] This represents the packet success rate of a baseline strategy (such as the standard deep Q-network algorithm). This represents the amplification factor, used to enhance sensitivity to changes in packet success rate.

[0159] The energy efficiency improvement item is expressed as follows:

[0160]

[0161] in The current energy efficiency can be calculated using the following formula, where Represents a node Number of data bits successfully delivered, Represents a node Energy consumed.

[0162]

[0163] This represents the energy efficiency value of the benchmark strategy. Sensitivity technology for energy efficiency items.

[0164] The adaptive weighting coefficients should be determined based on the characteristics of wireless body area networks in different scenarios. In static scenarios, energy efficiency should be given more importance; in dynamic transition scenarios, environmental adaptability should be given more importance; and in periodic motion scenarios, long-term stability should be given more importance, so as to ensure that the contribution of each indicator matches its importance.

[0165] The low-level reward function focuses on optimizing real-time power control, aiming to minimize energy consumption while satisfying communication quality constraints. Unlike higher-level reward functions, the low-level reward function needs to provide timely feedback signals on a millisecond-level timescale to guide the power adjustment decisions of each node. To simplify the design and improve computational efficiency, we adopt a single-objective optimization approach, transforming the multi-objective problem into a constrained single-objective problem.

[0166] The core form of the low-level reward function is:

[0167]

[0168] in This parameter represents a dynamic trade-off between energy consumption and reliability, aiming to enable lower-level agents to autonomously adjust the importance of energy saving based on the wireless body area network status. In practical applications, we dynamically adjust this parameter by estimating the marginal cost of constraint violations online. Value. Specifically, when When constraints are frequently violated, the system will decrease. This increases the emphasis on communication quality; when energy consumption is too high, the system will increase... This is to strengthen the energy-saving orientation. The parameter update formula is as follows:

[0169]

[0170] Value range restrictions (clipping function):

[0171]

[0172] Among them, For learning rate, For the target constraint violation rate, The actual constraint violation rate measured at the current moment. Define constraint violation rate:

[0173]

[0174] in To limit the number of violations, This represents the total number of decisions.

[0175] It is an energy urgency function, which is usually inversely proportional to the remaining energy of the network or node (the lower the remaining energy, the higher the urgency).

[0176]

[0177] in, This represents the critical energy threshold. This represents the lowest node energy at the current moment. This represents the total network energy at the current moment.

[0178] , These represent the importance of constraint violation and energy urgency, respectively. This represents the total power consumption of the network (i.e., the sum of the power consumption of each node). By normalizing, the value is compressed to the range of [0,1], so that the reward value will not fluctuate by orders of magnitude due to the absolute value being too large or too small, which is beneficial to the stable training of reinforcement learning algorithms.

[0179] (·) represents the smoothed RSSI reward function. This represents the minimum requirement for communication quality. The function transforms a rigid, uninformative communication quality constraint into a continuous, differentiable scoring signal that accurately measures the current network state and clearly guides the agent on how to optimize power. This solves one of the core challenges in power control during reinforcement learning—reward sparsity and lack of gradients—significantly accelerating convergence and improving final performance. The mathematical expression of this function is as follows:

[0180]

[0181] in,

[0182] Figure 3 This is a simulation result diagram of the low-level cumulative reward curve of the present invention;

[0183] Figure 4This is a simulation result diagram of the cumulative reward curve for the high-level phases of this invention;

[0184] In this simulation, the WBAN consists of a star network topology with one coordinator node and five sensor nodes, and the node data generation rate is 8-12 pkt / s. The comparison algorithms selected are representative methods in the current WBAN power control field, including the standard Deep Q-Network (DQN) algorithm and the actor-critic algorithm based on dual-delay deep deterministic policy gradient (TD3). To ensure a fair comparison, we independently tuned the hyperparameters of all comparison algorithms, finding a set of fixed hyperparameter configurations that performed best in the comprehensive scenario, and kept them unchanged throughout all tests.

[0185] This invention uses the cumulative reward curves of low-level and high-level interaction sequences of three algorithms in three scenarios as evaluation indicators to conduct experimental analysis. In order to quantify the learning efficiency and stability, the study defines the interaction sequence that first meets the condition as the convergence point when the relative change of the mean of the sliding window (size of 50 interaction sequences) is lower than the threshold (5%).

[0186] (1) Low-level cumulative reward curve

[0187] In static scenarios, the learning curve of the hierarchical RL algorithm exhibits exponential saturation characteristics, entering a stable region around 600 rounds. In contrast, the actor-critic algorithm shows larger curve oscillations and slower convergence speed, with the DQN algorithm being the slowest.

[0188] In dynamic transition scenarios, all three algorithms exhibit perturbations near the point of environmental change: the hierarchical RL algorithm initially experiences a slight drop in reward, but quickly recovers and reaches a new convergence point. This is because, during environmental changes, its higher-level policies can rapidly guide lower-level algorithms to switch to pre-trained sub-policies adapted to the new environment. The actor-critic algorithm, guided by continuous policy updates and a value function, experiences a slightly larger drop in reward range and a slightly slower recovery compared to the hierarchical RL algorithm. The DQN algorithm, relying on an experience replay buffer, may be affected by outdated experience during environmental changes, making it more susceptible to perturbations, and its convergence can be delayed by more than 1000 rounds. These results demonstrate the buffering effect of the hierarchical RL algorithm structure.

[0189] In the periodic motion scenario, the cumulative reward curves of all three algorithms show an overall upward trend, while also exhibiting visible periodic fluctuations. As learning progresses, the amplitude of these periodic fluctuations decays exponentially. Among them, the hierarchical RL algorithm, by learning and memorizing periodic patterns through high-level policies, achieves stability earlier and exhibits significantly smaller periodic oscillation amplitudes than the other two algorithms, with the largest decay coefficient. The actor-critic algorithm shows a clear upward trend in periodic fluctuations, and through its continuous policy update mechanism, it is more adaptable to periodic changes than the DQN algorithm. The DQN algorithm, relying on experience replay and greedy exploration, has the largest fluctuation amplitude in its cumulative reward curve and is relatively slow to adapt to periodicity. The results show that the hierarchical RL algorithm exhibits faster convergence speed, better stability, and a higher training reward ceiling in all three scenarios. This demonstrates that when facing complex human WBAN environments, the hierarchical RL algorithm can better balance overall network energy consumption and channel reliability quality without sacrificing basic communication requirements (such as minimum received signal strength and maximum allowable latency) to save energy.

[0190] (2) Cumulative reward curve for senior management

[0191] The results show that the cumulative reward curves of the high-level and low-level layers exhibit a highly similar trend, reflected in the following aspects: the relative performance rankings of the three algorithms are consistent; the convergence trends and fluctuation patterns are basically the same; and the performance characteristics are similar in different scenarios. This phenomenon is because the fundamental way to extend network lifetime, which the high-level reward function focuses on, is to optimize power control, which is precisely the core indicator that the low-level reward function focuses on. Therefore, the two reward functions essentially pursue the same optimization direction. Whether viewed from the high-level or low-level perspective, they are optimizing different layers of the same physical system, indirectly confirming the rationality of the hierarchical reinforcement learning architecture design of this invention.

Claims

1. A volume area network power control method based on hierarchical reinforcement learning, characterized in that, Specifically, the following steps are included: S1: Based on the CM3 and CM4 channel models, and combined with measured data, an improved path loss model for the human body channel is established. S2: To ensure WBAN communication reliability, the overall network energy consumption is minimized by adjusting the transmission power sequence of each node within the time window T in real time. S3: Hierarchical Markov Decision Process (MDP) Modeling; The high-level MDP focuses on detecting environmental changes and selecting adaptation strategies. Its decision cycle is set at the second level. It is responsible for monitoring changes in channel statistical characteristics and adjusting the low-level control strategies accordingly. It evaluates the criticality of each state variable by calculating the gradient norm of the Q-value function based on the criticality analysis method of the value function gradient. The state transition probability P describes the probability distribution of transitioning to the next state after performing an action in a certain state; A model-free reinforcement learning method is adopted; the reward r reflects the immediate reward obtained by the agent after performing a certain action in a specific state; the discount factor γ is used to adjust the agent's emphasis on long-term rewards when making decisions; S4: Hierarchical reinforcement learning architecture design; The high-level learning algorithm adopts the Option-Critic method and designs an Actor-Critic low-level network architecture based on the TD3 algorithm. The Actor network is responsible for selecting specific power actions based on the current state and high-level Options, while the Critic network is responsible for evaluating the value of the actions. Training is stabilized through delayed updates and target policy smoothing. Communication between higher and lower layers is achieved through action switching signals. When the higher layer makes an action switching decision, it sends corresponding parameter adjustment instructions to the lower layer. Communication between lower and higher layers is achieved through performance monitoring indicators. Key indicators such as average reward, policy entropy, and TD error variance are calculated periodically and fed back to the higher layer. When these indicators show abnormal fluctuations, immediately report to higher management to trigger the corresponding adjustment mechanism; S5: Layered reward function design; introduces short-term performance indicators as auxiliary rewards, including improvements in average packet success rate and energy efficiency.

2. The volume area network power control method based on hierarchical reinforcement learning according to claim 1, characterized in that, The path loss model is expressed as: , in, For frequency Down and node position Related at reference distance =Path loss at 10cm; The path loss exponent is related to location and frequency. For pose and position-dependent zero-mean Gaussian shading fading; In order to match human posture Location-related additional losses; For attitude transition dynamic loss; This includes losses due to environmental changes, such as temperature-related losses, humidity-related losses, and clothing-related losses.

3. The volume area network power control method based on hierarchical reinforcement learning according to claim 1, characterized in that, Suppose there are N sensor nodes in the network, i∈{1,2,...,N}, and time is discretized into T time slots, t∈{1,2,...,T}. The transmit power level selected by each sensor node is a discrete set P = { , , ..., }, where M=8; node i selects the transmit power ∈ P and normalized accordingly, the corresponding received signal strength is: , in, For time slots In posture The path loss is as follows. =2dBi and =2dBi represents the transmit and receive antenna gains, respectively. Measure noise for RSSI; Link reliability constraints require that the received signal strength not be lower than a preset threshold: , in, Determined based on the data type of node i; the energy consumption over a period of time is equivalent to the sum of the products of transmission power and time; Energy constraints ensure long-term network operation: , in Let i be the initial energy of node i. The minimum remaining energy required for node i to function properly; QoS constraints take into account the latency requirements of different data types: , in, Let i be the data transmission delay of node i in time slot t. The maximum tolerable delay; Under the premise of satisfying the above constraints, the multi-objective optimization problem is formulated as follows: Minimize WBAN transmission power consumption: , Maximizing reliability: Constraints: , , , ,in, Let represent the importance weight of node i, and (·)⁺ denote the positive part function.

4. The volume area network power control method based on hierarchical reinforcement learning according to claim 1, characterized in that, Compress the complete high-level five-dimensional state space into a three-dimensional vector. = { , (t), };in The time correlation coefficient of the channel at the current moment is used to quantify the smoothness of channel changes. It is calculated by hysteresis of continuous RSSI measurements. = Autocorrelation analysis over 50ms yielded the following results: , The calculation formula in step S2 is used to obtain the result; (t) represents the RSSI measurement value at a certain moment within the time window. = The standard deviation within 1 second reflects the severity of channel fluctuations and is a key indicator of environmental dynamics; The average energy consumption level of the network at the current decision-making moment is represented by an exponentially weighted moving average: , calculate, where = 0.1 is a smoothing factor used to evaluate the long-term energy efficiency performance of the current strategy; Indicates the time since the last decision. -1 to the current decision time The actual measured average energy consumption during this period; The high-level action space is defined as three discrete options. = {Hold, Switch, Parameter Tuning}, which correspond to different levels of strategy adjustment; The lower-level MDP consists of a quintuple (S, A, P, r, γ), where: the state S describes the current state of the system and involves a multi-dimensional feature vector, which specifically includes: features related to channel state, network performance, environmental context, and historical performance. Action A represents the action taken by the agent in the current state; the action is defined as the selection of the transmit power level for each time slot, and the specific action space is: A = {-25, -20, -15, -10, -5, 0} dBm, with a total of 6 discrete power levels; the transmit power remains constant within a single communication time slot; the specific mapping method is as follows: ,in, This represents the transmit power of the sensor node before normalization. This represents the normalized transmit power. This represents the upper limit of the transmission power, with a lower limit of 0.

5. The volume area network power control method based on hierarchical reinforcement learning according to claim 1, characterized in that, The complete high-level reward function is as follows: ,in This represents the actual measured network lifetime under the current strategy, estimated based on the current power consumption rate and remaining power, expressed as follows: ,in Represents a node The remaining energy, This represents the average power consumption of a node over a recent period of time. This represents the total number of nodes in the network; The adaptive baseline lifetime is represented by the following formula: ,in Indicates from time arrive Within the time window, the actual measured historical data set of network lifetime; the tanh function ensures that the reward value is always within the range of [-1, 1], avoiding the reward explosion problem; The expression for the package success rate improvement item is: ,in The current packet success rate is represented by the following formula, where Represents a node The number of data packets successfully delivered. Represents a node Total number of data packets sent; , This represents the packet success rate of the baseline strategy. This represents the amplification factor, used to enhance sensitivity to changes in packet success rate; The energy efficiency improvement item is expressed as follows: ,in The current energy efficiency is expressed by the following formula, where Represents a node Number of data bits successfully delivered, Represents a node Energy consumed; , This represents the energy efficiency value of the benchmark strategy. Sensitivity technology for energy efficiency items; The adaptive weighting coefficients should be determined based on the characteristics of the wireless body area network in different scenario tasks; The low-level reward function focuses on optimizing real-time power control, and its design goal is to minimize energy consumption while satisfying communication quality constraints. Unlike the high-level reward function, the low-level reward function needs to provide timely feedback signals on a millisecond time scale to guide the power adjustment decisions of each node. It adopts a single-objective optimization method to transform the multi-objective problem into a constrained single-objective problem.

6. The volume area network power control method based on hierarchical reinforcement learning according to claim 5, characterized in that, The core form of the low-level reward function is: ,in This is a dynamic trade-off parameter between energy consumption and reliability, when When constraints are frequently violated, the system will decrease. This increases the emphasis on communication quality; when energy consumption is too high, the system will increase... The value is intended to strengthen energy conservation; the parameter update formula is as follows: Value range restrictions (clipping function): , among which, among which For learning rate, For the target constraint violation rate, The actual constraint violation rate measured at the current moment; defining the constraint violation rate: ,in To limit the number of violations, Total number of decisions; It is an energy urgency function, which is inversely proportional to the remaining energy of the network or node. The lower the remaining energy, the higher the urgency. ,in, This represents the critical energy threshold. This represents the lowest node energy at the current moment. This represents the total network energy at the current moment; , These respectively represent the importance of constraint violation and energy urgency; This represents the total power consumption of the network, which is the sum of the power consumption of each node. Normalization compresses the value to the range [0,1].

Citation Information

Cited By

  • A method and system for temperature control in a feed pelleting production process

    CN122387227A

  • A method and system for temperature control in a feed pelleting production process

    CN122387227B