Time slot allocation and power control method and device for body area network, equipment and medium

By predicting path loss and optimizing time slot allocation and power control, the information age and energy efficiency problems caused by rapid changes in link state in wireless domain networks are solved, and data timeliness and energy efficiency is improved.

CN120264401APending Publication Date: 2025-07-04SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510354070.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Due to the rapid changes in the link status in the wireless domain network, it is difficult to optimize information age and energy efficiency in a coordinated manner, sensor nodes consume high energy and insufficient data timeliness.

Method used

By obtaining the historical path loss data and information age data of the body domain network, using the trained path loss prediction model and time slot allocation and power control model, predict path loss and optimize time slot allocation and transmission power to realize real-time network state observation and resource scheduling.

Benefits of technology

Reduces data transmission delay, improves data timeliness, reduces energy consumption, and improves energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264401A_ABST
    Figure CN120264401A_ABST
Patent Text Reader

Abstract

The invention relates to a time slot allocation and power control method and device for a body area network, equipment and a medium. The method comprises the following steps: acquiring historical path loss data and information age data of the body area network; according to the historical path loss data and a trained path loss prediction model, obtaining a predicted path loss sequence during data transmission between each sensor node and a sink node in the current superframe; and according to the predicted path loss sequence, the information age data and a trained time slot allocation and power control model, obtaining time slot allocation and power control results of each sensor node in the current superframe, thereby reducing data transmission delay, effectively reducing the information age and enabling the data to have more timeliness. Meanwhile, the transmission power can be adjusted more accurately, unnecessary energy consumption is reduced, and the energy efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of wireless communication networks, and particularly to a method, device, electronic device, and storage medium for time slot allocation and power control in a body area network. Background Art

[0002] A wireless body area network is a special sensor network composed of multiple sensor nodes arranged on the human body surface or implanted in the body. These sensor nodes can detect various physiological data of the human body in real time. Meanwhile, the sink node can organize and coordinate the communication between each sensor node, and is also responsible for integrating and sending the collected data.

[0003] However, the wireless body area network faces some challenges. On the one hand, since sensor nodes are usually small in size, the energy they carry is limited and it is not easy to replace the battery, and the transmission energy is the main energy consumption. Therefore, improving energy efficiency has become the key to extending the working time of sensor nodes. On the other hand, based on the requirements of medical diagnosis and detection, the physiological data collected by the wireless body area network must have high timeliness. Therefore, reducing the age of information and making the data timely are also problems that need to be solved urgently.

[0004] In the related art, due to the inability to adapt to the rapid change of the link state in the wireless body area network scenario, it is difficult to jointly optimize the age of information and energy efficiency in the wireless body area network. Summary of the Invention

[0005] Based on this, the purpose of the present application is to provide a method, device, electronic device, and storage medium for time slot allocation and power control in a body area network, which can reduce the age of information and improve energy efficiency.

[0006] According to the first aspect of the embodiments of the present application, a method for time slot allocation and power control in a body area network is provided, including the following steps:

[0007] Obtain the historical path loss data and age of information data of the body area network; the historical path loss data includes the true path loss measured when each sensor node transmits data to the sink node in several historical superframes before the current superframe, and the age of information data includes the age of information when each sensor node transmits data to the sink node in several historical superframes before the current superframe and the age of information when each sensor node transmits data to the sink node in the current superframe;

[0008] According to the historical path loss data and the trained path loss prediction model, obtain the predicted path loss sequence when each sensor node transmits data to the sink node in the current superframe; the predicted path loss sequence includes the predicted path loss when each sensor node transmits data to the sink node through each transmission time slot;

[0009] Based on the predicted path loss sequence, the age-of-information data, and the trained time slot allocation and power control model, the time slot allocation and power control results of each sensor node in the current superframe are obtained.

[0010] According to the second aspect of the embodiments of the present application, a time slot allocation and power control device for a body area network is provided, including:

[0011] An age-of-information data acquisition module, configured to acquire the historical path loss data and the age-of-information data of the body area network; the historical path loss data includes the true path loss measured when each sensor node transmits data to the sink node in a plurality of historical superframes before the current superframe, and the age-of-information data includes the age of information when each sensor node transmits data to the sink node in a plurality of historical superframes before the current superframe, and the age of information when each sensor node transmits data to the sink node in the current superframe;

[0012] A predicted path loss sequence acquisition module, configured to obtain a predicted path loss sequence when each sensor node transmits data to the sink node in the current superframe according to the historical path loss data and the trained path loss prediction model; the predicted path loss sequence includes the predicted path loss for each sensor node to transmit data to the sink node through each transmission time slot;

[0013] A time slot allocation and power control result acquisition module, configured to obtain the time slot allocation and power control results of each sensor node in the current superframe according to the predicted path loss sequence, the age-of-information data, and the trained time slot allocation and power control model.

[0014] According to the third aspect of the embodiments of the present application, an electronic device is provided, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the method in the first aspect.

[0015] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method in the first aspect are implemented.

[0016] In the embodiments of the present application, historical path loss data and age-of-information data of the body area network are obtained; according to the historical path loss data and the trained path loss prediction model, a predicted path loss sequence when transmitting data between each sensor node and the sink node in the current superframe is obtained; according to the predicted path loss sequence, age-of-information data, and the trained time slot allocation and power control model, the time slot allocation and power control results of each sensor node in the current superframe are obtained. In the embodiments of the present application, based on the trained path loss prediction model, the path loss of each sensor node in the current superframe is predicted; based on the trained time slot allocation and power control model, transmission time slots and transmission power are allocated to each sensor node. Since the path loss can be predicted, the problem in the related art of being unable to adapt to the rapid changes in the link state in the body area network is solved. The trained time slot allocation and power control model can observe the network state of the body area network in real time, optimize the time slot allocation and power control results, thereby reducing data transmission delay, effectively reducing the age of information, and making the data more time-sensitive. At the same time, the transmission power can be adjusted more accurately, reducing unnecessary energy consumption and improving energy efficiency.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application.

[0018] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0019] Figure 1 It is a schematic flowchart of a time slot allocation and power control method for a body area network provided by an embodiment of the present application;

[0020] Figure 2 It is a schematic structural diagram of a wireless body area network provided by an embodiment of the present application;

[0021] Figure 3 It is a structural block diagram of a time slot allocation and power control device for a body area network provided by an embodiment of the present application. Detailed Embodiments

[0022] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be described in further detail below with reference to the accompanying drawings.

[0023] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.

[0024] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of this application. The singular forms "a", "said", and "the" used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0025] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims. In the description of this application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0026] In addition, in the description of this application, unless otherwise specified, "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0027] Please refer to Figure 1 , which is a schematic flowchart of the time slot allocation and power control method for a body area network provided by an embodiment of this application. The time slot allocation and power control method for a body area network provided by the embodiments of this application includes the following steps:

[0028] S10: Obtain the historical path loss data and age of information data of the body area network; the historical path loss data includes the actual path loss measured when each sensor node transmits data to the sink node in several historical superframes before the current superframe, and the age of information data includes the age of information when each sensor node transmits data to the sink node in several historical superframes before the current superframe and the age of information when each sensor node transmits data to the sink node in the current superframe.

[0029] Among them, in a Wireless Body Area Network (WBAN for short), a superframe is used to schedule the medium access between nodes to ensure the effective transmission of data and the stable operation of the network.

[0030] Specifically, each superframe includes a beacon phase and a data transmission phase. The beacon phase includes an uplink beacon phase and a downlink beacon phase. The data transmission phase includes a number of transmission time slots and an ACK period. In the uplink beacon phase, each sensor node reports its own status to the sink node in turn and completes time synchronization. In the downlink beacon phase, the sink node informs each sensor node of the transmission time slots and transmission power allocated to each sensor node. In the data transmission phase, each sensor node transmits physiological data to the sink node through the corresponding transmission time slot. After receiving the physiological data, the sink node replies with an ACK frame to each sensor node during the ACK period.

[0031] Among them, path loss is used to measure the power attenuation of the communication link between each sensor node and the sink node during the transmission time slot. For example, when the communication environment of this transmission time slot is good, the path loss of this transmission time slot is small; when the communication environment of this transmission time slot is poor, the path loss of this transmission time slot is large. And path loss will affect the result of time slot allocation, because during the time slot allocation process, the transmission time slot with a small path loss will be selected as much as possible for data transmission. The smaller the path loss, the smaller the transmission power required for the sensor node to transmit data, so as to improve the network energy efficiency.

[0032] Among them, the age of information refers to the time elapsed from the generation of data to its reception at the destination. Specifically, the age of information is the time elapsed since the data newly received by the sink node from each sensor node was generated.

[0033] In the embodiment of the present application, please refer to Figure 2 , which is a schematic structural diagram of a wireless body area network. The wireless body area network consists of a sink node and several sensor nodes. The sink node serves as the central hub of the entire network. The sink node is usually arranged at the torso position of the human body and is responsible for coordinating and processing the data from sensor nodes in all directions. The sensor nodes are dispersedly arranged at various key parts of the human body, such as the head, heart, left arm, left hand, left leg, left foot, right arm, right hand, right leg, and right foot, etc. Each node continuously monitors the physiological data of the human body. These sensor nodes adopt wireless communication technology compliant with the IEEE 802.15.6 standard.

[0034] It can be that each sensor node sends a probe packet to the sink node. By detecting the received signal strength when the sink node receives the probe packet, the real path loss when transmitting data between each sensor node and the sink node can be obtained. Among them, the probe packet is a data packet used to detect and diagnose the network status.

[0035] It is also possible that each sensor node sends data packets to the sink node. By detecting the received signal strength when the sink node receives the data, the path loss of data transmission between each sensor node and the sink node can be obtained. Among them, the data packet is obtained by packing the physiological data collected by the sensor node.

[0036] Each sensor node sends data packets to the sink node through transmission time slots. When each sensor node generates a data packet, a timestamp is attached to record the generation time of the data packet. When the sink node receives the data packet, it records the reception time. Based on the reception time and the generation time, the age of information of data transmission between each sensor node and the sink node is obtained.

[0037] S20: According to the historical path loss data and the trained path loss prediction model, obtain the predicted path loss sequence when data is transmitted between each sensor node and the sink node in the current superframe; the predicted path loss sequence includes the predicted path loss of each sensor node transmitting data to the sink node through each transmission time slot.

[0038] Among them, the path loss prediction model is a temporal convolutional network model, including a causal convolution mechanism and a dilated convolution mechanism. The causal convolution mechanism can ensure the causality of information flow, so that the model output only depends on the input at the current and previous moments; the dilated convolution mechanism expands the model receptive field, enabling the convolution to cover more extensive information.

[0039] In the embodiment of the present application, the historical path loss data is input into the trained path loss prediction model to obtain the predicted path loss of each sensor node transmitting data to the sink node through each transmission time slot in the current superframe. Specifically, taking N sensor nodes and M transmission time slots as an example, the predicted path loss sequence of the nth sensor node output by the trained path loss prediction model is: Among them, represents the predicted path loss of the nth sensor node transmitting data to the sink node through the mth transmission time slot.

[0040] S30: According to the predicted path loss sequence, the age of information data, and the trained time slot allocation and power control model, obtain the time slot allocation and power control results of each sensor node in the current superframe.

[0041] Among them, the trained time slot allocation and power control model is a deep reinforcement learning model, which is used to allocate communication resources for each sensor node. The communication resources include transmission time slots and transmission power.

[0042] In the embodiment of the present application, the age of information of data transmission between each sensor node and the sink node in several historical superframes before the current superframe and the age of information of data transmission between each sensor node and the sink node in the current superframe are averaged to obtain the average age of information. The predicted path loss sequence and the average age of information are input into the trained time slot allocation and power control model to obtain the transmission time slots and transmission powers of each sensor node in the current superframe. Each sensor node transmits data to the sink node at the transmission power in the transmission time slot, which can reduce the average age of information and transmission energy loss of the entire body area network.

[0043] Applying the embodiment of the present application, by obtaining the historical path loss data and age of information data of the body area network; according to the historical path loss data and the trained path loss prediction model, obtaining the predicted path loss sequence of data transmission between each sensor node and the sink node in the current superframe; according to the predicted path loss sequence, age of information data and the trained time slot allocation and power control model, obtaining the time slot allocation and power control results of each sensor node in the current superframe. The embodiment of the present application predicts the path loss of each sensor node in the current superframe based on the trained path loss prediction model; based on the trained time slot allocation and power control model, allocates transmission time slots and transmission powers for each sensor node. Since the path loss can be predicted, the problem that the related technology cannot adapt to the rapid change of the link state in the body area network is solved. The trained time slot allocation and power control model can observe the network state of the body area network in real time, optimize the time slot allocation and power control results, thereby reducing the data transmission delay, effectively reducing the age of information, and making the data more timely. At the same time, the transmission power can be adjusted more accurately, reducing unnecessary energy consumption and improving the energy efficiency.

[0044] In one embodiment, the step of obtaining the path loss data of the body area network in step S10 includes steps S101 to S103, which are specifically as follows:

[0045] S101: In the uplink beacon phase of each superframe, obtain the probe packets transmitted by each sensor node to the sink node.

[0046] In the embodiment of the present application, the uplink beacon phase is divided into several time intervals, each sensor node corresponds to a time interval, and each sensor node transmits a probe packet to the sink node within the corresponding time interval.

[0047] S102: Obtain the received signal strength according to the probe packet.

[0048] Among them, the Received Signal Strength Indicator (RSSI) represents the strength of the received wireless signal. In wireless communication, the RSSI value is usually expressed in negative dBm. The closer the value is to zero, the stronger the signal; conversely, the weaker the signal.

[0049] In the embodiment of the present application, when the aggregation node receives a probe packet, it detects the strength of the wireless signal of the probe packet and obtains the corresponding received signal strength.

[0050] S103: Obtain the path loss data of the body area network according to the received signal strength and the preset path loss calculation model.

[0051] Among them, the path loss calculation model is a mathematical model for calculating the signal strength attenuation caused by various factors (such as distance, obstacles, environment, etc.) during the propagation of wireless signals. In a body area network (BAN) or similar application scenarios, the path loss calculation model needs to comprehensively consider various loss factors, including free space loss and human body loss, etc. The specific calculation formula of the path loss calculation model is prior art and will not be elaborated here.

[0052] In the embodiment of the present application, when the received signal strength when the aggregation node receives the probe packets sent by each sensor node is input into the preset path loss calculation model, the path loss of data transmission between each sensor node and the aggregation node can be obtained.

[0053] In one embodiment, before step S20, it includes steps S201 to S204, which are specifically as follows:

[0054] S201: Obtain a training data set; the training data set includes a first path loss sequence and a second path loss sequence; the first path loss sequence includes the true path losses measured when each sensor node transmits data to the aggregation node in several historical superframes, and the second path loss sequence includes the true path losses measured when each sensor node transmits data to the aggregation node in the current superframe.

[0055] In the embodiment of the present application, each sensor node transmits data to the aggregation node in the allocated transmission time slot in each historical superframe. For the communication link between the nth sensor node and the aggregation node, the first path loss sequence is expressed as: {x n,1 , x n,2 , …, x n,i , …, x n,I}. Among them, I is the number of historical superframes, and x n,i represents the true path loss measured when the nth sensor node transmits data in the ith historical superframe. The second path loss sequence is expressed as: {y n,1 , yn,2 , …, y n,m , …, y n,M}, where M is the number of transmission time slots, and y n,m represents the true path loss measured when the nth sensor node transmits data through the mth transmission time slot.

[0056] S202: Input the first path loss sequence into the path loss prediction model to be trained, and obtain the first predicted path loss sequence.

[0057] In an embodiment of the present application, the first path loss sequence is used as the input data of the path loss prediction model to be trained, and the path loss prediction model to be trained outputs the predicted path loss for data transmission between each sensor node and the sink node in the current superframe.

[0058] S203: Obtain the first loss function value according to the first predicted path loss sequence, the second path loss sequence, and a preset first loss function.

[0059] Among them, the preset first loss function can be set according to actual needs. Specifically, the first loss function can be a root mean square error function.

[0060] In an embodiment of the present application, the expression of the first loss function value is as follows:

[0061]

[0062] S204: Train the path loss prediction model to be trained according to the first loss function value, and obtain the trained path loss prediction model.

[0063] In an embodiment of the present application, according to the first loss function value, use the gradient descent algorithm to continuously optimize the network parameters of the path loss prediction model, and obtain the trained path loss prediction model.

[0064] In one embodiment, before step S30, steps S31 to S36 are included, specifically as follows:

[0065] S31: Obtain the current network state data of each sensor node in the body area network; the current network state data includes the current age of information, the first average age of information, and the predicted path loss for data transmission between the sensor node and the sink node in the current superframe; the current age of information is the age of information for data transmission between the sensor node and the sink node in the current superframe, and the first average age of information is the average of the age of information for data transmission between the sensor node and the sink node in the current superframe and the ages of information in several historical superframes before the current superframe.

[0066] In the embodiment of the present application, a state vector is used to describe the current real-time network state of each sensor node in a wireless body area network environment. Taking the nth sensor node as an example, the current network state data is expressed as: where, Δ n,m represents the age of information of the data packet from the nth sensor node to the sink node in the mth transmission time slot of the current superframe; represents the average age of information of the data packets from the nth sensor node saved by the sink node in the current network state, which reflects the long-term performance of data transmission; represents the predicted path loss between the nth sensor node and the sink node in the mth transmission time slot of the current superframe.

[0067] S32: Input the current network state data into the time slot allocation and power control model to be trained, and obtain the current time slot allocation and power control results of each sensor node.

[0068] In the embodiment of the present application, the current network state data is input into the time slot allocation and power control model to be trained, and the time slot allocation and power control model to be trained outputs the transmission time slots and transmission powers that each sensor node needs to be allocated.

[0069] S33: Obtain the next network state data of the current network state data according to the current time slot allocation and power control results.

[0070] In the embodiment of the present application, each sensor node transmits data to the sink node according to the allocated transmission time slots and transmission powers. Since each sensor node executes new transmission time slots and transmission powers, the network state data of each sensor node will also change, so as to obtain the next network state data of the current network state data. Specifically, the next network state data of the current network state data is expressed as: where, Δ n,p represents the age of information of the data packet from the nth sensor node to the sink node in the pth transmission time slot of the current superframe; represents the average age of information of the data packets from the nth sensor node saved by the sink node in the next network state of the current network state; represents the predicted path loss between the nth sensor node and the sink node in the pth transmission time slot of the current superframe.

[0071] S34: Obtain a reward value according to the first average age of information in the current network state data, the first average age of information of the previous network state data of the current network state data, and the current time slot allocation and power control results.

[0072] Among them, the reward value is used to quantify the quality of the slot allocation and power control results output by the slot allocation and power control model.

[0073] In the embodiment of the present application, the reward value can be obtained according to the change amount of the first average age of information of two adjacent network states and the proportion of the transmission power within the transmission power range.

[0074] S35: Store the current network state data, the current slot allocation and power control results, the next network state data, and the reward value as a set of sample data in the experience pool.

[0075] Among them, the experience pool is used to store the experience data generated by the interaction between the reinforcement learning model and the environment.

[0076] In the embodiment of the present application, the current network state data, the current slot allocation and power control results, the next network state data, and the reward value are stored as a set of sample data in the experience pool in the form of a quadruple for subsequent model training.

[0077] S36: When the number of groups of sample data in the experience pool is greater than or equal to a preset number, train the slot allocation and power control model to be trained according to the sample data in the experience pool to obtain a trained slot allocation and power control model.

[0078] Among them, the preset number can be set according to actual needs.

[0079] In the embodiment of the present application, when the experience data in the experience pool reaches the minimum training batch size, start training the model. Specifically, randomly extract a batch of experience data from the experience pool, and update the network parameters of the slot allocation and power control model according to the experience data to obtain a trained slot allocation and power control model.

[0080] In one embodiment, step S34 includes steps S341 to S345, specifically as follows:

[0081] S341: Calculate the difference between the first average age of information of each sensor node in the current network state data and the first average age of information of the corresponding sensor node in the previous network state data of the current network state data to obtain a number of differences;

[0082] S342: Sum the number of differences to obtain a first summation result; take the opposite of the first summation result to obtain a first reward.

[0083] In the embodiment of the present application, the expression of the first reward is:

[0084]

[0085] Among them, Denotes the average age of information of the data packet from the nth sensor node saved by the sink node at the previous network state of the current network state.

[0086] S343: Obtain the current transmission power from the current time slot allocation and power control results; calculate the first difference between the current transmission power and the first preset transmission power, and the second difference between the second preset transmission power and the first preset transmission power; divide the first difference by the second difference to obtain the second reward.

[0087] Wherein, the first preset transmission power is the minimum transmission power within the transmission power range, and the second preset transmission power is the maximum transmission power within the transmission power range.

[0088] In the embodiments of the present application, the expression of the first reward is:

[0089]

[0090] Wherein, represents the current transmission power of the nth sensor node in the mth transmission time slot of the current network state, p min Represents the first preset transmission power, p max Represents the second preset transmission power.

[0091] S344: Calculate the third difference between the second preset transmission power and the current transmission power; divide the third difference by the second difference to obtain a ratio; obtain the third reward according to the ratio, the first average age of information in the current network state data, and the preset reward function.

[0092] In the embodiments of the present application, the expression of the third reward is:

[0093]

[0094] Wherein, the preset reward function Γ represents that when the first average age of information in the current network state data Is less than or equal to the preset threshold, the value is 1; when Is greater than the preset threshold, the value is 0.

[0095] S345: Obtain the reward value according to the first reward, the second reward, and the third reward.

[0096] In the embodiments of the present application, according to the specific model training requirements, the first reward function, the second reward, or the third reward can be used as the reward value in the model training process. Specifically, in order to reduce the age of information, the first reward can be used as the reward value. In order to improve energy efficiency, the second reward can be used as the reward value. In order to simultaneously reduce the age of information and improve energy efficiency, the third reward can be used as the reward value.

[0097] In the embodiments of the present application, during the model training process, different reward values can be designed based on different optimization objectives, thereby improving the robustness of the model.

[0098] In one embodiment, the time slot allocation and power control model to be trained includes a first deep Q-network and a second deep Q-network. Step S35 includes steps S351 to S353, which are specifically as follows:

[0099] S351: When the number of sample data groups in the experience pool is greater than or equal to a preset number, extract a plurality of sample data from the experience pool, input the plurality of sample data into the first deep Q-network to obtain a first predicted Q value; input the plurality of sample data into the second deep Q-network to obtain a second predicted Q value; obtain a target Q value according to the second predicted Q value and the reward value in the sample data.

[0100] In the embodiments of the present application, first initialize the first deep Q-network and the second deep Q-network, assign the same network parameters to them, and then input a plurality of sample data into the first deep Q-network and the second deep Q-network respectively. Specifically, the expression of the target Q value is as follows:

[0101]

[0102] where r t+1 represents the immediate reward value, γ represents the discount factor, Q(s t+1 , a; θ) represents the first predicted Q value, represents the second predicted Q value.

[0103] S352: Obtain a second loss function value according to the first predicted Q value, the target Q value, and a preset second loss function.

[0104] In the embodiments of the present application, the expression of the second loss function value is as follows:

[0105]

[0106] S353: Train the first deep Q-network according to the second loss function value until the second loss function value meets a preset threshold, and use the trained first deep Q-network as the trained time slot allocation and power control model; wherein, when the network parameters of the first deep Q-network are updated every preset number of times, copy the network parameters of the first deep Q-network to the second deep Q-network.

[0107] In the embodiments of the present application, according to the second loss function value, use the gradient descent algorithm to update the network parameters of the first deep Q-network until the second loss function value meets a preset threshold, and use the trained first deep Q-network as the trained time slot allocation and power control model. Specifically, the update formula of the network parameters is:

[0108]

[0109] During the training process of the first deep Q-network, when the network parameters of the first deep Q-network are updated every preset number of times, the network parameters of the first deep Q-network are copied to the second deep Q-network to update the network parameters of the second deep Q-network, thereby updating the target Q value to assist the training of the first deep Q-network, ensuring the stability and effectiveness of the training, and enabling the time slot allocation and power control model to gradually learn the optimal resource scheduling strategy.

[0110] In one embodiment, step S30 includes steps S301 to S305, which are specifically as follows:

[0111] S301: Obtain the first preset transmission time slots allocated to each sensor node in the current superframe and the second preset transmission time slots allocated to each sensor node in several historical superframes before the current superframe.

[0112] Among them, the first preset transmission time slot and the second preset transmission time slot can be set according to actual needs.

[0113] In the embodiment of the present application, in the current superframe, obtain a first preset transmission time slot allocated by the sink node to each sensor node. In each historical superframe, obtain a second preset transmission time slot allocated by the sink node to each sensor node.

[0114] S302: Extract the corresponding predicted path loss from the predicted path loss sequence according to the first preset transmission time slot.

[0115] In the embodiment of the present application, the predicted path loss sequence includes the predicted path loss of each sensor node in each transmission time slot. By matching the first preset transmission time slot corresponding to each sensor node with each transmission time slot in the predicted path loss sequence, the corresponding predicted path loss can be obtained.

[0116] S303: Obtain the age of information of each sensor node transmitting data through the second preset transmission time slot in several historical superframes before the current superframe and the age of information of each sensor node transmitting data through the first preset transmission time slot in the current superframe.

[0117] In the embodiment of the present application, the process of obtaining the age of information is the same as the aforementioned step S10 and will not be elaborated here.

[0118] S304: Calculate the average value of the age of information of each sensor node transmitting data through the second preset transmission time slot in several historical superframes before the current superframe and the age of information of each sensor node transmitting data through the first preset transmission time slot in the current superframe to obtain the second average age of information.

[0119] In the embodiment of the present application, the age of information of each sensor node transmitting data through the second preset transmission time slot in several historical superframes before the current superframe is summed with the age of information of each sensor node transmitting data through the first preset transmission time slot in the current superframe, and the sum result is divided by the number of ages of information to obtain the second average age of information.

[0120] S305: Input the corresponding predicted path loss, the age of information of each sensor node transmitting data in the current superframe, and the second average age of information into the trained time slot allocation and power control model to obtain the time slot allocation and power control results of each sensor node in the current superframe.

[0121] In the embodiment of the present application, by using the trained time slot allocation and power control model, based on the predicted path loss, the age of information, and the second average age of information of each sensor node in the current superframe, the transmission time slots and transmission powers to be allocated for each sensor node in the current superframe are obtained.

[0122] The following is an embodiment of the apparatus of the present application, which can be used to execute the content of the method in the embodiment of the present application. For details not disclosed in the embodiment of the apparatus of the present application, please refer to the content of the method in the embodiment of the present application.

[0123] Please refer to Figure 3 , which shows a schematic structural diagram of a time slot allocation and power control device for a body area network provided by an embodiment of the present application. The time slot allocation and power control device 4 for a body area network provided by an embodiment of the present application includes:

[0124] An age of information data acquisition module 41, configured to acquire historical path loss data and age of information data of the body area network; the historical path loss data includes the true path loss measured when each sensor node transmits data to the sink node in several historical superframes before the current superframe, and the age of information data includes the age of information when each sensor node transmits data to the sink node in several historical superframes before the current superframe and the age of information when each sensor node transmits data to the sink node in the current superframe;

[0125] A predicted path loss sequence acquisition module 42, configured to obtain a predicted path loss sequence when each sensor node in the current superframe transmits data to the sink node according to the historical path loss data and the trained path loss prediction model; the predicted path loss sequence includes the predicted path loss of each sensor node transmitting data to the sink node through each transmission time slot;

[0126] The time slot allocation and power control result obtaining module 43 is configured to obtain the time slot allocation and power control results of each sensor node in the current superframe according to the predicted path loss sequence, the age of information data, and the trained time slot allocation and power control model.

[0127] It should be noted that when the time slot allocation and power control device of the body area network provided in the above embodiment executes the time slot allocation and power control method of the body area network, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the time slot allocation and power control device of the body area network provided in the above embodiment and the time slot allocation and power control method of the body area network belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.

[0128] This application also provides an electronic device, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the method steps of the above embodiment.

[0129] This application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method steps of the above embodiment are implemented.

[0130] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0131] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A time slot allocation and power control method for a body area network, characterized in that It includes the following steps: Obtain the historical path loss data and age-of-information data of the body area network; the historical path loss data includes the true path loss measured when data is transmitted between each sensor node and the sink node in several historical superframes before the current superframe, and the age-of-information data includes the age of information when data is transmitted between each sensor node and the sink node in several historical superframes before the current superframe, and the age of information when data is transmitted between each sensor node and the sink node in the current superframe; According to the historical path loss data and the trained path loss prediction model, obtain the predicted path loss sequence when data is transmitted between each sensor node and the sink node in the current superframe; the predicted path loss sequence includes the predicted path loss when each sensor node transmits data to the sink node through each transmission time slot; According to the predicted path loss sequence, the age-of-information data, and the trained time slot allocation and power control model, obtain the time slot allocation and power control results of each sensor node in the current superframe.

2. The time slot allocation and power control method for the body area network according to claim 1, wherein: The step of obtaining the path loss data of the body area network includes: In the uplink beacon phase of each superframe, obtain the probe packets transmitted from each sensor node to the sink node; According to the probe packets, obtain the received signal strength; According to the received signal strength and the preset path loss calculation model, obtain the path loss data of the body area network.

3. The time slot allocation and power control method for the body area network according to claim 1, wherein: Before the step of obtaining the predicted path loss sequence when each sensor node transmits data in the current superframe according to the path loss data and the trained path loss prediction model, it includes: Obtain the training data set; the training data set includes the first path loss sequence and the second path loss sequence; the first path loss sequence includes the true path loss measured when data is transmitted between each sensor node and the sink node in several historical superframes, and the second path loss sequence includes the true path loss measured when data is transmitted between each sensor node and the sink node in the current superframe; Input the first path loss sequence into the path loss prediction model to be trained, and obtain the first predicted path loss sequence; According to the first predicted path loss sequence, the second path loss sequence, and the preset first loss function, obtain the first loss function value; According to the first loss function value, train the path loss prediction model to be trained to obtain the trained path loss prediction model.

4. The time slot allocation and power control method for the body area network according to claim 1, wherein: Before the step of obtaining the time slot allocation and power control results of each sensor node in the current superframe according to the predicted path loss sequence, the age-of-information data, and the trained time slot allocation and power control model, it includes: Obtain the current network status data of each sensor node in the body area network; the current network status data includes the current age of information, the first average age of information, and the predicted path loss of data transmission between the sensor node and the sink node within the current superframe; the current age of information is the age of information of data transmission between the sensor node and the sink node within the current superframe, and the first average age of information is the average of the age of information of data transmission between the sensor node and the sink node in several historical superframes before the current superframe. Input the current network status data into the time slot allocation and power control model to be trained, and obtain the current time slot allocation and power control results of each sensor node. Obtain the next network status data of the current network status data according to the current time slot allocation and power control results. Obtain the reward value according to the first average age of information in the current network status data, the first average age of information of the previous network status data of the current network status data, and the current time slot allocation and power control results. Store the current network status data, the current time slot allocation and power control results, the next network status data, and the reward value as a set of sample data in the experience pool. When the number of groups of sample data in the experience pool is greater than or equal to the preset quantity, train the time slot allocation and power control model to be trained according to the sample data in the experience pool, and obtain the trained time slot allocation and power control model.

5. The time slot allocation and power control method for the body area network according to claim 4, wherein: The step of obtaining the reward value according to the first average age of information in the current network status data, the first average age of information of the previous network status data of the current network status data, and the current time slot allocation and power control results includes: Calculate the difference between the first average age of information of each sensor node in the current network status data and the first average age of information of the corresponding sensor node in the previous network status data of the current network status data to obtain several differences. Sum the several differences to obtain the first summation result; take the opposite of the first summation result to obtain the first reward. Obtain the current transmission power from the current time slot allocation and power control results; calculate the first difference between the current transmission power and the first preset transmission power, and the second difference between the second preset transmission power and the first preset transmission power; divide the first difference by the second difference to obtain the second reward. Calculate the third difference between the second preset transmission power and the current transmission power; divide the third difference by the second difference to obtain a ratio; obtain the third reward according to the ratio, the first average age of information in the current network status data, and the preset reward function. Obtain the reward value according to the first reward, the second reward, and the third reward.

6. The time slot allocation and power control method for the body area network according to claim 4, wherein: The to-be-trained time slot allocation and power control model includes a first deep Q-network and a second deep Q-network; The step of training the to-be-trained time slot allocation and power control model according to the sample data in the experience pool to obtain a trained time slot allocation and power control model when the number of groups of sample data in the experience pool is greater than or equal to a preset number includes: When the number of groups of sample data in the experience pool is greater than or equal to a preset number, extract a plurality of sample data from the experience pool, input the plurality of sample data into the first deep Q-network to obtain a first predicted Q value; input the plurality of sample data into the second deep Q-network to obtain a second predicted Q value; obtain a target Q value according to the second predicted Q value and the reward value in the sample data; Obtain a second loss function value according to the first predicted Q value, the target Q value, and a preset second loss function; Train the first deep Q-network according to the second loss function value until the second loss function value meets a preset threshold, and use the trained first deep Q-network as the trained time slot allocation and power control model; wherein, when the network parameters of the first deep Q-network are updated a preset number of times, copy the network parameters of the first deep Q-network to the second deep Q-network.

7. The time slot allocation and power control method for a body area network according to any one of claims 1 to 6, characterized in that: The step of obtaining the time slot allocation and power control results of each sensor node in the current superframe according to the predicted path loss sequence, the age of information data, and the trained time slot allocation and power control model includes: Obtain the first preset transmission time slots allocated to each sensor node in the current superframe and the second preset transmission time slots allocated to each sensor node in several historical superframes before the current superframe; Extract the corresponding predicted path loss from the predicted path loss sequence according to the first preset transmission time slot; Obtain the age of information of each sensor node transmitting data through the second preset transmission time slot in several historical superframes before the current superframe and the age of information of each sensor node transmitting data through the first preset transmission time slot in the current superframe; Calculate the average value of the age of information of each sensor node transmitting data through the second preset transmission time slot in several historical superframes before the current superframe and the age of information of each sensor node transmitting data through the first preset transmission time slot in the current superframe to obtain a second average age of information; Input the corresponding predicted path loss, the age of information of each sensor node transmitting data in the current superframe, and the second average age of information into the trained time slot allocation and power control model to obtain the time slot allocation and power control results of each sensor node in the current superframe.

8. A time slot allocation and power control device for a body area network, characterized in that, Including: An age of information data acquisition module, configured to acquire historical path loss data and age of information data of a body area network; The historical path loss data includes the true path losses measured when each sensor node transmits data to the sink node in a number of historical superframes before the current superframe, and the age-of-information data includes the age of information when each sensor node transmits data to the sink node in a number of historical superframes before the current superframe, and the age of information when each sensor node transmits data to the sink node in the current superframe; A predicted path loss sequence obtaining module, configured to obtain a predicted path loss sequence when each sensor node transmits data to the sink node in the current superframe according to the historical path loss data and a trained path loss prediction model; the predicted path loss sequence includes the predicted path losses for each sensor node to transmit data to the sink node through each transmission time slot; A time slot allocation and power control result obtaining module, configured to obtain the time slot allocation and power control results for each sensor node in the current superframe according to the predicted path loss sequence, the age-of-information data, and a trained time slot allocation and power control model.

9. An electronic device, characterized in that, Comprising: A processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the time slot allocation and power control method for the body area network according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the time slot allocation and power control method for the body area network according to any one of claims 1 to 7.