An Optimization Method for Energy-Capturing Mobile Heterogeneous Networks Based on Information Age Constraints

By optimizing mobile heterogeneous networks using digital twin technology and deep reinforcement learning frameworks, the problems of device energy limitations and information age constraints are solved, enabling efficient mode selection and resource allocation, and improving network performance and device energy supply.

CN119815373BActive Publication Date: 2025-10-31ZHEJIANG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510024111.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-10-31
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

In mobile heterogeneous networks, device energy limitations and information age constraints lead to network performance degradation, making it difficult to achieve efficient mode selection and resource allocation to meet long-term total throughput and device freshness requirements.

Method used

By employing digital twin technology and a deep reinforcement learning framework, and through the exchange and acquisition of state information, the mode selection and resource allocation are optimized. By combining the first and second deep reinforcement learning networks to coordinate the actions of the agent, reasonable power control and wireless energy capture time are achieved.

Benefits of technology

It improved the network's long-term total throughput while meeting device information age constraints, thus enhancing network performance and device energy supply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119815373B_ABST
    Figure CN119815373B_ABST
Patent Text Reader

Abstract

This invention discloses an optimization method for energy-harvesting mobile heterogeneous networks based on information age constraints. For the physical entity layer, including base stations, hybrid access points, transmitters, and receivers, a corresponding digital twin layer is established. During the time slot control phase, a trained first deep reinforcement learning network obtains transmitter mode selection and channel allocation strategies, while a trained second deep reinforcement learning network obtains transmitter power control and energy harvesting time strategies. These are then sent to the digital twin layer, which merges them into a complete mode selection and resource allocation strategy and synchronously provides the transmitter with the corresponding operational information. This invention achieves more efficient state information exchange and acquisition, coordinates the actions of agents in deep reinforcement learning, reduces competition for priority network resources, and optimizes power control and wireless energy harvesting time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of wireless transmission technology, and in particular relates to an energy-capturing mobile heterogeneous network optimization method based on information age constraints. Background Technology

[0002] With the rapid development of the Internet of Things (IoT) and communication technologies, the number of devices and data traffic are exploding in emerging applications and service scenarios such as smart cities and autonomous driving. The energy limitations of these devices are significant; when a large number of devices are deployed in locations unsuitable for battery charging and replacement, their limited lifespan degrades network performance. To meet the ever-increasing data traffic demands of emerging wireless communication services while providing high-quality services to these devices, new technologies are needed to improve network performance.

[0003] To improve network performance, a heterogeneous network architecture can be employed. This involves deploying multiple small base stations within the coverage area of ​​a macro base station, further enhancing data throughput and network performance. Furthermore, to increase device power supply, all devices utilize radio frequency (RF) signal-based wireless power transfer technology, allowing devices in the network to capture energy from RF signals and store it in batteries. Additionally, device-to-device (D2D) communication enables nearby devices to communicate directly without relays. Due to the shorter transmission distance, devices can use lower transmission power, reducing interference to other devices and improving network performance.

[0004] On the other hand, real-time applications and services in the network have strict requirements on data latency. Data arriving at the destination must be as fresh as possible. Information Age (AoI) measures the freshness of data and is defined as the time elapsed since the most recently received data at the destination was generated. Therefore, efficient mode selection and resource allocation are crucial to improving the network's long-term overall throughput while meeting the AoI constraints of devices.

[0005] To address the challenges of long-term performance optimization and device mobility in energy-capturing mobile heterogeneous networks, while simultaneously meeting device information age constraints, efficient mode selection and resource allocation are crucial. In dynamic and complex network scenarios with limited network resources, achieving efficient mode selection and resource allocation to improve long-term overall network throughput while satisfying device AoI constraints is of great significance. Summary of the Invention

[0006] The purpose of this application is to provide an optimization method for energy-capturing mobile heterogeneous networks based on information age constraints, enabling devices to rationally select modes and utilize limited network resources in dynamic and complex networks to improve the long-term total throughput of energy-capturing mobile heterogeneous networks, while satisfying the device information age (AoI) constraint.

[0007] To achieve the above objectives, the technical solution of this application is as follows:

[0008] An optimization method for energy-capturing mobile heterogeneous networks based on information age constraints includes:

[0009] For the physical entity layer, including base stations, hybrid access points, transmitters and receivers, establish corresponding digital twin layers;

[0010] During the time slot control phase, the transmitter and receiver move, and the transmitter inputs local state information into the trained first deep reinforcement learning network to obtain the corresponding mode selection and channel allocation strategy, which is then sent to the digital twin layer.

[0011] The base station inputs global state information and transmitter mode selection and channel allocation strategies into a trained second deep reinforcement learning network to obtain transmitter power control and energy capture time strategies, which are then sent to the digital twin layer.

[0012] The digital twin layer merges the transmitter's mode selection and channel allocation strategies with the transmitter's power control and energy capture time strategies into a complete mode selection and resource allocation strategy, and synchronizes it with the transmitter.

[0013] During the wireless power transfer phase, the transmitter captures energy based on the energy capture time. During the wireless data transmission phase, the transmitter transmits data based on mode selection, channel allocation, and power control.

[0014] Furthermore, the digital twin layer is deployed at the base station.

[0015] Furthermore, the local state information includes:

[0016] The coordinates of the transmitter after it moves in the current time slot, the coordinates of the receiver after it moves in the current time slot, the ratio of the transmitter's available energy to its battery capacity at the start of the current time slot, the information age of the transmitter in the current time slot, the binary variable indicating whether the transmitter in the current time slot generates data, the channel gain between the transmitter after it moves in the current time slot and its associated base station or hybrid access point, the channel gain between the transmitter after it moves in the current time slot and its moved receiver, the throughput of the transmitter in the previous time slot, and the channel allocation of the transmitter in the previous time slot.

[0017] Furthermore, the global state information includes:

[0018] The ratio of the transmitter's available energy to its battery capacity at the start of the current time slot, the information age of the transmitter in the current time slot, and the throughput of the transmitter in the previous time slot.

[0019] Furthermore, the energy-harvesting mobile heterogeneous network optimization method based on information age constraints also includes:

[0020] Train the first deep reinforcement learning network and the second deep reinforcement learning network;

[0021] The training of the first deep reinforcement learning network and the second deep reinforcement learning network includes:

[0022] Step F1: Initialize the local Q-network and target Q-network corresponding to the first deep reinforcement learning network, and initialize the parameters of the current action network, target action network, first current value network, first target value network, second current value network and second target value network corresponding to the second deep reinforcement learning network;

[0023] Step F2: Initialize the hybrid access point, transmitter, and receiver;

[0024] Step F3: The transmitter and receiver move to update the local state information of the current time slot in the digital twin layer;

[0025] Step F4: The transmitter sends the local status information to the local Q network to obtain the corresponding mode selection and channel allocation strategy, and then sends it to the digital twin layer;

[0026] Step F5: The base station receives the global state information and transmitter mode selection and channel allocation strategy from the digital twin layer, inputs the current action network, obtains the transmitter power control and energy capture time strategy, and sends it to the digital twin layer.

[0027] Step F6: The digital twin layer merges the transmitter's mode selection and channel allocation strategy with the transmitter's power control and energy capture time strategy into a complete mode selection and resource allocation strategy, and synchronizes it with the transmitter.

[0028] Step F7: The transmitter executes the complete mode selection and resource allocation strategy;

[0029] Step F8: The transmitter and base station calculate the corresponding rewards;

[0030] Step F9: The transmitter and base station store their own experience into their own experience replay pool and randomly extract a preset number of experience samples.

[0031] Step F10: The transmitter updates the parameters of the local Q network based on the extracted experience;

[0032] Step F11: If the result of taking the current iteration number modulo 2 is not equal to 0, no operation is performed; otherwise, the parameters of the target Q network are updated using the parameters of the local Q network.

[0033] Step F12: The base station updates the parameters of the first current value network and the second current value network based on the extracted experience;

[0034] Step F13: If the result of taking the current iteration number modulo 2 is not equal to 0, no operation is performed; otherwise, the parameters of the current action network and all target networks are updated based on the extracted experience.

[0035] Step F14: If the current time slot is less than the total number of time slots, then use the next time slot as the current time slot and return to step F3;

[0036] Step F15: If the current iteration count is less than the total iteration count, return to step F2 to continue iterative training; otherwise, end the training.

[0037] Furthermore, the transmitter's own experience includes the transmitter's local state information, mode selection and channel allocation strategies, the rewards received by the transmitter, and the local state information of the transmitter in the next time slot.

[0038] Furthermore, the base station's own experience includes global state information, transmitter mode selection and channel allocation strategies, transmitter power control and energy capture time strategies, rewards received by the base station, global state information for the next time slot, and transmitter mode selection and channel allocation strategies.

[0039] The energy-harvesting mobile heterogeneous network optimization method based on information age constraints proposed in this application has the following significant advantages compared with existing technologies:

[0040] This application integrates digital twin technology into the state information exchange and acquisition process during the training of a complex deep reinforcement learning framework. Twin communication between the digital twin layer and the physical entity layer enables more efficient state information exchange and acquisition, as well as the implementation of mode selection and resource allocation within the physical entity layer. Simultaneously, to further improve the performance of deep reinforcement learning for action outputs and avoid the convergence problem caused by the large state-action space, this application employs a first deep reinforcement learning approach to optimize mode selection and channel allocation. However, the first deep reinforcement learning approach leads to competition among agents for limited network resources. To coordinate the actions of the agents, this application incorporates a second deep reinforcement learning approach. By using global state information and receiving the output of the first deep reinforcement learning approach as its own state information, this approach coordinates the actions of the agents to reduce competition, while also optimizing power control and wireless energy capture time.

[0041] This application achieves reasonable pattern selection and resource allocation strategies through deep reinforcement learning, and improves the long-term total throughput of the network through training, while satisfying device AoI constraints. Attached Figure Description

[0042] Figure 1 This is a flowchart of the energy-capturing mobile heterogeneous network optimization method based on information age constraints in this application.

[0043] Figure 2 This is a schematic diagram of the energy-capturing mobile heterogeneous network structure of this application.

[0044] Figure 3 This is a schematic diagram of the time slot structure of this application. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0046] In one embodiment, such as Figure 1 As shown, an energy-capturing mobile heterogeneous network optimization method based on information age constraints is provided. Deep reinforcement learning is used to enable devices to rationally select modes and utilize limited network resources in a dynamic and complex network to maximize the long-term throughput that meets the device's AoI constraints.

[0047] The energy-harvesting mobile heterogeneous network optimization method based on information age constraints includes:

[0048] Step S1: Establish a corresponding digital twin layer for the physical entity layer, which includes base stations, hybrid access points, transmitters, and receivers.

[0049] The energy-capturing mobile heterogeneous network optimization method based on information age constraints provided in this application can be applied to, for example... Figure 2 In the application environment shown, the energy-harvesting mobile heterogeneous network comprises a physical entity layer and a digital twin layer, wherein the physical entity layer includes a multi-antenna base station, Multiple antenna hybrid access points and A D2D device pair (D2D pair for short) consists of a D2D transmitter (MDT) and a D2D receiver (MDR). The base station and hybrid access point transmit radio frequency signals for the transmitter to capture energy. The transmitter can choose between relay mode or D2D mode to transmit data to the corresponding receiver. The wireless channel can be divided into relay mode and D2D mode. A channel with equal bandwidth.

[0050] The digital twin layer establishes a mapping between physical entities in the physical entity layer and their corresponding digital twins in the digital twin layer. It is used to store the state information of physical entities and exchange real-time state information. The digital twin layer resides in the digital twin server and is deployed at the base station. Figure 2 The digital twin server is located next to the base station.

[0051] Each transmitter in the network is equipped with a data buffer and generates data with a certain probability in each time slot. If a transmitter generates data in the current time slot, the data is stored in the data buffer, and previously untransmitted data is discarded. Otherwise, the untransmitted data remains stored in the data buffer. When data transmission in the data buffer is successful, the transmitter's AoI decreases to the number of time slots elapsed since the successful data transmission, and the throughput achieved by the transmitter is included in the total network throughput of the current time slot. When data transmission fails, the transmitter's AoI increases by 1, and the throughput achieved by the transmitter is not included in the total network throughput of the current time slot.

[0052] Step S2: During the time slot control phase, the transmitter and receiver move. The transmitter inputs the local state information into the trained first deep reinforcement learning network to obtain the corresponding mode selection and channel allocation strategy, and sends it to the digital twin layer.

[0053] like Figure 3 As shown, each time slot includes a control phase t0, a wireless power transfer phase t1(k), and a wireless data transmission phase t2(k).

[0054] In the network, all hybrid access points, transmitters, and receivers follow a static Poisson point distribution. The transmitter and receiver move according to a Gauss-Markov random model during the control phase.

[0055] In the time slot ,transmitter speed and direction Updated to:

[0056]

[0057]

[0058] in It is a constant used to adjust the effect of the previous time slot velocity on the current time slot velocity. It is a constant used to adjust the effect of the previous time slot direction on the current time slot direction. It is the average speed of the transmitter and receiver. It is a transmitter The average direction, Follows the mean-variance pair Independent Gaussian distribution, Follows the mean-variance pair , independent Gaussian distribution;

[0059] In the time slot ,transmitter Position coordinates Updated to:

[0060]

[0061]

[0062] in It is the travel time of the transmitter and receiver during the control phase.

[0063] The digital twin layer and the physical entity layer communicate via twins, and the state information stored in the digital twin corresponding to the physical entity in the digital twin layer is updated. Each transmitter obtains the corresponding mode selection and channel allocation strategy by inputting its local state information into the trained first deep reinforcement learning network, and then sends it to the digital twin layer via twin communication.

[0064] In the time slot ,transmitter Local state information is defined as follows:

[0065]

[0066]

[0067] in and It is a time slot transmitter The coordinates of the position after the movement; and It is a time slot Receiver Position coordinates; It is a time slot Initial transmitter The ratio of available energy to battery capacity. It is a time slot transmitter AoI; It is a time slot transmitter Whether to generate binary variables of the data Indicates time slot transmitter Generate data, Indicates time slot transmitter No data is generated; It is a time slot The transmitter after it was moved Channel gain between the associated base station or hybrid access point, It is a time slot The transmitter after it was moved With the receiver after the move Channel gain between; It is a time slot Transmitter throughput It is a time slot Channel allocation for the transmitter.

[0068] In the time slot ,transmitter The mode selection and channel allocation strategy is defined as follows:

[0069]

[0070] in It is a time slot transmitter mode selection, Indicates time slot transmitter Select D2D mode. Indicates time slot transmitter Select relay mode. It is a time slot transmitter The assigned channel number.

[0071] Step S3: The base station inputs the global state information and the transmitter's mode selection and channel allocation strategy into the trained second deep reinforcement learning network to obtain the transmitter's power control and energy capture time strategy, and sends it to the digital twin layer.

[0072] The transmitter and receiver remain stationary during the wireless power transfer and wireless data transmission phases of each time slot.

[0073] The base station will transmit global status information and The mode selection and channel allocation strategies of each transmitter are input into a pre-trained second deep reinforcement learning network to obtain... The power control and energy capture time strategy of each transmitter is sent to the digital twin layer via twin communication.

[0074] In the time slot Global status information and The mode selection and channel allocation strategies for each transmitter are defined as follows:

[0075]

[0076] in It is a time slot The channel gain (corresponding to the mode selection) between the transmitter after the move and the receiver determined by the mode selection obtained from the first deep reinforcement learning network. It is a time slot The channel allocation (corresponding channel allocation strategy) is obtained based on the first deep reinforcement learning network. It is a time slot The initial ratio of the transmitter's available energy to the battery capacity. It is a time slot Transmitter's AoI It is a time slot Transmitter throughput.

[0077] Step S4: The digital twin layer merges the transmitter's mode selection and channel allocation strategy with the transmitter's power control and energy capture time strategy into a complete mode selection and resource allocation strategy, and synchronizes it with the transmitter.

[0078] The digital twin layer will The mode selection and channel allocation strategies of each transmitter and The power control and energy capture time strategies of each transmitter are merged into a complete mode selection and resource allocation strategy, and synchronized with... One transmitter.

[0079] In the time slot , The mode selection and channel allocation strategies for each transmitter are defined as follows:

[0080]

[0081] In the time slot , The power control and energy capture time strategy for each transmitter is defined as follows:

[0082]

[0083] in It is a time slot transmitter Energy and time slots consumed by data transmission The transmitter begins the wireless data transmission phase. The ratio of available energy, It is a time slot The ratio of transmitter energy capture time to the sum of the wireless energy transfer phase and the wireless data transfer phase.

[0084] In the time slot The digital twin layer will The mode selection and channel allocation strategies of each transmitter and The power control and energy capture time strategies of each transmitter are combined into a complete mode selection and resource allocation strategy, defined as:

[0085] .

[0086] Step S5: In the wireless power transfer phase, the transmitter performs power acquisition based on the power acquisition time. In the wireless data transmission phase, the transmitter performs data transmission based on mode selection, channel allocation, and power control.

[0087] The transmitter and receiver remain stationary during the wireless power transfer and wireless data transmission phases of each time slot.

[0088] During the wireless power transfer phase, Each transmitter performs energy capture based on the energy capture time obtained from the complete mode selection and resource allocation strategy;

[0089] During the wireless data transmission phase, Each transmitter performs corresponding data transmission based on the mode selection, channel allocation, and power control derived from the complete mode selection and resource allocation strategy.

[0090] In another embodiment, the training process for deep reinforcement learning is described, namely, training a first deep reinforcement learning and a second deep reinforcement learning, the training process including the following steps:

[0091] Step F1: Initialize the local Q-network and target Q-network corresponding to the first deep reinforcement learning network, and initialize the parameters of the current action network, target action network, first current value network, first target value network, second current value network and second target value network corresponding to the second deep reinforcement learning network.

[0092] Specifically, initialize the network parameters for each agent in the first deep reinforcement learning network, including the following parameters: The local Q network and parameters are The target Q-network is initialized with parameters for the second deep reinforcement learning network, including the following parameters: The current action network, parameters are The target action network, with parameters as follows The first current value network, parameters are The first target value network, with parameters as follows: The second current value network and parameters are The second objective is the value network, which initializes the total number of training iterations. Initialize the total number of time slots for each round. Set the iteration rounds .

[0093] Step F2: Initialize the hybrid access point, transmitter, and receiver.

[0094] Initialize the hybrid access points in the physical entity layer, including transmitter and mobile receiver distributions, with transmitter battery available energy at 0, throughput at 0, AoI at 1, and probability-based data generation. Set the time slots in the iteration rounds. .

[0095] In step F3, the transmitter and receiver move to update the local state information of the current time slot in the digital twin layer.

[0096] The transmitter and receiver move according to the Gauss-Markov random model, and the digital twin layer and the physical entity layer communicate in twin form. The state information stored in the digital twin corresponding to the physical entity in the digital twin layer is updated.

[0097] Each transmitter receives local state information of the current time slot stored in its own digital twin from the digital twin layer. .

[0098] Step F4: The transmitter sends the local status information to the local Q network to obtain the corresponding mode selection and channel allocation strategy, and then sends it to the digital twin layer.

[0099] Each transmitter will As input to the local Q network, the corresponding action is output. This refers to the mode selection and channel allocation strategy, which is sent to the digital twin layer via twin communication. The digital twin layer stores this information. The mode selection and channel allocation strategy for each transmitter.

[0100] Step F5: The base station receives the global state information and transmitter mode selection and channel allocation strategies from the digital twin layer, inputs them into the current action network, obtains the transmitter power control and energy capture time strategies, and sends them to the digital twin layer.

[0101] The base station receives the global state information stored in the corresponding digital twin in the digital twin layer and Transmitter mode selection and channel allocation strategies The base station will As the input to the current action network, the output action is... This refers to the transmitter's power control and energy capture time strategy, which is transmitted to the digital twin layer via twin communication.

[0102] Step F6: The digital twin layer merges the transmitter's mode selection and channel allocation strategy with the transmitter's power control and energy capture time strategy into a complete mode selection and resource allocation strategy, and synchronizes it with the transmitter.

[0103] The digital twin layer will and Merging into a complete pattern selection and resource allocation strategy And simultaneously give One transmitter.

[0104] Step F7: The transmitter executes the complete mode selection and resource allocation strategy.

[0105] One transmitter received Next, data analysis is performed to obtain relevant pattern selection and resource allocation strategies. Each transmitter according to The energy capture time in the time slot is first in the time slot. Energy capture is performed during the wireless power transfer phase. In time slots During the wireless data transmission phase, data is transmitted according to the strategy's mode selection, channel allocation, and power control.

[0106] Step F8: The transmitter and base station calculate the corresponding rewards.

[0107] Each transmitter and base station in the time slot During the control phase, the transmitter obtains relevant status information through twin communication. According to the time slot The achieved throughput and the corresponding reward for AoI calculation ,in The calculation formula is:

[0108]

[0109] in It is throughput weight. It is a time slot transmitter The throughput achieved It is the AoI weight. It is a time slot transmitter Penalties related to AoI. (When the time slot) transmitter When the AoI is less than the maximum tolerance threshold, ,otherwise, ,in Positive constants of balanced magnitude Maximum tolerance threshold; base station according to The reward calculation for each transmitter corresponds to the reward. ,in The calculation formula is:

[0110] .

[0111] In step F9, the transmitter and base station respectively store their own experience into their own experience replay pool and randomly extract a preset number of experience samples.

[0112] Each transmitter will have a time slot My own experience Stored in its own experience replay pool and randomly drawn This experience, among which, It is a time slot transmitter Local status information, It is a time slot for The mode selection and channel allocation strategy, It is a time slot transmitter The reward received It is a time slot transmitter Local status information.

[0113] Base station timeslots experience Stored in its own experience replay pool and randomly drawn This experience, among which, It is a time slot Global status information and The mode selection and channel allocation strategies for each transmitter. It is a time slot for of Power control and energy capture time strategies for each transmitter It is a time slot The rewards received by the base station It is a time slot Global status information and The mode selection and channel allocation strategy for each transmitter.

[0114] Step F10: The transmitter updates the parameters of the local Q network based on the extracted experience.

[0115] Each transmitter will randomly select... The local Q network parameters are updated based on the experience according to formulas (1)-(3).

[0116]

[0117]

[0118]

[0119] in It is a transmitter experience The target Q value for the corresponding time slot, It is a discount factor. It is to make experience Transmitter in the next time slot corresponding to the time slot The action with the highest estimated Q-value in the Q-network. It is a transmitter experience The state information of the next time slot corresponding to the time slot and the estimated Q value output by the target Q-network corresponding to the action. It is a transmitter The loss function of the Q-network, It's the learning rate. express right Differentiate, Indicates transmitter Randomly selected The first experience One experience.

[0120] Step F11: If the result of taking the current iteration number modulo 2 is not equal to 0, no operation is performed; otherwise, the parameters of the target Q network are updated using the parameters of the local Q network.

[0121] If the current iteration number If the result of modulo 2 is not equal to 0, no operation is performed; otherwise, the transmitter... Update the target Q network parameters to .

[0122] Step F12: The base station updates the parameters of the first current value network and the second current value network based on the extracted experience.

[0123] The base station randomly selects The current value network parameters are updated based on the experience according to formulas (4)-(8).

[0124]

[0125]

[0126]

[0127]

[0128]

[0129] in Base station experience The smaller value of the one-step time difference error of the first and second current value networks corresponding to the time slot. It is experience The next time slot status information corresponding to the time slot Action value of the target value network (j equals 1 corresponds to the first target value network, j equals 2 corresponds to the second target value network). and These are the loss functions for the first and second current value networks, respectively. It is experience The state information corresponding to the time slot is the action value of the first current value network. It is experience The state information corresponding to the time slot is the action value of the second current value network. express right Differentiate, express right Differentiate, This indicates that the base station randomly selected... The first experience One experience.

[0130] Step F13: If the result of taking the current iteration number modulo 2 is not equal to 0, no operation is performed; otherwise, the parameters of the current action network and all target networks are updated based on the extracted experience.

[0131] If the current iteration number If the result of modulo 2 is not equal to 0, no operation is performed; otherwise, the selected base station... The experience is used to update the parameters of the current action network and all target networks according to formulas (9)-(13).

[0132]

[0133]

[0134]

[0135]

[0136]

[0137] in It is a soft update factor in action networks. It is a soft update factor for the value network.

[0138] Step F14: If the current time slot is less than the total number of time slots, then the next time slot is taken as the current time slot, and the process returns to step F3.

[0139] If the current time slot K represents the total number of time slots, then Then proceed to step F3.

[0140] Step F15: If the current iteration count is less than the total iteration count, return to step F2 to continue iterative training; otherwise, end the training.

[0141] like and One iteration ends, And jump to step F2; if and Training is over.

[0142] It should be noted that after training, the parameters of the local Q-network are the same as the parameters of the trained first deep reinforcement learning network, and the parameters of the current action network are the same as the parameters of the second deep reinforcement learning network, which are used for energy capture mobile heterogeneous network optimization based on information age constraints.

[0143] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. An optimization method for energy-harvesting mobile heterogeneous networks based on information age constraints, characterized in that, The energy-harvesting mobile heterogeneous network optimization method based on information age constraints includes: For the physical entity layer, including base stations, hybrid access points, transmitters and receivers, establish corresponding digital twin layers; During the time slot control phase, the transmitter and receiver move, and the transmitter inputs local state information into the trained first deep reinforcement learning network to obtain the corresponding mode selection and channel allocation strategy, which is then sent to the digital twin layer. The base station inputs global state information and transmitter mode selection and channel allocation strategies into a trained second deep reinforcement learning network to obtain transmitter power control and energy capture time strategies, which are then sent to the digital twin layer. The digital twin layer merges the transmitter's mode selection and channel allocation strategies with the transmitter's power control and energy capture time strategies into a complete mode selection and resource allocation strategy, and synchronizes it with the transmitter. During the wireless power transfer phase, the transmitter captures energy based on the energy capture time. During the wireless data transmission phase, the transmitter transmits data based on mode selection, channel allocation, and power control. The local status information includes: The coordinates of the transmitter after it moves in the current time slot, the coordinates of the receiver after it moves in the current time slot, the ratio of the transmitter's available energy to its battery capacity at the start of the current time slot, the information age of the transmitter in the current time slot, the binary variable indicating whether the transmitter in the current time slot generates data, the channel gain between the transmitter after it moves in the current time slot and its associated base station or hybrid access point, the channel gain between the transmitter after it moves in the current time slot and its moved receiver, the throughput of the transmitter in the previous time slot, and the channel allocation of the transmitter in the previous time slot. The global state information includes: The ratio of the transmitter's available energy to its battery capacity at the start of the current time slot, the information age of the transmitter in the current time slot, and the throughput of the transmitter in the previous time slot.

2. The energy-harvesting mobile heterogeneous network optimization method based on information age constraints according to claim 1, characterized in that, The digital twin layer is deployed at the base station.

3. The energy-harvesting mobile heterogeneous network optimization method based on information age constraints according to claim 1, characterized in that, The energy-harvesting mobile heterogeneous network optimization method based on information age constraints also includes: Train the first deep reinforcement learning network and the second deep reinforcement learning network; The training of the first deep reinforcement learning network and the second deep reinforcement learning network includes: Step F1: Initialize the local Q-network and target Q-network corresponding to the first deep reinforcement learning network, and initialize the parameters of the current action network, target action network, first current value network, first target value network, second current value network and second target value network corresponding to the second deep reinforcement learning network; Step F2: Initialize the hybrid access point, transmitter, and receiver; Step F3: The transmitter and receiver move to update the local state information of the current time slot in the digital twin layer; Step F4: The transmitter sends the local status information to the local Q network to obtain the corresponding mode selection and channel allocation strategy, and then sends it to the digital twin layer; Step F5: The base station receives the global state information and transmitter mode selection and channel allocation strategy from the digital twin layer, inputs the current action network, obtains the transmitter power control and energy capture time strategy, and sends it to the digital twin layer. Step F6: The digital twin layer merges the transmitter's mode selection and channel allocation strategy with the transmitter's power control and energy capture time strategy into a complete mode selection and resource allocation strategy, and synchronizes it with the transmitter. Step F7: The transmitter executes the complete mode selection and resource allocation strategy; Step F8: The transmitter and base station calculate the corresponding rewards; Step F9: The transmitter and base station store their own experience into their own experience replay pool and randomly extract a preset number of experience samples. Step F10: The transmitter updates the parameters of the local Q network based on the extracted experience; Step F11: If the result of taking the current iteration number modulo 2 is not equal to 0, no operation is performed; otherwise, the parameters of the target Q network are updated using the parameters of the local Q network. Step F12: The base station updates the parameters of the first current value network and the second current value network based on the extracted experience; Step F13: If the result of taking the current iteration number modulo 2 is not equal to 0, no operation is performed; otherwise, the parameters of the current action network and all target networks are updated based on the extracted experience. Step F14: If the current time slot is less than the total number of time slots, then use the next time slot as the current time slot and return to step F3; Step F15: If the current iteration count is less than the total iteration count, return to step F2 to continue iterative training; otherwise, end the training.

4. The energy-harvesting mobile heterogeneous network optimization method based on information age constraints according to claim 3, characterized in that, The transmitter's own experience includes the transmitter's local state information, mode selection and channel allocation strategies, the rewards received by the transmitter, and the local state information of the transmitter in the next time slot.

5. The energy-harvesting mobile heterogeneous network optimization method based on information age constraints according to claim 3, characterized in that, The base station's own experience includes global state information, transmitter mode selection and channel allocation strategies, transmitter power control and energy capture time strategies, rewards received by the base station, global state information for the next time slot, and transmitter mode selection and channel allocation strategies.

Citation Information

Patent Citations

  • Mobile wireless energy supply Internet of Things resource allocation method based on deep reinforcement learning

    CN116684964A